Package evaluation to test TextSearch on Julia 1.14.0-DEV.2802 (918269a29b*) started at 2026-08-28T00:12:25.957 ################################################################################ # Set-up # Installing PkgEval dependencies (TestEnv)... Activating project at `~/.julia/environments/v1.14` Set-up completed after 14.35s ################################################################################ # Installation # Installing TextSearch... Resolving package versions... Installed TextSearch ─ v1.1.1 Updating `~/.julia/environments/v1.14/Project.toml` [7f6f6c8a] + TextSearch v1.1.1 Updating `~/.julia/environments/v1.14/Manifest.toml` [621f4979] + AbstractFFTs v1.5.0 [7d9f7c33] + Accessors v0.1.45 [66dad0bd] + AliasTables v1.1.3 [dce04be8] + ArgCheck v2.5.0 [7d9fca2a] + Arpack v0.5.4 [6309b1aa] + CodecInflate64 v0.1.3 [944b1d66] + CodecZlib v0.7.9 [38540f10] + CommonSolve v0.2.14 [a33af91c] + CompositionsBase v0.1.2 [187b0558] + ConstructionBase v1.6.0 [9a962f9c] + DataAPI v1.16.0 [864edb3b] + DataStructures v0.19.6 [b4f34e82] + Distances v0.10.12 [31c24e10] + Distributions v0.25.131 [ffbed154] + DocStringExtensions v0.9.5 [7a1cc6ca] + FFTW v1.10.0 [1a297f60] + FillArrays v1.17.0 [a0844989] + Gamma v1.2.0 [4a05ff16] + Hadamard v1.8.0 [34004b35] + HypergeometricFunctions v0.3.30 [0c81fc1b] + InputBuffers v1.1.1 [3587e190] + InverseFunctions v0.1.17 [92d709cd] + IrrationalConstants v0.2.6 [692b3bcd] + JLLWrappers v1.8.0 [682c06a0] + JSON v1.7.1 [0f8b85d8] + JSON3 v1.14.3 [2ab3a3ac] + LogExpFunctions v1.0.1 [1914dd2f] + MacroTools v0.5.16 [e1d29d7a] + Missings v1.2.0 [bac558e1] + OrderedCollections v2.0.1 [90014a1f] + PDMats v0.11.41 ⌅ [69de0a69] + Parsers v2.8.7 [aea7be01] + PrecompileTools v1.3.4 [21216c6a] + Preferences v1.5.2 [92933f4c] + ProgressMeter v1.11.0 [43287f4e] + PtrArrays v1.4.0 [1fd47b50] + QuadGK v2.11.3 [189a3867] + Reexport v1.2.2 [79098fc4] + Rmath v0.9.0 [f2b01f46] + Roots v3.0.7 [fdea26ae] + SIMD v3.7.2 [0e966ebe] + SearchModels v0.5.1 [053f045d] + SimilaritySearch v1.2.0 [a2af1166] + SortingAlgorithms v1.2.3 [276daf66] + SpecialFunctions v2.9.0 [10745b16] + Statistics v1.11.1 [82ae8749] + StatsAPI v1.8.0 [2913bbd2] + StatsBase v0.34.13 [4c63d2b9] + StatsFuns v2.2.1 [856f2bd8] + StructTypes v1.11.0 [ec057cc2] + StructUtils v2.8.5 [7f6f6c8a] + TextSearch v1.1.1 [3bb67fe8] + TranscodingStreams v0.11.3 [49080126] + ZipArchives v2.6.0 ⌅ [68821587] + Arpack_jll v3.5.2+0 [f5851436] + FFTW_jll v3.3.12+0 [1d5cc7b8] + IntelOpenMP_jll v2025.2.0+0 [856f044c] + MKL_jll v2025.2.0+0 [efe28fd5] + OpenSpecFun_jll v0.5.6+0 [f50d1b31] + Rmath_jll v0.5.2+0 [1317d2d5] + oneTBB_jll v2022.3.0+0 [0dad84c5] + ArgTools v1.2.0 [56f22d72] + Artifacts v1.11.0 [2a0f44e3] + Base64 v1.11.0 [ade2ca70] + Dates v1.11.0 [8ba89e20] + Distributed v1.11.0 [f43a241f] + Downloads v1.7.0 [7b1f6079] + FileWatching v1.11.0 [ac6e5ff7] + JuliaSyntaxHighlighting v1.13.0 [4af54fe1] + LazyArtifacts v1.11.0 [b27032c2] + LibCURL v1.0.0 [76f85450] + LibGit2 v1.11.0 [8f399da3] + Libdl v1.11.0 [37e2e46d] + LinearAlgebra v1.14.0 [56ddb016] + Logging v1.11.0 [d6f4376e] + Markdown v1.11.0 [a63ad114] + Mmap v1.11.0 [ca575930] + NetworkOptions v1.3.0 [44cfe95a] + Pkg v1.14.0 [de0858da] + Printf v1.11.0 [9a3f8284] + Random v1.11.0 [ea8e919c] + SHA v1.13.0 [9e88b42a] + Serialization v1.11.0 [6462fe0b] + Sockets v1.11.0 [2f01184e] + SparseArrays v1.13.0 [f489334b] + StyledStrings v1.13.0 [4607b0f0] + SuiteSparse [fa267f1f] + TOML v1.0.3 [a4e569a6] + Tar v1.10.0 [cf7118a7] + UUIDs v1.11.0 [4ec0a83e] + Unicode v1.11.0 [e66e0078] + CompilerSupportLibraries_jll v1.5.7+0 [deac9b47] + LibCURL_jll v8.21.0+0 [e37daf67] + LibGit2_jll v1.9.6+0 [29816b5a] + LibSSH2_jll v1.11.103+0 [14a3606d] + MozillaCACerts_jll v2026.7.16 [4536629a] + OpenBLAS_jll v0.3.34+0 [05823500] + OpenLibm_jll v0.8.7+0 [458c3c95] + OpenSSL_jll v3.5.7+0 [efcefdf7] + PCRE2_jll v10.47.0+0 [bea87d4a] + SuiteSparse_jll v7.10.1+0 [83775a58] + Zlib_jll v1.3.2+0 [3161d3a3] + Zstd_jll v1.5.7+1 [8e850b90] + libblastrampoline_jll v5.15.0+0 [8e850ede] + nghttp2_jll v1.69.0+0 [3f19e933] + p7zip_jll v17.8.0+0 Info Packages marked with ⌅ have new versions available but compatibility constraints restrict them from upgrading. To see why use `status --outdated -m` Installation completed after 6.45s ################################################################################ # Precompilation # Precompiling PkgEval dependencies... Precompiling package dependencies... Precompiling project... 11.4 s ✓ TextSearch 1 dependency successfully precompiled in 13 seconds. 116 already precompiled. Precompilation completed after 43.58s ################################################################################ # Testing # Testing TextSearch Status `/tmp/jl_2EIhZM/Project.toml` [7d9f7c33] Accessors v0.1.45 [4c88cf16] Aqua v0.8.16 [7d9fca2a] Arpack v0.5.4 [0f8b85d8] JSON3 v1.14.3 [92933f4c] ProgressMeter v1.11.0 [053f045d] SimilaritySearch v1.2.0 [2913bbd2] StatsBase v0.34.13 [7f6f6c8a] TextSearch v1.1.1 [49080126] ZipArchives v2.6.0 [f43a241f] Downloads v1.7.0 [37e2e46d] LinearAlgebra v1.14.0 [9a3f8284] Random v1.11.0 [2f01184e] SparseArrays v1.13.0 [8dfed614] Test v1.11.0 [4ec0a83e] Unicode v1.11.0 Status `/tmp/jl_2EIhZM/Manifest.toml` [621f4979] AbstractFFTs v1.5.0 [7d9f7c33] Accessors v0.1.45 [66dad0bd] AliasTables v1.1.3 [4c88cf16] Aqua v0.8.16 [dce04be8] ArgCheck v2.5.0 [7d9fca2a] Arpack v0.5.4 [6309b1aa] CodecInflate64 v0.1.3 [944b1d66] CodecZlib v0.7.9 [38540f10] CommonSolve v0.2.14 [34da2185] Compat v4.18.1 [a33af91c] CompositionsBase v0.1.2 [187b0558] ConstructionBase v1.6.0 [9a962f9c] DataAPI v1.16.0 [864edb3b] DataStructures v0.19.6 [b4f34e82] Distances v0.10.12 [31c24e10] Distributions v0.25.131 [ffbed154] DocStringExtensions v0.9.5 [7a1cc6ca] FFTW v1.10.0 [1a297f60] FillArrays v1.17.0 [a0844989] Gamma v1.2.0 [4a05ff16] Hadamard v1.8.0 [34004b35] HypergeometricFunctions v0.3.30 [0c81fc1b] InputBuffers v1.1.1 [3587e190] InverseFunctions v0.1.17 [92d709cd] IrrationalConstants v0.2.6 [692b3bcd] JLLWrappers v1.8.0 [682c06a0] JSON v1.7.1 [0f8b85d8] JSON3 v1.14.3 [2ab3a3ac] LogExpFunctions v1.0.1 [1914dd2f] MacroTools v0.5.16 [e1d29d7a] Missings v1.2.0 [bac558e1] OrderedCollections v2.0.1 [90014a1f] PDMats v0.11.41 ⌅ [69de0a69] Parsers v2.8.7 [aea7be01] PrecompileTools v1.3.4 [21216c6a] Preferences v1.5.2 [92933f4c] ProgressMeter v1.11.0 [43287f4e] PtrArrays v1.4.0 [1fd47b50] QuadGK v2.11.3 [189a3867] Reexport v1.2.2 [79098fc4] Rmath v0.9.0 [f2b01f46] Roots v3.0.7 [fdea26ae] SIMD v3.7.2 [0e966ebe] SearchModels v0.5.1 [053f045d] SimilaritySearch v1.2.0 [a2af1166] SortingAlgorithms v1.2.3 [276daf66] SpecialFunctions v2.9.0 [10745b16] Statistics v1.11.1 [82ae8749] StatsAPI v1.8.0 [2913bbd2] StatsBase v0.34.13 [4c63d2b9] StatsFuns v2.2.1 [856f2bd8] StructTypes v1.11.0 [ec057cc2] StructUtils v2.8.5 [7f6f6c8a] TextSearch v1.1.1 [3bb67fe8] TranscodingStreams v0.11.3 [49080126] ZipArchives v2.6.0 ⌅ [68821587] Arpack_jll v3.5.2+0 [f5851436] FFTW_jll v3.3.12+0 [1d5cc7b8] IntelOpenMP_jll v2025.2.0+0 [856f044c] MKL_jll v2025.2.0+0 [efe28fd5] OpenSpecFun_jll v0.5.6+0 [f50d1b31] Rmath_jll v0.5.2+0 [1317d2d5] oneTBB_jll v2022.3.0+0 [0dad84c5] ArgTools v1.2.0 [56f22d72] Artifacts v1.11.0 [2a0f44e3] Base64 v1.11.0 [ade2ca70] Dates v1.11.0 [8ba89e20] Distributed v1.11.0 [f43a241f] Downloads v1.7.0 [7b1f6079] FileWatching v1.11.0 [b77e0a4c] InteractiveUtils v1.11.0 [ac6e5ff7] JuliaSyntaxHighlighting v1.13.0 [4af54fe1] LazyArtifacts v1.11.0 [b27032c2] LibCURL v1.0.0 [76f85450] LibGit2 v1.11.0 [8f399da3] Libdl v1.11.0 [37e2e46d] LinearAlgebra v1.14.0 [56ddb016] Logging v1.11.0 [d6f4376e] Markdown v1.11.0 [a63ad114] Mmap v1.11.0 [ca575930] NetworkOptions v1.3.0 [44cfe95a] Pkg v1.14.0 [de0858da] Printf v1.11.0 [9a3f8284] Random v1.11.0 [ea8e919c] SHA v1.13.0 [9e88b42a] Serialization v1.11.0 [6462fe0b] Sockets v1.11.0 [2f01184e] SparseArrays v1.13.0 [f489334b] StyledStrings v1.13.0 [4607b0f0] SuiteSparse [fa267f1f] TOML v1.0.3 [a4e569a6] Tar v1.10.0 [8dfed614] Test v1.11.0 [cf7118a7] UUIDs v1.11.0 [4ec0a83e] Unicode v1.11.0 [e66e0078] CompilerSupportLibraries_jll v1.5.7+0 [deac9b47] LibCURL_jll v8.21.0+0 [e37daf67] LibGit2_jll v1.9.6+0 [29816b5a] LibSSH2_jll v1.11.103+0 [14a3606d] MozillaCACerts_jll v2026.7.16 [4536629a] OpenBLAS_jll v0.3.34+0 [05823500] OpenLibm_jll v0.8.7+0 [458c3c95] OpenSSL_jll v3.5.7+0 [efcefdf7] PCRE2_jll v10.47.0+0 [bea87d4a] SuiteSparse_jll v7.10.1+0 [83775a58] Zlib_jll v1.3.2+0 [3161d3a3] Zstd_jll v1.5.7+1 [8e850b90] libblastrampoline_jll v5.15.0+0 [8e850ede] nghttp2_jll v1.69.0+0 [3f19e933] p7zip_jll v17.8.0+0 Info Packages marked with ⌅ have new versions available but compatibility constraints restrict them from upgrading. Testing Running tests... Test Summary: | Pass Total Time Unbound type parameters | 1 1 0.1s Test Summary: | Pass Total Time Undefined exports | 1 1 0.0s Test Summary: | Pass Total Time Compare Project.toml and test/Project.toml | 1 1 0.0s Test Summary: | Pass Total Time Compat bounds | 4 4 0.7s Test Summary: | Pass Total Time Persistent tasks | 1 1 1m18.2s Test Summary: | Pass Total Time Search | 15000 15000 0.8s Test Summary: | Pass Total Time SVS | 601 601 1.2s Test Summary: | Pass Total Time BK | 905 905 1.6s Test Summary: | Pass Total Time Merge/Union | 602 602 0.5s Test Summary: | Pass Total Time threshold | 16 16 0.7s Test Summary: | Pass Total Time individual tokenizers | 4 4 5.0s Test Summary: | Pass Total Time message vectors | 1 1 0.5s gettoken(A) = ["hello", ";)", "#jello", "world", "."] gettoken(B) = ["hello", ";)", "#jello", "world", "."] [ Info: (2, 1) (gettrainsize(A), gettrainsize(B), gettrainsize(C)) = (2, 1, 3) Test Summary: | Pass Total Time vocabulary of different kinds of docs | 7 7 15.4s Test Summary: | Pass Total Time Normalize and tokenize | 1 1 0.2s Test Summary: | Pass Total Time Normalize and tokenize bigrams and trigrams | 1 1 0.0s Test Summary: | Pass Total Time Normalize and tokenize | 1 1 0.0s Test Summary: | Pass Total Time TokenPipeline runs its stages in the order that matters | 5 5 0.3s Test Summary: | Pass Total Time custom AbstractTokenGenerator extends tokenize_ without editing TextConfig | 2 2 0.7s Test Summary: | Pass Total Time overridable regex/emoji set | 2 2 0.1s Test Summary: | Pass Total Time nested partial config updates | 2 2 0.1s Test Summary: | Pass Total Time paragraph and sentence tokenizers | 15 15 0.6s Test Summary: | Pass Total Time vocabulary | 4 4 5.2s [ Info: =================== Test Summary: | Pass Total Time Vocabulary and BOW | 1 1 1.4s Test Summary: | Pass Total Time numtokens counts occurrences, not distinct tokens | 5 5 0.4s Test Summary: | Pass Total Time stopword_candidates | 15 15 4.3s Test Summary: | Pass Total Time derive_variants | 55 55 4.0s text1 => x = "hello world!! @user;) #jello.world :)" => sparsevec(Int32[1, 2, 3, 4, 5, 7, 8, 9], Float32[0.30151135, 0.6030227, 0.30151135, 0.30151135, 0.30151135, 0.30151135, 0.30151135, 0.30151135], 9) corpus = ["hello world :)", "@user;) excellent!!", "#jello world."] text1 = "hello world!! @user;) #jello.world :)" text2 = "a b c d e f g h i j k l m n o p q" text2 => v = "a b c d e f g h i j k l m n o p q" => sparsevec(Int32[], Float32[], 9) [ Info: sparsevec(Int32[], Float32[], 9) Test Summary: | Pass Total Time Tokenizer, Dict-based vectors, and vectorize | 1 1 3.8s Test Summary: | Pass Total Time tokenize list of strings as a single message | 1 1 0.5s (length(corpus), length(corpus_bows)) = (3, 5) Test Summary: | Pass Total Time Tokenizer, Dict-based vectors, and vectorize | 2 2 3.0s (gw, lw, dot_, dot(x, y), x, y) = (BinaryGlobalWeighting(), FreqWeighting(), 0.3162, 0.31622776f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.70710677, 0.70710677], 8)) (gw, lw, dot_, dot(x, y), x, y) = (BinaryGlobalWeighting(), TfWeighting(), 0.3162, 0.31622776f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.70710677, 0.70710677], 8)) (gw, lw, dot_, dot(x, y), x, y) = (BinaryGlobalWeighting(), TpWeighting(), 0.3162, 0.31622776f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.70710677, 0.70710677], 8)) (gw, lw, dot_, dot(x, y), x, y) = (IdfWeighting(), BinaryLocalWeighting(), 0.3668, 0.3668393f0, sparsevec(Int32[4, 5], Float32[0.85490215, 0.5187891], 8), sparsevec(Int32[5, 6], Float32[0.70710677, 0.70710677], 8)) (gw, lw, dot_, dot(x, y), x, y) = (IdfWeighting(), TfWeighting(), 0.2053, 0.2053078f0, sparsevec(Int32[4, 5], Float32[0.9569208, 0.29034907], 8), sparsevec(Int32[5, 6], Float32[0.70710677, 0.70710677], 8)) (gw, lw, dot_, dot(x, y), x, y) = (EntropyWeighting(), FreqWeighting(), 0.44456, 0.44455594f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.9940573, 0.10885762], 8)) (gw, lw, dot_, dot(x, y), x, y) = (EntropyWeighting(), TfWeighting(), 0.44456, 0.44455594f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.9940573, 0.10885762], 8)) (gw, lw, dot_, dot(x, y), x, y) = (EntropyWeighting(), TpWeighting(), 0.44456, 0.44455594f0, sparsevec(Int32[4, 5], Float32[0.8944272, 0.4472136], 8), sparsevec(Int32[5, 6], Float32[0.9940573, 0.10885762], 8)) (gw, lw, dot_, dot(x, y), x, y) = (EntropyWeighting(), BinaryLocalWeighting(), 0.7029, 0.70290464f0, sparsevec(Int32[4, 5], Float32[0.70710677, 0.70710677], 8), sparsevec(Int32[5, 6], Float32[0.9940573, 0.10885762], 8)) [ Info: ====== weight: [ Info: Float32[1.0, 1.0, 1.0, 1.0, 1.0, 0.109508395, 1.0, 1.0] [ Info: Float32[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0] [ Info: ("====== token:", ["me", "gusta", "encanta", "lo", "odio", "lol", "!"]) [ Info: ("lo lo odio", "odio esto") ("=========", x, y, norm(x), norm(y)) = ("=========", sparsevec(Int32[4, 5], Float32[0.70710677, 0.70710677], 7), sparsevec(Int32[5], Float32[1.0], 7), 0.99999994f0, 1.0f0) (gw, lw, dot(x, y), dot_, x, y) = (EntropyWeighting(), BinaryLocalWeighting(), 0.70710677f0, 0.7071067690849304, sparsevec(Int32[4, 5], Float32[0.70710677, 0.70710677], 7), sparsevec(Int32[5], Float32[1.0], 7)) [ Info: ====== weight: [ Info: Float32[0.6520767, 1.8744692, 1.1375035, 1.8744692, 1.1375035, 1.1375035, 1.8744692, 1.8744692] [ Info: Float32[1.8744692, 1.8744692, 1.8744692, 1.8744692] [ Info: ("====== token:", ["gusta", "lo", "lol", "!"]) [ Info: ("lo lo odio", "odio esto") ("=========", x, y, norm(x), norm(y)) = ("=========", sparsevec(Int32[2], Float32[1.0], 4), sparsevec(Int32[], Float32[], 4), 1.0f0, 0.0f0) (gw, lw, dot(x, y), dot_, x, y) = (IdfWeighting(), TfWeighting(), 0.0f0, 0.0, sparsevec(Int32[2], Float32[1.0], 4), sparsevec(Int32[], Float32[], 4)) Test Summary: | Pass Total Time Weighting schemes | 15 15 8.1s Test Summary: | Pass Total Time SparseVector-based vectors | 99 99 5.7s Test Summary: | Pass Total Time SparseVector adaptive dot: small/large/ratio thresholds agree with native dot | 19 19 0.5s Test Summary: | Pass Total Time SparseVector centroid/sum matches Dict-based centroid/sum | 211 211 0.9s LOG add! sp=1 ep=7 n=7 mem=164MB max-rss=876MB 2026-08-28T00:17:17.738 Test Summary: | Pass Total Time invindex | 1 1 11.0s Test Summary: | Pass Total Time centroid computing | 1 1 5.1s [ Info: 1 => "la casa roja" [ Info: 2 => "la casa verde" [ Info: 3 => "la casa azul" [ Info: 4 => "la manzana roja" [ Info: 5 => "la pera verde esta rica" [ Info: 6 => "la manzana verde esta rica" [ Info: 7 => "la hoja verde" LOG add! sp=1 ep=7 n=7 mem=170MB max-rss=876MB 2026-08-28T00:17:32.154 invfile.voc = Vocabulary: vocsize: 6 trainsize: 7 numtokens: 25 avgdoclen: 3.5714285714285716 TextConfig: NormalizationConfig: del_diac: true del_dup: false del_punc: false group_num: true group_url: true group_usr: false group_emo: false lc: true re_user: r"@[^;:,.@#&\\\-\"'/:\*\(\)\[\]\¿\?\¡\!\{\}~\<\>\|\s]+" re_url: r"(http|ftp|https)://\S+" re_num: r"[-+]?(\d+\.?\d*)|(\.\d+)" emojis: 3923 emoji chars TokenizationConfig: nlist: Int8[1] mark_token_type: true generators: AbstractTokenGenerator[] pipeline: TokenPipeline(lemmas=nothing, stopwords=nothing) language: unknown invfile.bm25 = BM25Scorer: k1: 1.2 b: 0.75 avgdoclen: 3.571429 δ: 1.0 trainsize: 7 Test Summary: | Pass Total Time bm25 invindex | 3 3 11.3s LOG add! sp=1 ep=7 n=7 mem=252MB max-rss=876MB 2026-08-28T00:17:37.637 Test Summary: | Pass Total Time BM25InvertedFile query-time expansion | 4 4 0.3s Test Summary: | Pass Total Time BM25InvertedFile LOG events: exactly-once :add! | 3 3 6.4s LOG add! sp=1 ep=7 n=7 mem=143MB max-rss=876MB 2026-08-28T00:17:44.531 LOG add! sp=8 ep=8 n=8 mem=172MB max-rss=876MB 2026-08-28T00:17:45.748 Test Summary: | Pass Total Time BM25InvertedFile decoupled index! | 11 11 1.3s LOG add! sp=1 ep=7 n=7 mem=188MB max-rss=876MB 2026-08-28T00:17:46.717 Test Summary: | Pass Total Time bm25score matches BM25InvertedFile search, both for SparseVecView and SparseVector | 15 15 2.5s Test Summary: | Pass Total Time expand_query! | 20 20 4.0s LOG add! sp=1 ep=61 n=61 mem=352MB max-rss=967MB 2026-08-28T00:19:16.780 Test Summary: | Pass Total Time query pipeline | 48 48 1m24.9s Test Summary: | Pass Total Time TextProfile: policy / artifact split | 30 30 3.5s Test Summary: | Pass Total Time variants are not an artifact: derived from the vocabulary, never stored | 7 7 31.2s Test Summary: | Pass Total Time a pre-trained profile drives an index | 16 16 2.7s ┌ Warning: `trainsize(voc::Vocabulary)` is deprecated, use `gettrainsize(voc)` instead. │ caller = _mapreduce(f::typeof(trainsize), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:448 └ @ Base reduce.jl:448 ┌ Warning: `trainsize(voc::Vocabulary)` is deprecated, use `gettrainsize(voc)` instead. │ caller = _mapreduce(f::typeof(trainsize), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:448 └ @ Base reduce.jl:448 ┌ Warning: `numtokens(voc::Vocabulary)` is deprecated, use `getnumtokens(voc)` instead. │ caller = _mapreduce(f::typeof(numtokens), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:448 └ @ Base reduce.jl:448 ┌ Warning: `numtokens(voc::Vocabulary)` is deprecated, use `getnumtokens(voc)` instead. │ caller = _mapreduce(f::typeof(numtokens), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:448 └ @ Base reduce.jl:448 Test Summary: | Pass Total Time fit_profile | 56 56 36.9s ┌ Warning: `token(voc::Vocabulary, tokenID::Integer)` is deprecated, use `gettoken(voc, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `token(voc::Vocabulary)` is deprecated, use `gettoken(voc)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `occs(voc::Vocabulary, tokenID::Integer)` is deprecated, use `getoccs(voc, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `occs(voc::Vocabulary)` is deprecated, use `getoccs(voc)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `ndocs(voc::Vocabulary, tokenID::Integer)` is deprecated, use `getndocs(voc, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `ndocs(voc::Vocabulary)` is deprecated, use `getndocs(voc)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `trainsize(voc::Vocabulary)` is deprecated, use `gettrainsize(voc)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `numtokens(voc::Vocabulary)` is deprecated, use `getnumtokens(voc)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `token(model::VectorModel, tokenID::Integer)` is deprecated, use `gettoken(model, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `occs(model::VectorModel, tokenID::Integer)` is deprecated, use `getoccs(model, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `ndocs(model::VectorModel, tokenID::Integer)` is deprecated, use `getndocs(model, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `trainsize(model::VectorModel)` is deprecated, use `gettrainsize(model)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `weight(model::VectorModel, tokenID::Integer)` is deprecated, use `getweight(model, tokenID)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 ┌ Warning: `weight(model::VectorModel)` is deprecated, use `getweight(model)` instead. │ caller = macro expansion at Test.jl:781 [inlined] └ @ Core /opt/julia/share/julia/stdlib/v1.14/Test/src/Test.jl:781 Test Summary: | Pass Total Time deprecated accessor aliases | 22 22 0.6s Test Summary: | Pass Total Time save_profile / load_profile / zip_profile | 74 74 13.9s ┌ Warning: `trainsize(voc::Vocabulary)` is deprecated, use `gettrainsize(voc)` instead. │ caller = _mapreduce(f::typeof(trainsize), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:451 └ @ Base reduce.jl:451 ┌ Warning: `numtokens(voc::Vocabulary)` is deprecated, use `getnumtokens(voc)` instead. │ caller = _mapreduce(f::typeof(numtokens), op::typeof(Base.add_sum), ::IndexLinear, A::Vector{Vocabulary}) at reduce.jl:451 └ @ Base reduce.jl:451 Test Summary: | Pass Total Time merge_profiles | 127 127 5.9s Test Summary: | Pass Total Time refit_profile | 89 89 15.6s LOG add! sp=1 ep=4 n=4 mem=244MB max-rss=989MB 2026-08-28T00:21:21.295 doc 1 (cosine distance 0.4764): Machine learning and natural language processing in Julia doc 3 (cosine distance 0.6959): Natural language text retrieval with BM25 and inverted files doc 4 (cosine distance 0.8773): Julia programming language for scientific computing and machine learning LOG add! sp=1 ep=4 n=4 mem=245MB max-rss=989MB 2026-08-28T00:21:21.422 Doc 1: the quick brown fox jumps over the lazy dog (distance: -3.3781) Doc 2: brown fox jumps high over the lazy fence (distance: -2.086) one call: vocsize=28 base=true profile: vocsize=28 query_expansion=28 lemmas=1 base=true refitted: vocsize=13 tuned=true lineage: fit(outdim=4, trainsize=6) -> refit(kappa=3.0, lemmas_applied=false, sample_trainsize=3, trainsize=6) bridged: ["clasica" => ["clásica"], "manana" => ["mañana"], "musica" => ["música"]] corrected: ["música"] musica is in 1 document against música's 3, so it reads as a misspelling; searched as música (variant) instead untouched: ["sol"] literal: ["musica"] Test Summary: | Pass Total Time README examples run | 10 10 6.1s LOG add! sp=1 ep=7 n=7 mem=324MB max-rss=989MB 2026-08-28T00:21:27.176 LOG add! sp=1 ep=7 n=7 mem=370MB max-rss=989MB 2026-08-28T00:21:31.976 LOG add! sp=1 ep=7 n=7 mem=149MB max-rss=989MB 2026-08-28T00:21:35.856 LOG add! sp=1 ep=7 n=7 mem=154MB max-rss=989MB 2026-08-28T00:21:36.747 Test Summary: | Pass Total Time FullText and TextInvertedFile | 12 12 9.6s LOG add! sp=1 ep=4 n=4 mem=160MB max-rss=989MB 2026-08-28T00:21:37.578 Test Summary: | Pass Total Time avgdoclen agrees with the lengths BM25 actually measures | 6 6 0.3s LOG add! sp=1 ep=1 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=155MB max-rss=989MB 2026-08-28T00:21:51.066 LOG add! sp=2 ep=2 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=173MB max-rss=989MB 2026-08-28T00:21:54.942 LOG add! sp=1 ep=1 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=146MB max-rss=989MB 2026-08-28T00:22:07.350 LOG add! sp=1 ep=1 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=284MB max-rss=989MB 2026-08-28T00:22:19.173 Test Summary: | Pass Total Time LatentSemanticIndexing (LSI) | 558 558 42.1s Test Summary: | Pass Total Time query expansion is filtered by purpose, not by quality | 13 13 0.8s Test Summary: | Pass Total Time lemma_clusters | 31 31 9.8s LOG add! sp=1 ep=1 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=178MB max-rss=989MB 2026-08-28T00:22:38.244 LOG add! sp=1 ep=1 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=165MB max-rss=995MB 2026-08-28T00:22:58.256 LOG add! sp=2 ep=2 BeamSearch(bsize=4, Δ=1.0, maxvisits=1000000) mem=176MB max-rss=995MB 2026-08-28T00:23:01.799 Test Summary: | Pass Total Time RandomIndexing (RI) | 50 50 36.7s [ Info: FINISH Testing TextSearch tests passed Testing completed after 561.55s PkgEval succeeded after 659.74s