We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model's tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated registry of 90 models across 15 providers and six answer-selection methods, and needs only three core dependencies. For chunking, ChunkRank avoids context-window overflow automatically from the model name, whereas character-based splitters overflow or waste the budget, and a fidelity study across 11 languages shows why token-exact budgets matter beyond English. For answer selection we report a negative result: on NaturalQuestions, TriviaQA and HotpotQA, with extractive and generative readers, no content-based ranker reliably beats taking the first non-empty answer. The reason is reader abstention on chunks that lack the answer, not answer position. A long-context baseline shows that chunking matches single-call reading on single-hop questions, so ChunkRank targets small-window and beyond-window settings. Code, registry and evaluation harness are released.
Figures & tables
Tool
Model-
Auto
Post-chunk
Light-
Async
aware
registry
ranking
weight
stream
LangChain splitter
No
No
No
No
No
LlamaIndex splitter
No
No
No
No
No
Chonkie
No
No
No
Yes
No
ChunkRank
Yes
Yes
Yes
Yes
Yes
Table 1: Feature comparison with the widely-used chunking utilities. “Model-aware” means chunk size is derived from the target model’s tokenizer and context window without manual configuration (LangChain and Chonkie can be pointed at a tokenizer, but the user must supply it; neither resolves it automatically from a model name); “Auto registry” means the library ships pre-configured parameters for named models. Quantitative overflow/utilisation results are in Table 3 .
Reader / dataset
Empty chunks
No-answer ex.
NQ, extractive
52%
22%
TriviaQA, extractive
57%
16%
HotpotQA, extractive
85%
38%
NQ, generative
36%
2%
TriviaQA, generative
51%
7%
HotpotQA, generative
79%
18%
Table 2: Abstention rates. “Empty chunks” is the share of chunks for which the reader returned no answer; “No-answer ex.” is the share of examples with no non-empty candidate at all (every method scores 0). The extractive reader reads full chunks. These rates drive the results: first-non-empty exploits exactly this abstention.
Splitter
Overflow
Mean util.
Max tok.
LangChain char (naive 4:1)
0.64%
75.0%
7,375
Chonkie char (default)
0.60%
76.0%
7,397
LangChain char (aggressive)
0.00%
39.3%
4,559
LangChain tiktoken (manual)
0.00%
84.5%
5,979
Chonkie tiktoken (manual)
0.00%
86.2%
6,000
ChunkRank (auto)
0.00%
75.9%
6,000
Table 3: Chunk overflow rate, mean budget utilisation, and largest chunk across 500 TriviaQA documents ( gpt-4o-mini , 6,000-token budget). Character-based splitters (top two) overflow the budget; token-aware ones (middle) do not, but require manual per-model tokenizer configuration. ChunkRank guarantees zero overflow automatically from the model name, at lower utilisation ( 75.9% vs. 84 – 86% ), a deliberate safety margin.
Method
NaturalQuestions
TriviaQA
EM
F1
EM
F1
First (non-empty)
22.4
29.1
45.4
51.0
Last
14.0
20.7
40.8
47.1
Random
18.8
24.8
43.8
49.4
BM25
21.8
28.6
45.2
50.5
TF-IDF
20.4
27.4
44.8
51.2
Table 4: Exact Match (EM) and token-level F1 for answer selection over the non-empty candidate set (chunk budget = 6,000 tokens, 500 examples per dataset; extractive reader reading full chunks). Bold marks the best non-oracle result per column. No content-based method (BM25, TF-IDF, embedding, cross-encoder) significantly beats first-non-empty: on NQ, embedding and cross-encoder are significantly below it ( p<0.02 ), and on TriviaQA all differences from first-non-empty are non-significant (reader-confidence is numerically highest, p=0.18 ). The reader-based selectors (first-non-empty, reader-confidence) match or beat the content-based ones throughout. Oracle headroom remains but is captured by no method.
Method
NQ
TriviaQA
HotpotQA
Chunked (first-non-empty)
36.0
59.0
34.6
Long-context (single call)
35.0
57.0
55.1
p vs. chunked
0.87
0.70
< 0.001
Table 5: Whole-document single-call (Claude Haiku 4.5) vs. chunked first-non-empty, EM. Single-hop NQ/TriviaQA ( n=100 ): a tie, at essentially equal input-token cost. Multi-hop HotpotQA ( n=136 subset; chunked matches its full-set 33.5 ): long-context wins, as expected when chunking splits multi-hop evidence (Limitation 1). All documents fit the 200k-token window.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Language
Script
ch/tok
len/4 err.
English
Latin
5.15
+27%
Spanish
Latin
4.50
+11%
French
Latin
4.54
+12%
German
Latin
4.32
+8%
Russian
Cyrillic
3.81
−5%
Hindi
Devanagari
3.46
−15%
Appendix
Table 6: Characters per o200k_base token and the error of the len/4 character-ratio estimate on UDHR Article 1 (negative error = under-count = overflow risk). Latin scripts are safe; logographic and abugida scripts under-count by up to 71% , so a character budget calibrated for English overflows the token budget.
Registry model
True tokenizer
True tok.
Proxy err.
qwen2.5-*
Qwen2.5-7B
1,549
−2.3%
deepseek-v3
deepseek-llm-7b
1,561
−3.1%
mistral-*
Mistral-7B-v0.2
1,705
−11.3%
(older BPE)
gpt2
1,517
−0.3%
Appendix
Table 7: Token count of the o200k_base proxy (1,513) vs. the true tokenizer on an 8,000-character English document. Negative error = proxy under-count = the true model emits more tokens than budgeted (overflow risk, absorbed by the reserve).
Method
Deterministic
Dependencies
Semantic
Relative speed
BM25
✓
rank-bm25
No
Fastest
TF-IDF
✓
scikit-learn
No
Fast
Embedding
∼
sentence-transformers
Yes
Moderate
Cross-Encoder
∼
sentence-transformers
Yes
Slow
Appendix
Table 8: Ranker method comparison. “Deterministic” marks methods whose output is identical across runs given the same input. Embedding and cross-encoder results may vary across model versions.
Chunking (6k budget)
1 chunk
6 chunks
11 chunks
gpt-4o-mini
0.7 ms
6.6 ms
14.2 ms
llama-4-scout
0.7 ms
7.6 ms
13.9 ms
Selection
5 cand.
10 cand.
20 cand.
First / Reader-conf.
< 0.01 ms
< 0.01 ms
< 0.01 ms
BM25
0.06 ms
0.08 ms
0.12 ms
TF-IDF
1.8 ms
1.8 ms
2.2 ms
Appendix
Table 9: Wall-clock latency (CPU only, median of 100 runs, chunkrank 2.0.0). Chunking scales with the number of chunks and is essentially model-independent (tiktoken vs. Hugging Face tokenizer overhead is small); selection scales with candidate count. The default first and confidence selectors are effectively free; only the neural methods add tens of milliseconds. Reader/LLM inference itself (not shown) dominates total pipeline latency by orders of magnitude.
Method
EM
F1
First (non-empty)
21.8
27.9
Last
20.0
26.6
Random
20.5
26.9
BM25
21.8
27.8
TF-IDF
21.8
28.0
Embedding
22.2
29.0
Appendix
Table 10: HotpotQA (distractor) EM/F1, 400 examples, 256-token budget, over the non-empty candidate set. Bold marks the best non-oracle result per column, but no method significantly beats first-non-empty (all paired-bootstrap p>0.4 ). Even with answers that are not front-loaded, first-non-empty is not beaten, so the baseline’s strength is due to reader abstention, not answer position.