When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations
Organizations: Independent Researcher
Abstract
Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned units: four embedding pipelines, a direct LLM call, an autonomous agent, and the agent revised by an independent auditor. Generative workflows recover 93-94% of reference correspondences, against at most 77% for embeddings. A ceiling analysis shows that sentence boundaries make some references unrepresentable by the embedding pipelines. Reference recovery is similar across generative workflows: the agent's advantage is 0.5 percentage points (95% CI -0.02 to 1.17), and auditing adds no established benefit. Agents nevertheless produce structurally valid output for all 452 texts, against 437 for direct calls. A blinded three-LLM panel assesses every generative mismatch against the source and human reference. Most mismatches are labelled defensible editorial variation; consensus major-error labels cover only 0.06-0.14% of units. The panel labels significantly fewer residual defects for agents than direct calls (0.7% versus 1.4%), suggesting that reference recovery alone understates alignment quality. On ten long Pali discourses taken as published online, agents and audited agents raise recovery from the direct call's 71% to 84% and 92%. Identical reference-located chunks bring all three to 93%. Agents thus improve structural reliability and reduce judged defects on short passages, while their large recovery advantage on long documents disappears after chunking. In this setting, independent auditing offers little measurable additional benefit on prepared passages.
Figures & tables
| Source segment | Reference English | Direct LLM | MITRA-E + Vecalign | |
|---|---|---|---|---|
| 1 | Aṅguttara Nikāya 4.5 | Numbered Discourses 4.5 | Numbered Discourses 4.5 ✓ | Numbered Discourses 4.5 1. ✗ |
| 2 | 1. Bhaṇḍagāmavagga | 1. At Wares Village | 1. At Wares Village ✓ | At Wares Village With the Stream "These four individuals are found in the world. (segments 2–4 grouped) ✗ |
| 3 | Anusotasutta | With the Stream | With the Stream ✓ | (in group above) ✗ |
| 4 | "Cattārome, bhikkhave, puggalā santo saṁvijjamānā lokasmiṁ. | "These four individuals are found in the world. | "These four individuals are found in the world. ✓ | (in group above) ✗ |
| 5 | Katame cattāro? | What four? | What four? ✓ | What four? ✓ |
| Corpus | Language | Texts | Parent works | Units | Units per text, median (range) | English words | Licence (source / English) |
|---|---|---|---|---|---|---|---|
| Bilara | Pāli | 50 | 50 | 2,350 | 45 (19–79) | 28,860 | public domain / CC0 |
| Itihāsa | Sanskrit | 74 | 2 | 2,499 | 30 (9–80) | 75,472 | Apache-2.0 |
| Sefaria | Hebrew | 279 | 63 | 2,502 | 9 (5–23) | 239,709 | public domain / CC BY 3.0 |
| 84000 | Tibetan | 49 | 25 | 2,482 | 47 (14–111) | 59,157 | CC BY 4.0 |
| Total | 452 | 140 | 9,833 | 403,198 |
| Model | Base model | Domain | Weights | Instruction added to source segments |
|---|---|---|---|---|
| LaBSE [ 16 ] | BERT encoder | general | fp32, local | none |
| F2LLM-v2-1.7B [ 63 ] | Qwen3-1.7B LLM | general | bf16, local | retrieve the English translation |
| Qwen3-Embedding-8B [ 62 ] | Qwen3-8B LLM | general | hosted API | retrieve the English translation |
| MITRA-E [ 43 ] | Gemma 2 9B LLM | Buddhist languages | Q8_0, cloud GPU | find the most similar English text |
| Direct | Agent | Agent + auditor | |
|---|---|---|---|
| Model calls | one | chosen by the agent | agent's calls, then review rounds |
| Tools | none | shell, code, validator, monitored web | as agent, for both aligner and auditor |
| Feedback | none | structural errors, if any | structural errors; up to 5 reviews and 4 revisions |
| Access | OpenRouter API | OpenCode harness | OpenCode harness |
| System | Pāli | Sanskrit | Hebrew | Tibetan | Equal-corpus mean | Valid outputs |
|---|---|---|---|---|---|---|
| LaBSE | 32.4 | 53.8 | 72.7 | 82.3 | 60.3 | 452/452 |
| F2LLM-v2-1.7B | 46.1 | 73.8 | 54.7 | 57.9 | 58.1 | 452/452 |
| Qwen3-Embedding-8B | 46.5 | 76.0 | 58.3 | 74.6 | 63.8 | 452/452 |
| MITRA-E | 60.3 | 83.1 | 77.5 | 86.4 | 76.8 | 452/452 |
| Embedding ceiling | 70.2 | 98.8 | 98.8 | 93.6 | 90.3 | — |
| Direct LLM | 93.4 | 90.9 | 99.8 | 89.4 | 93.4 | 437/452 |
| Workflow | Primary | Strict exact | Valid-output | Valid outputs |
|---|---|---|---|---|
| Direct LLM | 93.4 | 91.5 | 86.6 | 437/452 |
| Agent | 93.9 | 92.7 | 93.9 | 452/452 |
| Agent + auditor | 93.9 | 92.7 | 93.9 | 452/452 |
| Contrast | Primary | Strict exact | Valid-output |
|---|---|---|---|
| Agent − direct | +0.54 [−0.02, 1.17] | +1.19 [0.25, 2.39] | +7.28 [3.90, 11.10] |
| Agent + auditor − agent | +0.03 [−0.13, 0.24] | +0.01 [−0.14, 0.21] | +0.03 [−0.13, 0.24] |
| Workflow | Judged mismatches | Major | Minor | Defensible | Unresolved | Defect rate | Pāli | Sanskrit | Hebrew | Tibetan |
|---|---|---|---|---|---|---|---|---|---|---|
| Direct LLM | 653 | 14 | 127 | 505 | 7 | 1.44 | 2.04 | 1.68 | 0.04 | 2.01 |
| Agent | 599 | 9 | 59 | 530 | 1 | 0.70 | 1.06 | 1.04 | 0.00 | 0.68 |
| Agent + auditor | 596 | 6 | 53 | 536 | 1 | 0.61 | 1.06 | 0.84 | 0.00 | 0.52 |
| Label and segment | Source | Reference English | Candidate English | Remark |
|---|---|---|---|---|
| Major error: Mahābhārata (vol. II, ch. index 180), s0033 | \devanagari अहं हि दिवि दिव्येन विमानेन चरन् पुरा। अभिमानेन मत्तः सन् कंचिन्नान्यमचिन्तयम्॥ | In days of yore when mounted on a celestials chariot I used to range through the heavens, elated with pride, I thought of nothing else. | O tormentor of foes, O mighty monarch, you have rendered me a good service. (direct) | The candidate is English from the preceding verse; this verse's own English, on the celestial chariot and pride, was left unaligned. |
| Minor error: Bhikkhunī Pācittiya 75, s0026 | Dutiyampi yācitabbā. | And she should ask a second | And she should ask a second and a third time. (direct) | The two segments mean "she should ask a second time" and "a third time". The English sentence divides cleanly between them, but the whole of it was given to both, so each carries its neighbour's clause. No meaning is lost. |
| s0027 | Tatiyampi yācitabbā. | and a third time. | And she should ask a second and a third time. (direct) | |
| Defensible variation: AN 4.5, s0034 | Sa ve muni vusitabrahmacariyo, | they've completed the spiritual journey and gone to the end of the world, | they've completed the spiritual journey (agent; agent + auditor) | "Gone to the end of the world" renders lokantagū , which stands in s0035, so the candidate follows the Pāli more closely than the reference. |
| s0035 | Lokantagū pāragatoti vuccatī"ti. | they're called 'one who has gone beyond'." | and gone to the end of the world, they're called 'one who has gone beyond'." (agent; agent + auditor) |
| Workflow | Run 0 | Run 1 | Run 2 | Mean | SD (points) | Valid outputs (runs 0 / 1 / 2) |
|---|---|---|---|---|---|---|
| Direct LLM | 92.7 | 93.3 | 93.1 | 93.0 | 0.31 | 35 / 40 / 37 of 40 |
| Agent | 94.5 | 93.8 | 93.7 | 94.0 | 0.44 | 40 / 40 / 40 of 40 |
| Agent + auditor | 94.7 | 93.7 | 93.4 | 93.9 | 0.69 | 40 / 40 / 40 of 40 |
| Workflow | Model calls per text (median) | Prompt tokens (M) | Served from cache | Completion tokens (M) | Median time per text (s) | Cost (USD) |
|---|---|---|---|---|---|---|
| Direct LLM | 1 | 2.1 | 1% | 4.7 | 56 | 1.16 billed |
| Agent | 9 | 75.3 | 85% | 3.9 | 80 | 2.03 equivalent |
| Agent + auditor | 16 | 142.2 | 86% | 8.0 | 135 | 6.02 equivalent |
| Workflow | Whole documents | Valid documents | Identical chunks | Valid chunks |
|---|---|---|---|---|
| Direct LLM | 71.5 | 8/10 | 93.4 | 31/36 |
| Agent | 84.3 | 10/10 | 93.5 | 36/36 |
| Agent + auditor | 92.1 | 10/10 | 93.6 | 36/36 |
| Segment | Pāli | Literal sense | Reference English | All three generative workflows |
|---|---|---|---|---|
| s0005 | Kodhaṁ jahe vippajaheyya mānaṁ, | give up anger, abandon conceit | Give up anger, get rid of conceit, | Give up anger, get rid of conceit, |
| s0006 | Saṁyojanaṁ sabbamatikkameyya; | escape every fetter | and escape every fetter. | and escape every fetter. |
| s0007 | Taṁ nāmarūpasmimasajjamānaṁ, | him, not clinging to name and form | Sufferings don't befall one who has nothing, | not clinging to name and form. |
| s0008 | Akiñcanaṁ nānupatanti dukkhā. | sufferings do not befall one who has nothing | not clinging to name and form. | Sufferings don't befall one who has nothing, |
| Text | Observation |
|---|---|
| Mahābhārata vol. iii, ch. idx 133, ref. 1 | Sanskrit "Duryodhana said", English "Dhritarashtra said" |
| Mahābhārata vol. vi, ch. idx 27, refs 41–43 | Content spills across references; the named actor changes |
| Mishnah Ta'anit 2:1 | Malformed English ("fast days?They"; parenthesis cut off in "(av bet.") |
| SN 16.11:9.5–9.6 | Translation of 9.6 contained in 9.5; 9.6 empty |
| Toh 226, TU-2 | Malformed English title |
| LaBSE | F2LLM-v2-1.7B | Qwen3-Embedding-8B | MITRA-E | |
|---|---|---|---|---|
| Checkpoint | sentence-transformers/LaBSE @836121a | codefuse-ai/F2LLM-v2-1.7B @c5650fe | qwen/qwen3-embedding-8b via OpenRouter (Nebius) | buddhist-nlp/gemma-2-mitra-e , GGUF @952ad99 |
| Precision | fp32 | bf16 | provider-undisclosed | Q8_0 |
| Source prefix | none | Instruct: Retrieve the English translation of the given passage.\nQuery: | same as F2LLM | <instruct>Please find the semantically most similar text in English.\n<query> |
| Pooling | released pipeline, per 254-token window; windows averaged | last token | provider | last token |
| Token limit / longest input | windows cover all tokens | 8,192 / 7,278 | 32,000 / 7,278 | 4,096 / 3,537 |
| Batch | 12 | 8 | 16 (4 concurrent requests) | 8 |
| Model | Input | Cached input | Output |
|---|---|---|---|
| Muse Spark 1.3 Contributor | 0.10 | 0.002 | 0.20 |
| DeepSeek V4.1 Flash | 0.15 | 0.003 | 0.60 |
| Contrast | Equal-corpus | Text-level interval | Pāli | Sanskrit | Hebrew | Tibetan |
|---|---|---|---|---|---|---|
| Agent − Direct LLM | +0.54 [−0.02, +1.17] | [−0.07, +1.23] | +0.00 [−1.22, +1.20] | +1.04 [−0.36, +3.11] | +0.08 [+0.00, +0.21] | +1.05 [+0.05, +2.07] |
| Agent + auditor − Agent | +0.03 [−0.13, +0.24] | [−0.14, +0.23] | +0.26 [+0.00, +0.82] | +0.08 [−0.24, +0.48] | −0.04 [−0.13, +0.00] | −0.16 [−0.53, +0.27] |
| Direct LLM − LaBSE | +33.05 [+30.44, +35.73] | [+30.59, +35.54] | +60.98 [+54.11, +67.60] | +37.05 [+33.35, +40.65] | +27.06 [+23.39, +30.65] | +7.09 [+1.29, +13.75] |
| Direct LLM − F2LLM-v2-1.7B | +35.22 [+31.17, +39.18] | [+31.96, +38.60] | +47.28 [+40.30, +54.30] | +17.05 [+13.88, +20.19] | +45.08 [+41.30, +48.79] | +31.47 [+17.74, +44.91] |
| Direct LLM − Qwen3-Embedding-8B | +29.51 [+26.54, +32.48] | [+26.92, +32.16] | +46.89 [+40.29, +53.31] | +14.89 [+12.00, +17.74] | +41.49 [+37.57, +45.12] | +14.79 [+6.05, +23.74] |
| Direct LLM − MITRA-E | +16.51 [+14.05, +19.09] | [+14.19, +18.97] | +33.06 [+25.82, +40.78] | +7.76 [+5.73, +9.99] | +22.26 [+19.35, +25.31] | +2.94 [−2.17, +8.59] |
| Workflow | Corpus | Primary | Strict exact | Valid-output | Valid outputs | Punctuation-rule units |
|---|---|---|---|---|---|---|
| Direct LLM | Pāli | 93.4 | 93.2 | 88.5 | 48 | 4 |
| Direct LLM | Sanskrit | 90.9 | 90.9 | 82.4 | 68 | 0 |
| Direct LLM | Hebrew | 99.8 | 99.8 | 99.5 | 278 | 0 |
| Direct LLM | Tibetan | 89.4 | 82.1 | 76.1 | 43 | 180 |
| Agent | Pāli | 93.4 | 93.4 | 93.4 | 50 | 0 |
| Agent | Sanskrit | 91.9 | 91.9 | 91.9 | 74 | 0 |
| Measure | Contrast | Equal-corpus | Pāli | Sanskrit | Hebrew | Tibetan |
|---|---|---|---|---|---|---|
| Strict exact | Agent − Direct LLM | +1.19 [+0.25, +2.39] | +0.17 [−1.10, +1.39] | +1.04 [−0.36, +3.11] | +0.08 [+0.00, +0.21] | +3.46 [+0.44, +7.72] |
| Strict exact | Agent + auditor − Agent | +0.01 [−0.14, +0.21] | +0.26 [+0.00, +0.82] | +0.08 [−0.24, +0.48] | −0.04 [−0.13, +0.00] | −0.24 [−0.56, +0.06] |
| Valid-output | Agent − Direct LLM | +7.28 [+3.90, +11.10] | +4.85 [−0.47, +12.89] | +9.56 [+2.83, +17.54] | +0.40 [+0.00, +1.26] | +14.30 [+5.10, +25.43] |
| Valid-output | Agent + auditor − Agent | +0.03 [−0.13, +0.24] | +0.26 [+0.00, +0.82] | +0.08 [−0.24, +0.48] | −0.04 [−0.13, +0.00] | −0.16 [−0.53, +0.27] |
| Workflow | Text | Units |
|---|---|---|
| Direct LLM | 431-toh805-v4 | 43 |
| Direct LLM | 414-toh44-38-v4 | 31 |
| Direct LLM | 439-toh220-v4 | 29 |
| Direct LLM | 423-toh220-v4 | 26 |
| Direct LLM | 419-toh805-v4 | 23 |
| Direct LLM | 428-toh112-v4 | 16 |
| Pipeline | Recovery (%) | Singleton rows (pp) | Grouped rows (pp) | English ≤ 6 (%) |
|---|---|---|---|---|
| LaBSE | 60.3 | 55.6 | 4.7 | 59.0 |
| F2LLM-v2-1.7B | 58.1 | 53.9 | 4.2 | 56.9 |
| Qwen3-Embedding-8B | 63.8 | 59.1 | 4.7 | 62.0 |
| MITRA-E | 76.8 | 68.5 | 8.3 | 74.2 |
| Corpus | Ceiling (≤3 / ≤12) | Ceiling (≤3 / ≤6) | Single-segment ceiling | LaBSE | F2LLM | Qwen3 | MITRA-E |
|---|---|---|---|---|---|---|---|
| Pāli | 70.2 | 70.1 | 42.6 | 46 | 64 | 65 | 83 |
| Sanskrit | 98.8 | 98.3 | 91.8 | 54 | 75 | 77 | 84 |
| Hebrew | 98.8 | 87.2 | 98.6 | 74 | 55 | 59 | 79 |
| Tibetan | 93.6 | 93.2 | 77.0 | 86 | 60 | 78 | 90 |
| Equal-corpus mean | 90.3 | 87.2 | 77.5 | 66 | 64 | 70 | 84 |
| Collection | Texts | Units | Ceiling | LaBSE | F2LLM | Qwen3 | MITRA-E | Direct | Agent | Audited |
|---|---|---|---|---|---|---|---|---|---|---|
| Dīgha Nikāya | 11 | 405 | 95.8 | 51.1 | 66.4 | 70.9 | 80.5 | 94.8 | 96.3 | 96.3 |
| Majjhima Nikāya | 6 | 374 | 80.5 | 41.2 | 54.5 | 44.4 | 66.8 | 96.8 | 96.3 | 96.3 |
| Saṃyutta Nikāya | 10 | 388 | 76.5 | 39.7 | 56.7 | 54.1 | 68.0 | 93.3 | 96.6 | 96.6 |
| Aṅguttara Nikāya | 9 | 372 | 79.6 | 26.9 | 58.9 | 56.7 | 72.3 | 97.3 | 94.6 | 94.6 |
| Udāna | 1 | 24 | 79.2 | 45.8 | 50.0 | 50.0 | 70.8 | 83.3 | 83.3 | 83.3 |
| Itivuttaka | 1 | 32 | 56.2 | 9.4 | 37.5 | 37.5 | 43.8 | 87.5 | 87.5 | 87.5 |
| Corpus | Gained | Lost | Net | Recovered in both | Recovered in neither | Valid → valid |
|---|---|---|---|---|---|---|
| Pāli | 6 | 0 | 6 | 2194 | 150 | 50 |
| Sanskrit | 6 | 4 | 2 | 2293 | 196 | 74 |
| Hebrew | 0 | 1 | -1 | 2498 | 3 | 279 |
| Tibetan | 4 | 8 | -4 | 2236 | 234 | 49 |
| Total | 16 | 13 | 3 | 9221 | 583 | 452 |
| Run | Workflow | Text | Corpus | Departure from supplied English | Omitted words | Fixed by normalization | Recovered |
|---|---|---|---|---|---|---|---|
| 0 | Direct LLM | 434-toh61-v4 | Tibetan | quotation mark added or removed (1) | 0 | no | 78/82 |
| 0 | Direct LLM | 420-toh47-v4 | Tibetan | space between quotation marks removed (1) | 0 | no | 30/35 |
| 0 | Direct LLM | 452-toh1-1-v4 | Tibetan | quotation mark added or removed (8); space between quotation marks removed (1) | 0 | no | 20/44 |
| 0 | Direct LLM | 416-toh75-v4 | Tibetan | space between quotation marks removed (1) | 0 | no | 77/88 |
| 0 | Direct LLM | 438-toh555-v4 | Tibetan | supplied text omitted (1) | 12 | no | 33/48 |
| 0 | Direct LLM | 112-ramayana_vol-i_chapter-index-109 | Sanskrit | diacritic or letter form changed (1) | 0 | no | 23/26 |
| Judge | Workflow | Major | Minor | Defensible | Major rate | Defect rate |
|---|---|---|---|---|---|---|
| MiMo | Direct LLM | 18 | 111 | 524 | 0.18 | 1.32 |
| MiMo | Agent | 6 | 63 | 530 | 0.06 | 0.71 |
| MiMo | Agent + auditor | 2 | 57 | 537 | 0.02 | 0.61 |
| GLM | Direct LLM | 20 | 119 | 514 | 0.20 | 1.42 |
| GLM | Agent | 10 | 65 | 524 | 0.10 | 0.77 |
| GLM | Agent + auditor | 8 | 62 | 526 | 0.08 | 0.72 |
| Judge | Severity | Contrast | Reduction |
|---|---|---|---|
| Consensus | major + minor | Direct LLM − Agent | +0.75 [+0.17, +1.42] |
| Consensus | major + minor | Agent − Agent + auditor | +0.09 [−0.01, +0.22] |
| Consensus | major | Direct LLM − Agent | +0.05 [−0.05, +0.18] |
| Consensus | major | Agent − Agent + auditor | +0.03 [+0.00, +0.08] |
| MiMo | major + minor | Direct LLM − Agent | +0.61 [+0.09, +1.23] |
| MiMo | major + minor | Agent − Agent + auditor | +0.10 [−0.02, +0.26] |
| Pair | Measure | Value |
|---|---|---|
| All three | Fleiss' κ | 0.373 |
| MiMo/GLM | raw agreement | 0.864 |
| MiMo/GLM | Cohen's κ | 0.490 |
| MiMo/GLM | Cohen's κ, linear weights | 0.510 |
| MiMo/Luna | raw agreement | 0.757 |
| MiMo/Luna | Cohen's κ | 0.367 |
| Split | Group | Workflow | Major | Minor | Defensible | Unresolved |
|---|---|---|---|---|---|---|
| English shared | not shared | Agent | 6 | 36 | 468 | 0 |
| English shared | shared | Agent | 3 | 23 | 62 | 1 |
| English shared | not shared | Agent + auditor | 3 | 31 | 480 | 0 |
| English shared | shared | Agent + auditor | 3 | 22 | 56 | 1 |
| English shared | not shared | Direct LLM | 9 | 46 | 475 | 6 |
| English shared | shared | Direct LLM | 5 | 81 | 30 | 1 |
| Run | Workflow | Corpus | Recovery | Valid |
|---|---|---|---|---|
| 0 | Direct LLM | Pāli | 90.2 | 9/10 |
| 0 | Direct LLM | Sanskrit | 87.9 | 8/10 |
| 0 | Direct LLM | Hebrew | 100.0 | 10/10 |
| 0 | Direct LLM | Tibetan | 92.7 | 8/10 |
| 0 | Agent | Pāli | 88.5 | 10/10 |
| 0 | Agent | Sanskrit | 96.7 | 10/10 |
| Pipeline | Hardware | Inputs | Input tokens (M) | Embedding (h) | Vecalign (h) | API (USD) | GPU rental (USD) |
|---|---|---|---|---|---|---|---|
| LaBSE | local NVIDIA GeForce RTX 4080 Laptop GPU (WSL) | 226,705 | 40.2 | 0.37 | 0.005 | 0.00 | — |
| F2LLM-v2-1.7B | local NVIDIA GeForce RTX 4080 Laptop GPU (WSL) | 227,157 | 43.9 | 1.48 | 0.014 | 0.00 | — |
| Qwen3-Embedding-8B | OpenRouter hosted embedding, Nebius route | 227,157 | 43.8 | 3.44 | 0.025 | 0.44 | — |
| MITRA-E | rented Vast.ai NVIDIA RTX 5090 (all layers on GPU) | 227,157 | 39.6 | 1.82 | 0.011 | 0.00 | 1.20 |
| Judge | Model | Contexts | Prompt tokens (M) | Completion tokens (M) | OpenRouter (USD) | Upstream (USD) |
|---|---|---|---|---|---|---|
| MiMo | xiaomi/mimo-v2.6-flash | 232 | 4.45 | 1.94 | 1.05 | — |
| GLM | z-ai/glm-5.3-flash | 232 | 3.29 | 0.27 | 0.63 | — |
| Luna | openai/gpt-5.6-luna | 232 | 3.39 | 0.19 | 0.00 | 1.01 |
| Run | Workflow | Model calls | Prompt tokens (M) | Completion tokens (M) | USD | Median time (s) |
|---|---|---|---|---|---|---|
| 1 | Direct LLM | 40 | 0.2 | 0.57 | 0.14 | 69 |
| 1 | Agent | 473 | 10.6 | 0.55 | 0.26 | 144 |
| 1 | Agent + auditor | 803 | 19.5 | 1.18 | 0.84 | 295 |
| 2 | Direct LLM | 40 | 0.2 | 0.58 | 0.14 | 67 |
| 2 | Agent | 445 | 9.9 | 0.56 | 0.27 | 176 |
| 2 | Agent + auditor | 779 | 19.3 | 1.17 | 0.84 | 279 |
| Document | Units | Whole: direct | Whole: agent | Whole: audited | Chunks: direct | Chunks: agent | Chunks: audited |
|---|---|---|---|---|---|---|---|
| MN40 | 59 | 36 | 43 | 43 | 50 | 47 | 48 |
| MN52 | 50 | 41 | 42 | 42 | 42 | 43 | 42 |
| MN55 | 79 | 77 | 77 | 77 | 77 | 77 | 77 |
| MN71 | 45 | 43 | 41 | 41 | 43 | 43 | 43 |
| MN124 | 66 | 64 | 64 | 64 | 64 | 64 | 64 |
| MN131 | 75 | 72 | 72 | 72 | 72 | 72 | 72 |
| Condition | Workflow | Model calls | Prompt tokens (M) | Completion tokens (M) | Summed minutes | USD |
|---|---|---|---|---|---|---|
| whole documents | Direct LLM | 12 | 0.2 | 0.40 | 44 | 0.10 billed |
| whole documents | Agent | 342 | 25.4 | 0.47 | 96 | 0.35 equivalent |
| whole documents | Agent + auditor | 606 | 54.9 | 1.13 | 177 | 1.16 equivalent |
| identical gold-located chunks | Direct LLM | 36 | 0.2 | 0.67 | 63 | 0.15 billed |
| identical gold-located chunks | Agent | 559 | 13.3 | 0.77 | 124 | 0.33 equivalent |
| identical gold-located chunks | Agent + auditor | 905 | 24.2 | 1.65 | 212 | 1.04 equivalent |