The Last Mile Is the File: OfficeEditBench for Preservation-Aware Office Editing
Organizations: Peking University
Abstract
A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped maintenance of spreadsheets, presentations, and documents. Task contracts specify required updates, protected state, native structures, and applicable interaction requirements. Across 510 archived task-system outcomes from WorkBuddy, Doubao, and Codex, we distinguish file delivery, target completion, and verifier-defined acceptance. Hard package-valid delivery ranges from 92% to 100%, yet no selected output satisfies the complete contract. Case analysis highlights why local correctness is insufficient: an updated value can lose its generating formula, a revised rule can fail to reach related conclusions, and a new deadline can omit a retained prerequisite. These mechanisms connect artifact-level checks to the continued maintainability of Office files. We analyze maintenance failures while distinguishing frozen automatic verdicts from human acceptability. OfficeEditBench provides a testbed for completing required changes while preserving the logic and scope of existing work.
Figures & tables
| Tier | Tasks | Artifact coverage | Contract emphasis |
|---|---|---|---|
| Easy v2 | 50 | 20 XLSX, 20 PPTX, 10 DOCX; one file | Local target/type correctness; non-target content and structure |
| Medium | 70 | 30 spreadsheet-, 20 presentation-, 20 document-primary | Contextual business edits; linked native objects and protected narrative |
| Hard | 50 | XLSX + PPTX + DOCX per task; 2–3 evidence files | Cross-file consequences, approval boundaries, source authority |
| Tier | System | Strict pass | Rate (%) | Mean | [ ] | [ ] |
|---|---|---|---|---|---|---|
| Easy | WorkBuddy † | 43/50 | 86.00 | 91.60 | – | – |
| Easy | Doubao | 29/50 | 58.00 | 63.60 | – | – |
| Easy | Codex | 42/50 | 84.00 | 98.40 | – | – |
| Medium | WorkBuddy † | 37/70 | 52.86 | 82.43 | 85.97 [70] | 71.81 [70] |
| Medium | Doubao | 35/70 | 50.00 | 78.36 | – | – |
| Medium | Codex | 44/70 | 62.86 | 88.96 | 88.08 [70] | 91.63 [70] |
| Task | Reported outcome | Maintenance obligation | |
|---|---|---|---|
| OBM-032 | 100 | Approved total synchronized across workbook, conclusions, native chart, and embedded data | Complete the authorized chain while retaining other states, months, and source records. |
| OBM-034 | 76.92 | Current values correct; operator reports replacing AVERAGE with a constant | Preserve the relationship to source data, not only today’s displayed value. |
| OBM-058 | 62.50 | Rule says delay minutes; a 15-minute flight still labeled on time in summary/timeline | Propagate a changed definition to its dependent statements. |
| OBM-062 | 0 | Deadline changed to 30 days after acceptance; compliant-invoice prerequisite omitted | Update the requested term without deleting a retained condition. |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| System / tier | Inspected evidence | Selection and uncertainty |
|---|---|---|
| WorkBuddy / Easy | Original score/check JSON and generated task report; 50 tasks | Final formal UI file cards; retries included. 50 readable outputs in stored checks. |
| WorkBuddy / Medium | Original score/check JSON with ; 70 tasks | Formal delivery for all 70; DNS retries on 055/056; authorized candidate B on 063. |
| WorkBuddy / Hard | Original score/check JSON; 50 tasks; exploratory designation retained | 46 package-valid triples; four missing-submission zeros. |
| Doubao / Easy | Score/check JSON and progress metadata | Fresh conversation; first delivery selected according to metadata. |
| Doubao / Medium | Two progress files: 20 + 50 task scores | Completed 001–070; component means not available in the imported records. |
| Doubao / Hard | Final score/check JSON and historical summary | Final corrected snapshot after reruns and interaction corrections. |
| Scenario | Task IDs | Spreadsheet rows | Slides | DOCX pages |
|---|---|---|---|---|
| Public-company reporting | 13 | 160,000 | 112 | 86 |
| Retail operations | 13 | 220,000 | 128 | 92 |
| Capital/procurement | 12 | 220,000 | 118 | 90 |
| Environmental compliance | 12 | 220,000 | 142 | 104 |
| Quantity | Recorded value |
|---|---|
| Target-check specifications / task | 9–43 |
| All gold operation specifications / task | 14–43 |
| Explicit preservation specifications / task | 0–8 |
| Declared depth labels (tasks) | 4 (16); 5 (28); 7 (6) |
| Dependency edges / distinct edit sites | Not enumerated |
| Protected-object count | Not enumerated |
| WorkBuddy | Doubao | Codex | ||||
|---|---|---|---|---|---|---|
| Format | Pass | Mean | Pass | Mean | Pass | Mean |
| XLSX | 19/20 | 96.00 | 19/20 | 99.50 | 12/20 | 96.00 |
| PPTX | 19/20 | 96.00 | 1/20 | 14.50 | 20/20 | 100.00 |
| DOCX | 5/10 | 74.00 | 9/10 | 90.00 | 10/10 | 100.00 |
| WorkBuddy | Doubao | Codex | ||||
|---|---|---|---|---|---|---|
| Task IDs | Pass | Mean | Pass | Mean | Pass | Mean |
| 001–010 | 8/10 | 97.00 | 8/10 | 98.48 | 10/10 | 100.00 |
| 011–030 | 11/20 | 80.40 | 11/20 | 82.31 | 13/20 | 88.53 |
| 031–050 | 13/20 | 94.84 | 8/20 | 78.08 | 11/20 | 94.29 |
| 051–070 | 5/20 | 64.75 | 8/20 | 64.63 | 10/20 | 78.56 |
| Failed predicate | WorkBuddy ( ) | Doubao ( ) | Codex ( ) |
|---|---|---|---|
| Non-target document text | 34 | 45 | 49 |
| Non-target spreadsheet cells/formulas | 33 | 44 | 35 |
| Non-target presentation notes | 25 | 23 | 18 |
| Non-target presentation text | 20 | 26 | 30 |
| Source | System | Strict/ | Mean | |||
|---|---|---|---|---|---|---|
| Public-company | WorkBuddy | 13/13 | 0/13 | 52.48 | 44.37 | 83.93 |
| Doubao | 13/13 | 0/13 | 46.70 | 39.28 | 71.61 | |
| Codex | 12/13 | 0/13 | 51.10 | 48.83 | 81.96 | |
| Retail | WorkBuddy | 10/13 | 0/13 | 36.91 | 38.36 | 84.40 |
| Doubao | 13/13 | 0/13 | 41.97 | 32.91 | 73.27 | |
| Codex | 13/13 | 0/13 | 53.44 | 45.42 | 82.21 |
| Omitted source | Retained | WorkBuddy | Doubao | Codex |
|---|---|---|---|---|
| None (full snapshot) | 50 | 45.89 | 43.52 | 51.29 |
| Public-company | 37 | 43.58 | 42.40 | 51.36 |
| Retail | 37 | 49.05 | 44.06 | 50.54 |
| Capital/procurement | 38 | 45.57 | 43.99 | 51.14 |
| Environmental | 38 | 45.40 | 43.61 | 52.12 |
| Check family | Measured property | Not established by that check alone |
|---|---|---|
| Readability/package | Expected file is parseable or ZIP-valid under the checker | Native Office opening, recalculation, visual fidelity, accessibility |
| Target | Specified value/type, text, formula, or native object meets a predicate | Every business claim is true; all equivalent answers are accepted |
| Preservation | Selected non-target state matches the starting artifact | Severity of a mismatch; byte or pixel identity of the full file |
| Structure | Selected sheets, objects, paragraph/table structure or identifiers match | A failed structural signature means the output is unusable |
| Interaction | Required concepts and fixed reply satisfy the transcript check | Full pre-reply behavior, historical isolation, universal first attempt |
| WorkBuddy | Doubao | Codex | WorkBuddy | Doubao | Codex | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ID | Score | Pass | Score | Pass | Score | Pass | ID | Score | Pass | Score | Pass | Score | Pass |
| 001 | 100.00 | Y | 100.00 | Y | 90.00 | N | 026 | 100.00 | Y | 10.00 | N | 100.00 | Y |
| 002 | 100.00 | Y | 100.00 | Y | 90.00 | N | 027 | 100.00 | Y | 100.00 | Y | 100.00 | Y |
| 003 | 100.00 | Y | 100.00 | Y | 90.00 | N | 028 | 100.00 | Y | 10.00 | N | 100.00 | Y |
| 004 | 100.00 | Y | 100.00 | Y | 90.00 | N | 029 | 100.00 | Y | 10.00 | N | 100.00 | Y |
| 005 | 100.00 | Y | 100.00 | Y | 100.00 | Y | 030 | 20.00 | N | 10.00 | N | 100.00 | Y |
| WorkBuddy | Doubao | Codex | WorkBuddy | Doubao | Codex | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ID | Score | Pass | Score | Pass | Score | Pass | ID | Score | Pass | Score | Pass | Score | Pass |
| 001 | 85.00 | N | 100.00 | Y | 100.00 | Y | 036 | 87.50 | N | 100.00 | Y | 100.00 | Y |
| 002 | 85.00 | N | 100.00 | Y | 100.00 | Y | 037 | 100.00 | Y | 100.00 | Y | 83.93 | N |
| 003 | 100.00 | Y | 100.00 | Y | 100.00 | Y | 038 | 100.00 | Y | 100.00 | Y | 100.00 | Y |
| 004 | 100.00 | Y | 100.00 | Y | 100.00 | Y | 039 | 76.56 | N | 35.27 | N | 90.62 | N |
| 005 | 100.00 | Y | 100.00 | Y | 100.00 | Y | 040 | 100.00 | Y | 100.00 | Y | 100.00 | Y |
| WorkBuddy | Doubao | Codex | WorkBuddy | Doubao | Codex | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ID | Score | Pass | Score | Pass | Score | Pass | ID | Score | Pass | Score | Pass | Score | Pass |
| 001 | 72.14 | N | 70.29 | N | 66.71 | N | 026 | 72.95 | N | 43.00 | N | 52.09 | N |
| 002 | 61.82 | N | 46.55 | N | 56.73 | N | 027 | 47.76 | N | 33.85 | N | 33.21 | N |
| 003 | 49.63 | N | 39.25 | N | 48.27 | N | 028 | 28.25 | N | 33.14 | N | 55.11 | N |
| 004 | 49.63 | N | 39.25 | N | 64.27 | N | 029 | 40.60 | N | 35.21 | N | 40.60 | N |
| 005 | 62.88 | N | 56.06 | N | 61.21 | N | 030 | 47.12 | N | 34.91 | N | 47.12 | N |
| Benchmark | Evaluation unit | Relationship to the present study |
|---|---|---|
| SWE-bench (ICLR 2024) | Repository patch evaluated with issue-resolution and regression tests | Prior maintenance principle; our protected observations concern native Office state. |
| WebArena (ICLR 2024) | Web task evaluated by functional outcomes | Persistent-environment evaluation; our deliverables are existing files after scoped edits. |
| -bench (ICLR 2025) | User/tool episode, database end state, repeated-trial reliability | State and interaction constraints; our fixed clarification checks do not measure the same reliability statistic. |
| MLE-bench (ICLR 2025) | ML competition submission evaluated against held-out targets | Submission-based interface and explicit execution reporting; nominal limits do not replace per-attempt records. |
| ScienceAgentBench (ICLR 2025) | Scientific programming task with automatic and human assessment | Validity evidence beyond automatic outcomes; our construction QA is not independent human validation. |
| OfficeEditBench (this work) | Authorized edit to native Office artifacts with protected state | 170-task suite, decomposed predicates, and a descriptive audit of selected archived outputs. |
| Benchmark | Artifact scope and scoring | Preservation and validity evidence |
|---|---|---|
| PPTArena ( Ofengenden et al., 2026 ) | Existing PPTX editing; structure and visual judges | Instruction/quality assessment with human alignment; focused on presentations. |
| PPT-Eval ( Gandhi et al., 2026 ) | PPTX creation/editing; code/model rubrics and partial credit | Extra-change penalties; meta-evaluation on human-constructed completion attempts. |
| DeckEdit-Bench ( Kim et al., 2026 ) | Native PPTX editing; page targeting and object preservation | Preservation of out-of-scope objects, distinct from task-level conjunctive acceptance. |
| OmegaUse-OfficeVal ( Zhou et al., 2026 ) | Multi-format Office/PDF deliverables; usability gate then weighted rubrics | Negative criteria for unintended damage; expert–code calibration and a shared execution scaffold. |
| OfficeEditBench | Native XLSX/PPTX/DOCX; scoped updates and protected-state predicates | Conjunctive acceptance plus component diagnostics; construction QA, but no completed independent agent-output adjudication. |