What Happens During Autonomous Deep Research After the User Steps Away?
Organizations: University of Southern California · University of Michigan, Ann Arbor · Institute of Automation, Chinese Academy of Sciences · Zhejiang University · Beijing University of Posts and Telecommunications · Shanghai University of Finance and Economics
Abstract
In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommendations. We introduce DRaligned, a counterfactual behavioral evaluation framework built on PDR-Bench. By varying one task-relevant user factor while keeping the remaining context fixed, we compare acquisition requests, working drafts, and final reports. Source-grounded extraction, blinded local judgments, and deterministic aggregation yield coarse directional measurements while leaving ambiguous cases unresolved. Our experiments show that strong user-specific delivery can emerge from a largely shared research process: agents investigate similar broad questions but allocate requests differently, and final recommendations distinguish user conditions more clearly than explicit requests do. Reports can also integrate user factors that were not jointly visible during acquisition. In readable draft-to-report comparisons, recommendations often retain their coarse user-specific direction despite substantial rewriting. Final directional differences recur across tested agent models, execution harnesses, and evaluator models, even as execution paths vary. These findings describe how initial user information shapes autonomous research and clarify the relationship between the process an agent follows and the recommendations it delivers.
Figures & tables
| Agent runs | |||||
|---|---|---|---|---|---|
| Study | Tasks | Models / setting | Interface | Pairs | Runs |
| Initial randomized | 10 | DeepSeek / off | Report tool | 60 | 180/180 |
| Exploratory | 8 | DeepSeek / off | 4 setups | 64 | 127/128 |
| Main | 24 | Five families / lowest | Report tool | 110 | 272/272 |
| Reasoning follow-up | 24 | Five families / high | Report tool | 88 | 175/176 |
| Factor follow-up | 12 | Five families / high | Report tool | 128 | 236/256 |
| Quantity | Count | Rate |
|---|---|---|
| Items per round | 312 | 100.0% |
| Instrument accepted | 96 | 30.8% |
| Instrument abstained | 216 | 69.2% |
| Principal-reviewer exact agreement | 192 | 61.5% |
| Principal-reviewer disagreement | 120 | 38.5% |
| Pre-gate agreement among abstentions | 168/216 | 77.8% |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Comparison | Factor changed | Factor held fixed | Use |
|---|---|---|---|
| PP vs. QP | factor 1 | factor 2 at | primary pair (main study and follow-ups) |
| PQ vs. QQ | factor 1 | factor 2 at | additional profiles (main study) |
| PP vs. PQ | factor 2 | factor 1 at | main study; not in the main-text estimates |
| QP vs. QQ | factor 2 | factor 1 at | main study; not in the main-text estimates |
| Candidate | / endpoints | Recorded checks | Outcome |
|---|---|---|---|
| Accepted: main study, t16, persona gender field | : Also, I am male. : Also, I am female. | Validator passed; neither reviewer flagged a fatal problem | ✓ PASS used in the archived study |
| Built, not selected: main study, t16 second factor (source Recently started planning to travel solo) | : This time I plan to go by myself; I want to try traveling solo. : This time I plan to go with two friends, traveling together. | Validator passed; no fatal flag; not selected by the frozen two-factor assignment rule | NOT SELECTED in no run input |
| Rejected: initial randomized study pipeline, persona PDR_PERSONA_017 | : Consults professionals for advice before making investment decisions : Does not consult professionals for advice before making investment decisions, preferring to rely on own judgment | Both LLM verifiers failed gate G2 (single dimension) | ✗ FAIL not sent to human review |
| Split review: main study, t23 | : I currently live in Hangzhou, in a rented studio apartment in the city center, fairly close to my company. : I currently live in Nanjing, in a rented studio apartment in the city center, fairly close to my company. | One reviewer flagged a contradiction, the other did not; exclusion required both | RETAINED flag registered |
| Disputed: factor follow-up, t46_x1 | : On spending: we are quite willing to invest in our child’s education, and generally don’t care much about price when signing up for classes or buying learning resources. : On spending: we hope to save as much as possible on our child’s education, and generally pick the cheaper options first when signing up for classes or buying learning resources. | Designed as a soft preference; one validator judged the endpoints mutually exclusive, the other agreed with the design | DISPUTED run, not in the primary factor contrast |
| Cell | Roles (tools) | Run status | ||
|---|---|---|---|---|
| H00 | coupled | coupled | one persistent executor ( search , read , write_report , submit_report ) | RUN : main study, both follow-ups, H-R; also the exploratory study’s report-tool setups |
| H10 | handoff | coupled | W-controller ( write_report , submit_report , delegate_acquisition ); a fresh A-realizer per delegation ( search , read ) | NOT ANALYZED : earlier development study |
| H01 | coupled | handoff | upstream executor ( search , read , write_report , handoff_to_finalizer ); separate Finalizer ( submit_report , request_upstream ) | NOT ANALYZED : earlier development study |
| H11 | handoff | handoff | W-controller, A-realizer and Finalizer as above; the controller has neither retrieval nor submission | NOT ANALYZED : earlier development study |
| Implementation (studies) | Tools (arguments) | State visible to the agent | Drafts and final |
|---|---|---|---|
| Initial loop (initial randomized) | search(query) , read(url) , write_report(content) , submit() with no argument | One context: system prompt, a scripted clarification exchange carrying the user condition, the task request, own history, raw tool returns | Each write_report overwrites one draft; submit() delivers the current draft, so the final equals the last saved draft by construction |
| H00 report tool (exploratory S00/S10, main, both follow-ups; H-R) | search(query) , read(url) , write_report(content) , submit_report(content) | One context: system prompt, one user message, own history, raw tool returns (search: 6 results with 400-character snippets; read: first 6,000 characters of normalized page text) | Each write_report is a new saved version; the final is the submit_report content or a text-only reply, a separate object from the drafts |
| Direct submission (exploratory S01/S11) | search(query) , read(url) , submit_report(content) | As H00, without a draft tool | No drafts; final as in H00 |
| H-W file workspace (harness comparison) | search(query) , read(url) , shell(command) , view_file(path, offset, length) , write_file(path, content) , edit_file(path, old_text, new_text) , submit_file(path) | One root context plus a persistent sandboxed /workspace ; the shell has no network, so external material arrives only through search and read | Files are snapshotted after every tool call; /workspace/deliverable.md forms the report-draft chain; the final is the exact file text read at submit_file , or a text-only reply |
| H-D delegation (harness comparison) | H-W tools plus delegate_batch(tasks) , where each task carries { assignment , input_files }; workers: search , read , shell , view_file , write_file , edit_file | Root as in H-W. A worker sees only its assignment text, a read-only snapshot of the selected files, and a private scratch directory; not the root history, sibling workers or other files | As in H-W; only the root can submit. Worker replies return verbatim in a fixed template. Offered, but no recorded run used delegation |
| Family | Requested ID | Protocol | Lowest setting (main, harness comparison) | High setting (both follow-ups) |
|---|---|---|---|---|
| GPT | gpt-5.4 | chat completions; responses in the follow-ups | reasoning_effort=none (reasoning tokens not visible) | reasoning_effort=high ; dated echo gpt-5.4-2026-03-05 accepted |
| Claude | claude-sonnet-5 | chat completions | no reasoning field (no extended thinking by default) | reasoning_effort=high |
| Gemini | gemini-3.8-flash | chat completions | reasoning_effort=low (lowest accepted) | reasoning_effort=high |
| Grok | grok-4.6 | responses | reasoning.effort=none ; still reports reasoning tokens | not evaluated at the higher setting |
| DeepSeek | deepseek-v4-pro | chat completions (provider API) | thinking: disabled ; zero reasoning tokens required per call | thinking: enabled , reasoning_effort=high |
| Check | Recorded evidence | Outcome |
|---|---|---|
| Harness comparison: before the first main call | ||
| Same input | Identical user message in all three harnesses, P and Q differing only in the frozen endpoint (A01; no recorded mismatch at the end); H-R shows the agent the same messages as the frozen H00 loop on a simulated response sequence (A09) | ✓ PASS |
| Endpoints: no forced final | A text-only reply is a natural final; budget end makes no extra call and never promotes a draft (A13, A15) | ✓ PASS |
| Sandbox: process environment | The sandbox’s first process inherited the orchestrator environment, including a provider credential, and a command line naming host paths and the run identifier, which contains the user-condition label (D19). Fixed before any main call; no development shell call had read it | ✗ FAIL |
| Isolation and budget | After the fixes: no path traversal or host links, secrets, other-arm text or shell network (A17–A20, A28); a worker sees only its assignment and selected files, cannot submit or delegate, and returns verbatim (A21–A25); one shared ledger that the shell cannot bypass (A29, A30, A33). In all, 219 tests for 64 requirements passed and 52 injected mutations were detected | ✓ PASS |
| Harness comparison: during and after the main runs | ||
| (a) Studies and harnesses: differences from the main study | |
| Setting | System message, user information, and tools |
| Initial randomized | Own system message, which asks the agent to keep an initially empty plan draft current (do not wait until the end to start writing); user information given as a clarification exchange before the task request, with the factor answer omitted in the task-only arm; the request adds fixed research and decision-summary blocks; tools search , read , write_report , submit (run-time schemas not located) |
| Exploratory | P0, or P1 (a rewording of P0); the task text begins with a reference-date line; setup H_DIRECT offers no write_report |
| Main | P0; template; in two-factor runs both endpoint texts fill the last slot; tools as in the box above |
| Follow-ups | As main; the reasoning follow-up uses the main-study cards, the factor follow-up its own cards |
| Harness H-R | As main |
| Information | Request evaluation | Document extraction |
|---|---|---|
| Task text given to the agent | ||
| Factor: endpoint texts as ALPHA/BETA, type, scope, unknown cases | with request_support_rubric | with final_action_rubric |
| Request texts (query or URL) of the run | up to 32 per packet, 12 with two factors | |
| Document text | one document, numbered paragraphs | |
| Tool returns; other document versions | ||
| Shared background, user message, assigned endpoint, run, model, condition, paired run | (checked at packet build) | (checked at packet build) |
| Field | Allowed values (checked at ingest) | Use in the decision |
|---|---|---|
| axis_id | a factor id of the packet | groups items by factor |
| kind | ACTOR_ADVICE , USER_RESTATEMENT , THIRD_PARTY_OR_BACKGROUND | only ACTOR_ADVICE counts; restatements set the basis |
| status | MAIN , BACKUP , COMPARE , EXPLICIT_REJECT , MENTION_ONLY , DEFERRED , UNKNOWN | only MAIN counts |
| side | ALPHA_SIDE , BETA_SIDE , BOTH , NEITHER , UNKNOWN | restored to P_SIDE / Q_SIDE ; BOTH counts for both |
| source_ids , entity_key , condition_status | existing paragraph ids (non-empty list); non-empty string; EXPLICIT_NONE , EXPLICIT_CONDITION , UNDETERMINED | validation only |
| free text | condition , exception , explicit_priority , actual_time_or_budget | not used |
| Category | Frozen handling |
|---|---|
| No request | No request packet; request readout undefined and retained as a missing behavioral state, not zero. |
| No identified request | Request readout undefined; not imputed as neutral. |
| No document for a role | Missing document recorded explicitly; not converted to None . |
| No qualifying main advice | Unknown ; None is reserved for extracted main advice that expresses neither endpoint. |
| Disputed reads | Different labels or evidentiary bases become Unknown . |
| Unreadable response | Failed read remains missing; no semantic repair or choice between attempts. |
| Primary | Auxiliary | Exploratory study | |
|---|---|---|---|
| Model | deepseek-flash, provider API | gpt-5.4, claude-sonnet-5 via a local gateway; served identity not verified | deepseek-v4.1-flash, hosted plan of another provider |
| Settings | thinking disabled, zero reasoning tokens checked per call; one user message; max_tokens 16384 | provider-default reasoning, not recorded; otherwise same; timeout 900 s, not 600 s | thinking disabled; system and user message; max_tokens 8192 |
| Prompts | A_READ_PROMPT ; item extraction | same packets | requests: batched events with three preceding requests as context, field mentioned_side ; documents: same item text with P/Q slots, task-level axes, limit 12 then 30 |
| Two reads | orientation by hash, inverted in read 2; label only if both agree | same packets and rule, own labels | read 2 swaps conditions and side definitions; each read analyzed separately, agreement as sensitivity |
| Coverage | main-study packets | selected main-study packets | exploratory packets |
| Review | Object and reviewers | Items | Recovered | Source | Experiment |
|---|---|---|---|---|---|
| Construction review | Candidate cards; CR1, CR2; 5 criteria | 94 | 94 2 | RAW_RECOVERED | COMPLETED |
| Construction pilot | Four fictitious task fixtures; PR1, PR2; 7 questions each | 28 | 28 2 | RAW_RECOVERED | COMPLETED (diagnostic) |
| Real-task certification | Planned real-task questionnaire; marked do-not-send | – | – | DESIGN_ONLY | NOT_RUN |
| Directional audit | Exploratory-study outputs; R1, R2, and ADJ on disagreements | 13 | 13 2 + 6 | RAW_RECOVERED | COMPLETED |
| Criterion | CR1 | CR2 | Agreement | Kappa |
|---|---|---|---|---|
| H1 faithful to source | 69 / 24 / 1 | 94 / 0 / 0 | 0.7340 | 0.0000 |
| H2 same core factor | 67 / 27 / 0 | 94 / 0 / 0 | 0.7128 | 0.0000 |
| H3 one factor changed | 51 / 43 / 0 | 93 / 1 / 0 | 0.5532 | 0.0252 |
| H4 clear and plausible | 94 / 0 / 0 | 94 / 0 / 0 | 1.0000 | undefined |
| H5 no answer leakage | 88 / 6 / 0 | 93 / 1 / 0 | 0.9255 | |
| All five passed (cards) | 39 | 92 | 0.4149 (all five equal) | – |
| Fixture | R1 src | R1 cf | R2 src | R2 cf | R3a src | R3a cf | R3b | Same |
|---|---|---|---|---|---|---|---|---|
| 1 clean | 3/7 | |||||||
| 2 polarity defect | 5/7 | |||||||
| 3 binding defect | 5/7 | |||||||
| 4 mapping defect | 6/7 |
| Item | Packet | Axis view | Instrument | R1 | R2 | Stage 2 | ADJ | Final |
|---|---|---|---|---|---|---|---|---|
| A1 | ccdb4cfa | D05.G1 F | P, P UNK | P | P | no/no | – | P |
| A2 | d707606e | D05.G1 F | P, P UNK | P | P | may/no | Q (no) | P |
| A3 | 51f1e541 | D05.G1 A | Q, Q Q | Q | Q | – | – | Q |
| A4 | 2b08dccd | D05.G1 A | P, P P | P | P | – | – | P |
| A5 | 17a324cc | D05.G1 A | P, P P | P | P | – | – | P |
| A6 | 86efaf6a | D05.G1 F | Q, Q UNK | Q | Q | no/no | – | Q |
| Arm | Event | Tool | Request or argument | Return | Reads 1 / 2 |
|---|---|---|---|---|---|
| P | e0002 | search | car insurance comprehensive reform, theft insurance, vehicle damage insurance, roadside assistance, value-added services | 6 results | NEITHER / NEITHER |
| P | e0004 | search | PICC, Ping An, China Pacific, car insurance, roadside assistance, remote areas, Tibet, Yunnan, self-drive travel | 6 results | NEITHER / NEITHER |
| P | e0006 | search | new car first-year insurance, driving habits, parents’ car, accident-free, discount, NCD coefficient | 6 results | NEITHER / NEITHER |
| P | e0008 | write_report | (content: 5020 chars) | Draft saved. | — |
| P | e0010 | submit_report | (content: 4991 chars) | Report submitted. | — |
| Q | e0002 | search | car insurance, premium reform, vehicle damage insurance, whole-vehicle theft insurance, spontaneous combustion, water wading, third party cannot be found | 6 results | NEITHER / NEITHER |
| Arm | Observation | Read 1 / read 2 | Program decision | State |
| P | 3 requests | NEITHER / NEITHER for each request | identified, no side | NONE |
| P | draft W1 | P_LEAN (5 / 0 / 0) / P_LEAN (6 / 0 / 3) | agree, item basis | OWN |
| P | final report | P_LEAN (4 / 0 / 2) / P_LEAN (4 / 0 / 7) | agree, item basis | OWN |
| Q | 3 requests | NEITHER / NEITHER for each request | identified, no side | NONE |
| Q | draft | — | no document | NO_DOCUMENT |
| Q | final report | Q_LEAN (0 / 3 / 0) / Q_LEAN (0 / 2 / 0) | agree, item basis | OWN |
| Case | Decisive text (excerpt) | Reads 1 / 2 | Decision |
|---|---|---|---|
| Task 28, Claude, arm P, final report; factor 2, residence (P Suzhou, Q Nanjing) | [p022] ### 3. Offline events ( Suzhou /Yangtze River Delta) … Follow AI Meetups and university AI Talks in Suzhou /Shanghai (e.g., NLP-related salons in Shanghai/ Nanjing … | each read: main items P, P and Q; BOTH_APPLIED twice | BOTH |
| Task 13, Grok, arm Q, final report, a refusal; household (P married with a daughter, Q single) | [p002] Please consult qualified clinical staff such as an endocrinologist or a general practitioner and a registered dietitian as soon as possible, and have them draw up a plan based on your test results, medication, and risk of complications. … | main items (3 and 2) all on neither side; NO_DIRECTION twice | NONE |
| Task 13, Grok, arm P, final report, a refusal (same pair) | [p002] Please consult a qualified endocrinologist or your attending physician, and have them draw up a plan based on your test results and medical history; do not use any online content as a substitute for proper medical care. | no items in either read; basis NO_EVIDENCE twice | UNKNOWN no main advice |
| Task 14, DeepSeek, arm Q, draft W1; sex (P female, Q male) | [p010] Age/sex: 20-year-old male , strong ability to recover; this is your biggest advantage. | one item each: USER_RESTATEMENT, MENTION_ONLY (Q); basis RESTATEMENT_ONLY twice | UNKNOWN restatement only |
| Task 3, GPT, arm P, final report; occupation (P ad copywriter, Q product manager) | [p160] Interviewees can come from: - your current clients in the advertising industry, or upstream and downstream partners; - Shanghai entrepreneur communities; - Hangzhou e-commerce/brand practitioners; - brand/platform/investment-circle people you meet on business trips to Beijing. | read 1 (ALPHA = Q): 8 main items, all P, incl. p160; P_LEAN. Read 2 (ALPHA = P): 6 P and 2 BOTH, incl. p160; BOTH_APPLIED | UNKNOWN reads disagree |
| Task 31, Gemini, arm Q, final report; age (P about 30, Q about 45) | [p001] # Anti-aging and brightening skincare plan and product buying guide for 45-year-old combination skin [p014] [First choice] Proya Double-Resist Essence, 3rd generation (morning early anti-aging and brightening) [p016] [Backup] L’Oreal “Little Honey Jar” face cream (light version) …better fits the entry-level anti-sagging needs of a 45-year-old | p001 MENTION_ONLY (Q); p014 MAIN (P); p016 BACKUP (Q); other main items neither; P_LEAN twice | OTHER |
| Case | Assigned endpoint | Opposite-side request | Main recommendation |
|---|---|---|---|
| Parenting DeepSeek | International schooling track | double-reduction policy family education after-school tutoring restrictions subject training impact (request 17 of 39; “double-reduction policy, family education, tutoring limits”) | A home plan “highly isomorphic to IB PYP”; bilingual-school budget tiers |
| Career Claude | Business-management track | technical-to-management transition leadership book recommendations and two similar requests (requests 1, 2, 11 of 13; “technical-to-management leadership books”) | Plan titled “From algorithm engineer to business manager”; technical course only as a backup |
| Exchange study GPT | Help with domestic job search | “official LSAC JD degree first degree in law official” (request 59 of 76; original in English) | Destinations ranked by domestic job-search value and transferable courses |
| Travel DeepSeek | Budget cap ¥15,000 (vs. ¥6,000) | Southeast Asia backpacking two weeks Vietnam Thailand budget travel cost budget (request 2 of 13; “SE Asia backpacking two weeks, budget travel costs”) | Budget “about ¥8,000–9,500, well below your ¥15,000 cap” |
| Study | Model, interface | Decisions | Requests | Saved drafts | Final |
|---|---|---|---|---|---|
| Main | Gemini, report tool | 1 | 0 | 0 | OWN |
| Main | Grok, report tool | 3 | 2 | 0 | OWN |
| Reasoning follow-up | GPT (high), report tool | 31 | 30 | 0 | OWN |
| Harness comparison | Grok, H-D | 5 | 3 | 1 | OWN |
| Comparison | Observed process change | Interpretive boundary |
|---|---|---|
| Reasoning setting | Higher reasoning generally lengthens execution and increases acquisition activity. | No uniform corresponding improvement in the coarse Final directional readout was established. |
| Report-tool vs. file interfaces | File-based interfaces change drafting and submission behavior. | In file-submit interfaces the saved report can be the submitted report, so draft–Final identity may be structural. |
| Delegation-enabled interface | The interface makes isolated worker delegation available. | No recorded run used delegation; delegated-context behavior was therefore not tested. |
| Draft timing | Main reports usually appear late relative to evidence acquisition in the eligible archived subsets. | This does not imply that all workspace use or private reasoning occurs late. |
| Follow-up | What it tested | Frozen outcome |
|---|---|---|
| Rule search | Whether simple observable features define a portable loss-risk rule. | NO RULE SELECTED |
| Factor design | Whether factor properties identify where directional expression appears. | NOT IDENTIFIABLE |
| Locked loss rule | Whether a development association transfers unchanged to confirmation data. | NOT TRANSPORTED |
| Checking step | Whether the tested reminder policy establishes a benefit after its trigger. | NO ESTABLISHED BENEFIT |
| Counterfactual return | Proposed intervention on a tool return, stopped at construction. | NOT RUN |
| Approach (aim) | What exists (scope) | Experiment status | Source status |
|---|---|---|---|
| G1: Representations; persona echo (score behavior, not restated background) | Text-encoder tests on constructed sentences across multiple encoders; evaluator screening with restatement-conflict items. | CORRECTED_OR_VOID (encoder) COMPLETED (screening) NOT_RUN (hidden states) | RAW_RECOVERED; DESIGN_ONLY (hidden states) |
| G2: Source links, graphs (trace evidence into reports) | Return-to-draft census (54 runs); two 8-run provenance graphs; edge ontology. | COMPLETED (census, graphs) NOT_RUN (semantic graph) | RAW_RECOVERED; DESIGN_ONLY (ontology) |
| G3: Multi-level readouts (where use stops) | Design with five levels; two piloted separately (G4). | NOT_RUN | DESIGN_ONLY |
| G4: Prompted recall, production (stating vs. using) | July pilots with repeated multiple-choice probes, agent runs, and forced deliveries from saved prefixes. | COMPLETED (pilots) | RAW_RECOVERED |
| G5: Reference-report scoring (score against both users) | Cross-scoring design; one-user rubric; cross-report unit matching (35,640 judgments). | NOT_RUN (cross-scoring) RETIRED (matching) | DESIGN_ONLY; RAW_RECOVERED (matching) |
| G6: Direct compliance, holistic judges (judge fit or violation) | Compliance measurement (116 units, intermediate study); pairwise holistic judge (65 pairs). | RETIRED CORRECTED_OR_VOID (one pipeline) | RAW_RECOVERED; SUMMARY_ONLY (holistic judge) |
| Unit | Link rule | Linked / total | Share |
|---|---|---|---|
| Return units before first draft (54 runs) | URL or verbatim 20 chars | 662 / 9,666 | 0.0685 |
| search results | same | 550 / 9,356 | 0.0588 |
| web-page reads | same | 112 / 310 | 0.3613 |
| First-draft units (54 runs) | same | 287 / 11,241 | 0.0255 |
| strict fragments | verbatim 40 chars | 57 / 11,241 | – |
| Final documents, graph 1 (8 runs) | mechanical edges | 0 / 8 | – |
| Component | Archive location | Identity key |
|---|---|---|
| Tasks, users, factors | PDR-Bench/data/task_data/tasks_zh.jsonl ; /USER_PROFILES.jsonl , /FACTOR_CARDS.jsonl | PDR commit 5b43f9f188c7 ; profile_sha256 ; factor_id |
| Contract, routes, code | ROUND2_PROSPECTIVE_V1/HARNESS_CONTRACTS/H00.json ; */ACTOR_ROUTE_REGISTRY.json ; */DESIGN_FREEZE.json ; */scripts/ | file sha256 (H00 50eb581e… ); per-file code hashes |
| Run records | exp_v1/runs/formal_v1/ ; runs/natural_discovery_v1/MAIN/ ; /RUNS/ | run_uid , record_sha256 , slot_id |
| Evaluator records | /READ/ ( packets/ , raw_primary/ , raw_aux/ , ingested/ ) | packet_id , payload sha256 |
| Row tables, texts | ZERO_CALL_ANALYSIS_V1/ROW_TABLES/ , TEXT_INDEX/ ; HARNESS_TRANSPORT_V1/ROW_TABLES/ | MANIFEST.json ; artifact_text_sha256 |
| Analyses, tables, figures | */FINAL_ANALYSIS/ ; exp_v1/runs/formal_v1/_FORMAL_ANALYSIS.json ; experiments_overleaf/plot_sources/ | plan sha256; % SOURCE comments |
| Original issue | Scope of impact | Current handling | Old conclusion |
|---|---|---|---|
| Split gateway replies: dropped Claude tool calls. | Main study: 4 attempts. Earlier study: all 8 Claude runs. | Replies merged; 4 attempts replaced; earlier runs excluded, tables regenerated. | VOID |
| Invalid evaluator batches. | Earlier study: 627 readings by an uncalibrated evaluator; 366 item readings with reversed rubric (agreement 33.7% aligned vs. 52.1% presented). | Quarantined; items rerun with the corrected builder, later frozen for the reported studies. | VOID |
| BOTH and NONE merged as “attenuation”. | Earlier study: 11 of 12 “net-zero” first drafts were BOTH. | States never merged; not used to describe loss. | RETRACTED |
| UNKNOWN read as inactive. | Earlier study: re-entry through unknown gaps (one task 6/8 to 2/8). Exploratory: 368 uncertain requests coded “no side”. | Re-entry needs a fully readable gap; uncertain kept separate; reported numbers use the corrected analysis. | CORRECTED |
| Subset read as propagation: “63/64 (98.4%)”. | Earlier study: all 64 finals in the subset were on the assigned side. | Full table: 63/106 identical, 70/106 with zero states merged, 1/106 reversed; joint reading only. | WITHDRAWN |
| Risk-rule lookup with a wrong key. | Main study: three continuous features always missing. | Corrected lock on the discovery split. | UNCHANGED : no rule |