Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents
Organizations: IdeaHorizon Team
Abstract
We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that lineage's failure mode: under pressure to finish, they fabricate, skip, or smooth over. Our design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable. We encode research discipline as mechanically enforced laws (commitment before measurement; unforgeable freezing; reports are not facts; evidence persists but verdicts do not; negative results are first-class; mechanical questions to the framework and semantic judgment to the model), organized around three time horizons: a minimal set of research nodes within a run, an inquiry contract with frozen closure conditions and a hash-chained artifact ledger within a project, and a two-tier knowledge base with promotion by rewriting across projects. This is a system description written under one rule: each mechanism appears in exactly one place, with the invariant it enforces, the failure it prevents, the way it is realized, and the cost it imposes. It covers the node contract, the write-path gates, the two-tier memory, and the runtime substrate. Three traces walk real failure attempts through the mechanisms that catch them, and two closed campaigns are included as worked illustrations rather than as an evaluation. We report no benchmark: a process-integrity suite that would support quantitative comparison is under construction, and what we can measure today is only the operating cost of the machinery.
Figures & tables
| Where the rule lives | Example | Observed strength |
|---|---|---|
| In the prompt | “You must freeze the preregistration before running.” | Weakest. Under pressure the rule is restated and then worked around (§ 1 ). |
| Post-hoc check | After the run, a check tests whether a freeze happened. | Weak. Found too late; killing the run does not undo what ran. |
| Unrepresentable | The dispatch gate refuses to start experiment unless a frozen preregistration exists. | Strong. The illegal state cannot be written. |
| Law | Failure it prevents | Described in | |
|---|---|---|---|
| L1 | Commitment before measurement. A question with a written proposition is a hypothesis; closure conditions are frozen before work starts. | Thresholds invented to pass a falsifiability gate; a verdict chosen after the result was seen. | § 5.1 |
| L2 | Freezing is irreversible and unforgeable; revision leaves a trace. | A refused freeze bypassed by re-saving with frozen:true ; a frozen plan silently edited. | § 5.2 |
| L3 | Reports are not facts. Anything mechanically computable is computed by the framework from ledgers, not read from the model’s self-report. | status: verified written by hand with no verifier call; “11 of 12 fulfilled” with zero matching keys. | § 5.1 , § 8 |
| L4 | Evidence persists; verdicts do not. Detectors record; only the reviewer and the adjudicating node judge; judgments are recomputed, never stored. | A check layer that killed runs 87% of which never recovered; a node adjudicating its own experiment. | § 4.6 , § 5.3 |
| L5 | Negative results are first-class. Refuted stays refuted in the manuscript; withdrawn questions enter Limitations; dead ends become knowledge cards. | “H3 is refuted” asserted for a quantity never measured; a null result smoothed into a trend. | § 5.3 , § 6.2 |
| L6 | Mechanical questions to the framework, semantic judgment to the model, in both directions. | An orchestrator busy-waiting 47 turns on whether a node should continue; a quota of five citations standing in for relevance. | § 4.5 , § 4.6 |
| Node | Distinctive mechanisms |
|---|---|
| Analysis ( hypothesis ) | Four axioms with no “hypothesis” in them; a question with a proposition is a hypothesis, one without is not, with no yes/no switch to misdeclare. Seven pre-freeze audits: goal alignment (multi-pillar requests may not collapse to one), definition lock (the user’s ontology is copied verbatim), threshold grounding (every numeric threshold needs a stated source or must become qualitative), comparison protocol, resource feasibility, cost instrumentation. Exploratory questions may not list “nothing found” as a failure condition. Sole writer of research_state . |
| experiment | 55-word system prompt; knowledge lives in 30 loop hooks and 7 skills loaded after scope is known. execution_params are compared key-by-key with the frozen preregistration before anything runs. No proxy or toy may be reported as the original task; the only exits are probe, acquire, ask, or infeasible. Tree-sitter semantic analysis of shell commands (dynamic execution refused). Irreversible-window recovery for external job submission that trusts only the pre-submission intent record. Reproducibility snapshot; three-part evidence (log, clean, raw with SHA-256); mandatory sediment attempt or an auditable “no sediment” addendum. |
| writing | Two stages, plan then prose; the gate asks the ledger, not the disk: zero discharged closure conditions blocks writing and produces a material-insufficiency report; partial discharge writes with material_gaps . Mechanical blockers are separated from heuristic gaps, which are routed to the node that can fill them. Methods may not be completed from domain common sense; conflict-of-interest and authorship are not invented. Phantom claim identifiers fail the citation-integrity hook; internal identifiers may not appear in the clean PDF. TeX exit code 0 is not deliverable: layout gates on overfull boxes, unresolved references, and oversized floats. Freeze is the author’s signature; the orchestrator cannot sign for it. |
| literature | Four-mode contract (landscape, targeted lookup, contradiction check, method lookup) with per-mode required outputs and turn budgets, so that a “find five papers” request cannot run as a full survey. Research gaps may not be phrased as gaps in the knowledge base. Writes facts (chunks, concepts with external anchors) but never assertions. Author wiring is a hard completion gate; missing authors are recorded, not invented. |
| data | Planning gate with a three-role loop (requirement analyst, designer, critic); no asset is produced before an approved plan. Two authorities (plan-bound and request-bound) compiled into a locked work order. Structured terminal states ( completed , needs_input , externally_blocked , fatal ). May not refuse a redirected request on the grounds that it is “a computation task” (otherwise a two-sided deadlock). |
| postprocess | Provenance binding by three hashes: source artifact, rendering script, output file, recomputed by consumers; a file not written by this sandboxed execution is refused. Verdict fields removed entirely; only two mechanical facts remain, and the absence of a vision reviewer is a readable fact that downstream must disclose. Read-only scientific semantics: no aggregation, smoothing, outlier removal, or inference of graph structure. Generative images must be marked non-evidence-bearing. |
| # | Band | Source | Changes |
|---|---|---|---|
| 1 | node system prompt, rules, guidelines | the node contract | never |
| 2 | index of enabled skills (bodies on demand) | skill registry | never |
| 3 | workspace, provenance and blocker rights | framework | never |
| 4 | instruction layers: organization, profile, project | session-frozen snapshot, hashed | per session |
| 5 | research laws (constitution) | MEMORY.md | rarely |
| explicit prefix boundary marker | |||
| Layer | What it is | Strength and failure mode |
|---|---|---|
| guidelines | a soft recommendation in the contract | weak; violation is free |
| rules | a hard-sounding sentence in the contract | weak; “must” is not a mechanism |
| skill | a reusable piece of craft, shared across nodes | medium; the model must load it |
| hook injection | a mechanical fact placed in front of the model | strong; cannot be missed, can be ignored |
| tool validation | bad arguments refused at the call, with guidance | strong; scoped to one call |
| write gate | a non-conforming record cannot be written | strongest; misfires block legal states |
| Node | Runs | Turns | Tokens (M) | of which cache reads (M) |
|---|---|---|---|---|
| experiment | 11 | 642 | 68.31 | 55.29 |
| reviewer | 20 | 459 | 42.54 | 37.36 |
| writing | 10 | 326 | 33.26 | 26.57 |
| orchestrator | 1 | 143 | 23.92 | 12.07 |
| Analysis | 9 | 184 | 15.71 | 14.28 |
| postprocess | 9 | 172 | 11.31 | 10.11 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Node | Role | Post-run flow | Tools | Local skills | Required outputs / callable nodes |
|---|---|---|---|---|---|
| Analysis ( hypothesis ) | producing | full | 34 | 3 | preregistration, research plan, research state; calls literature |
| experiment | producing (high risk) | review, curate | 60 | 7 | experiment log, clean results, raw results; calls data |
| observation | producing | full | 20 | 0 | observation log with search protocol; calls literature , data |
| derivation | producing | review, curate | 28 | 0 | derivation log; calls literature , data |
| writing | producing | full | 28 | 7 | preflight plan, manuscript, validation report; calls postprocess |
| literature | service | none | 25 | 1 | per mode: survey report and index, or evidence package |
| Channel | Trigger | Budget |
|---|---|---|
| constitution | every turn, all nodes | source capped at 2048 bytes |
| situation | each run start | 3072 bytes |
| onboarding slice | turn one, all nodes | 4096 bytes, matched on applies_to |
| first-use tool brief | first call of a given tool | at most 2 entries, 512 bytes |
| knowledge-base opening injection | project opening | constant: 10 conclusions, 5 dead ends |