When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing
Organizations: The Hong Kong University of Science and Technology (Guangzhou), China
Abstract
Modern learned systems increasingly combine learned components with search, repair, or external solvers. Benchmarks often measure the resulting end-to-end system, while scientific claims may concern only one component, creating an attribution problem: evidence can fail to support the requested component-level claim while still supporting a positive conclusion about the larger system. Existing evidence-to-claim methods primarily calibrate claim strength. We argue that composite systems require a second dimension: scientific subject. We address this problem with subject-typed claim licensing, which separates weaker conclusions about the requested subject from positive but non-substitutive credit about another subject. We instantiate this idea in SCOPE-Routing for preference-conditioned multigraph routing. Non-authors reproducibly apply the declared semantics; held-out review yields fewer reference-relative upward deviations than unstructured review, while the difference from a strong evidence checklist remains unresolved; and a controlled routing study shows that score-optimal and claim-eligible methods can differ while valid hybrid-system credit is preserved. These results motivate treating claim strength and scientific subject as distinct dimensions of evidence-based evaluation.
Figures & tables
| Level | Licensed learned-primary conclusion | Incremental evidence |
| Scalar decision quality | Primary-only or intervention-free regret against a declared comparator. | |
| Feasibility-audited routing | Exact feasibility and invalid-route accounting. | |
| Identity-preserving multigraph routing | Edge identity, bundle-aware actions, and collapse checks. | |
| Preference-conditioned routing | Preference sampling, coverage/hypervolume, and held-out preference shift. | |
| Intervention-audited constrained routing | Constraint masks, fallback/repair taxonomy, and primary versus end-to-end decomposition. | |
| Dynamic robust routing | Drift stress, route-change evidence, tail regret, and latency. |
| Outcome | Careful review | Evidence checklist | SCOPE-Routing |
| Agreement with operational reference | 24/36 | 27/36 | 31/36 |
| Reference-relative upward deviation | 8/36 | 5/36 | 2/36 |
| Reference-relative downward deviation | 1/36 | 3/36 | 4/36 |
| Median time per claim | 18.1 min | 13.4 min | 15.8 min |
| Decision rule | Selected row | Regret | Key evidence | Licensed interpretation |
| Score only | Fallback-Heavy Safe | 0.011091 | 86.1% backup rescue | Hybrid credit retained; no learned-primary . |
| Target , then regret | Shock-Safe Online Repair | 0.016189 | Dynamic and attribution evidence complete | learned-primary. |
| Coverage-oriented diagnostic | Preference Hypernetwork | 0.020865 | HV | ; dynamics incomplete. |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Prior line of work | Primary output | Composition role in claim licensing |
| Argument-based and AI-evaluation validity | Evidence-backed interpretations, uses, and claims. | Supplies validity arguments; the contract operationalizes one bounded row-level conclusion decision. |
| Evidence-licensed calibration and claim replay | Calibrated assertion scope, maximal licensed frontiers, replayable inference, and semantic stability. | Supplies evidence-relative claim semantics; subject-typed licensing separates requested-subject answers from non-substitutive credit. |
| Structured assurance cases | Traceable claim–argument–evidence graph. | Supplies explicit support structure that can be encoded as claim obligations. |
| Claim-evidence retrieval and assessment | Evidence links, support labels, justifications, or graded overstatement scores. | Supplies evidence and support judgments consumed by obligation evaluators. |
| Datasheets and model cards | Documentation of data, models, intended uses, and limitations. | Supplies inspectable evidence conditions and intended-scope fields. |
| BenchmarkCards and benchmark documentation | Standardized benchmark metadata and risk reporting. | Supplies benchmark-level evidence records and scope declarations. |
| Adjudicator | Evidence diagnosis | Claim order | Source semantics | Incomparable credit | Primary output |
| Regret-only ranking | Scalar score | Absent | Absent | Absent | Score-ordered row. |
| Missing-field checklist | Structured gap prompts | Reviewer-authored | Prompted or implicit | Reviewer-authored | Gap record plus reviewer-selected claim. |
| Designated-claim assessment | Evidence links or support scores | Usually absent | Source-bound or implicit | Usually absent | Support label, score, or justification. |
| Unstructured human review | Contextual | Implicit | Implicit | Implicit | Review prose and judgment. |
| SCOPE-Routing | Typed obligations | Explicit | Explicit | Explicit relation | . |
| Architectural role | Routing instantiation | Selection-task replacement |
| Decision quality | Route regret | Downstream selected-set objective or regret |
| Feasibility | Valid source–destination path | Budget, cardinality, or packing feasibility |
| Action identity | Parallel-edge identity | Item identity and duplicate/substitution handling |
| Coverage | Preference-space hypervolume | Coverage across budgets, preferences, or item strata |
| Attribution | Fallback solver or repair | Optimizer, repair, abstention, or fallback disclosure |
| Robustness | Drift, route change, tail regret | Distribution shift, set change, tail decision loss |
| Channel | Representative fields | Claim-validity role |
| Quality | scalarized regret, tail regret, oracle objective, comparator identity | Establishes C0 scoring evidence. |
| Feasibility | aggregate success, feasible-set-nonempty flag, conditional invalid-route count, hidden-repair flag | Determines whether C1+ feasibility-audited claims can be credited; rounded aggregate rates are not the gate. |
| Fallback | claim subject, fallback reason/rate/trigger, primary-only and end-to-end quality, fallback-credit flag | Separates learned-component evidence from hybrid rescue and safety-control credit. |
| Structure | edge identity, bundle id, retained-edge set, collapse-ablation id | Prevents simple-graph evaluation from certifying C2 multigraph claims. |
| Preference | preference vector, sampler, hypervolume, held-out shift split | Tests whether scalar gains preserve C3 preference-space behavior. |
| Dynamics | route identity, route-change W/T/L, drift severity, non-static-equivalence flag | Distinguishes dynamic learned behavior from static-equivalent rows. |
| Evidence source | Instantiated structure | Default admissible claim scope |
| Controlled multigraph generator | Score, edge identity, preference vectors, hard feasibility, and controlled drift. | Controlled C1–C5 claims under generator assumptions. |
| Same-contract RouteScope-Lite suite | Multidomain semisynthetic graphs, shared routing contract, 15 policy families, 10 seeds, telemetry-complete rows. | C1–C5 same-contract evidence; C6 additionally requires deployment evidence. |
| Public shortest-path bridges | Score transfer under a public shortest-path or contextual path contract. | C0 learned-scoring bridge credit; routing C3 additionally requires inherited feasibility and multigraph-identity obligations. |
| Diagnostic realistic slices | Realistic topology or dynamics with limited slice size and mechanism probes. | Mechanism diagnosis unless the slice is large enough and telemetry-complete enough for confirmatory realistic claims. |
| Native or external ML4CO adapters | Task-dependent objectives, metrics, and solver-interface behavior. | Compatibility, native-metric, efficiency, or out-of-contract claims; matching route objects and telemetry enable routing claims. |
| External public-claim corpus | Four blinded non-author annotation streams and executable labels of claims from published evidence snippets. | Contract-relative operational reference for the specified evidence packets. |
| Layer | What it tests | Main anti-failure role |
| Fixed packets | Whether non-author experts apply the supplied contract consistently. | Conditional applicability of supplied objects. |
| Object construction | Whether non-authors reconstruct inputs from prespecified sections. | Upstream author dependence. |
| Reviewer comparison | Whether the method changes agreement, overpromotion, rationales, and time. | Practical benefit and conservative error. |
| Same-contract suite | Whether target-claim eligibility changes the selected row. | Consequence under shared reporting. |
| Scope and rule stress | Whether source credit and distinct gates behave as intended. | Transfer discipline and functional distinctness. |
| Versioned artifact | Whether core and extension-stage evidence can be traced and reanalyzed. | Chronological and computational closure. |
| Probe | What is removed | Failure mode exposed |
| Regret-only ranking | Source scope, inherited gates, and record closure. | Selects Fallback-Heavy Safe in RouteScope-Lite because it has the lowest scalar regret; 86.1% fallback assigns the row to rescue-dominated hybrid evidence. |
| Missing-field checklist | Cumulative claim semantics and narrowing. | Can detect absent fields but cannot decide whether a row should be unsupported , narrowed , or credited at source-native scope. |
| Source-scope-only rule | Row-level telemetry gates. | Enforces bridge/native/out-of-contract transfer boundaries; fallback attribution, graph identity, dynamic behavior, and latency remain outside this reduced rule. |
| Leave-one-gate-out rules | One inherited gate at a time. | Each omitted gate enables a distinct false-support mode: fallback-heavy rescue, multigraph collapse, coverage loss, static equivalence, hidden latency, source upgrade, or unlocked artifacts. |
| Human-majority reference | Deterministic execution and manifest linkage. | Provides external claim-scope evidence, but is not a rerunnable row-level adjudicator. |
| Stratum | Count | Sampling rule |
| Learned routing and neural routing | 24 | Public routing papers with empirical route-quality claims. |
| Decision-focused and predict-then-optimize learning | 20 | Claims about downstream decision quality, regret, or solver-aware training. |
| ML4CO benchmarks and native solver suites | 22 | Claims about benchmark coverage, solver-interface evidence, or native metrics. |
| Multiobjective optimization and preference learning | 18 | Claims involving hypervolume, coverage, Pareto quality, or preference conditioning. |
| Evaluation science, documentation, and stress testing | 24 | Claims about benchmark validity, intended use, stress tests, or evaluation frameworks. |
| External bridge tasks and out-of-contract comparators | 12 | Claims where task scope differs from SCOPE routing but evidence is relevant to source-scope behavior. |
| Stage | Object or count | Reviewer check in the artifact |
| Candidate collection | Candidate claim records | Candidate IDs, source-family tags, claim location, inclusion/exclusion decision, and exclusion reason. |
| Corpus sample | 120 claims from 40 papers | Stable Claim Cards with anonymized source identifiers and evidence excerpts. |
| Human annotation | Four blinded human labels per claim | Raw annotation matrix, missing-gate selections, anonymized annotator metadata, blinding/COI fields, and validation report. |
| Consensus construction | Majority label with tie rule | Scripted majority vote and tie-resolution log. |
| SCOPE-Routing adjudication | Executable label and missing gates | Adjudicator output linked to evidence source, telemetry/prose evidence, and artifact fields. |
| Disagreement audit | 19 exact disagreements | Disagreement taxonomy and representative anonymized examples. |
| Field or outcome | Agreement | Interpretation |
| Comparator identity | 23/24 | Comparator extraction is highly reproducible. |
| Source family | 22/24 | Residual ambiguity concerns solver-interface credit. |
| Requested claim family | 21/24 | Dynamic wording creates the main family ambiguity. |
| Intended scope | weighted | High but non-perfect ordinal agreement. |
| Required/missing gates | mean Jaccard | Most evidentiary obligations are reconstructed similarly. |
| Final status | 20/24 exact; 24/24 within one | Remaining differences are adjacent. |
| ID | Claim family | Human majority | SCOPE-Routing | Reason for disagreement |
|---|---|---|---|---|
| ECC-012 | Shortest-path bridge | C1 | C0 | Humans credited feasibility from the shortest-path task; SCOPE-Routing kept source scope at C0 bridge because multigraph identity and fallback telemetry were absent. |
| ECC-019 | Learned routing | C5 | C4 | Claim used dynamic language, but route-change and drift-stress telemetry were incomplete. |
| ECC-027 | Multiobjective method | C3 | C2 | Strong scalarized results were reported, but hypervolume or preference coverage was missing. |
| ECC-038 | Benchmark paper | C6 | C5 | Humans interpreted broad realistic coverage as deployment-style evidence; SCOPE-Routing required non-author usability and deployment telemetry. |
| ECC-044 | Native ML4CO suite | C3 | out-of-contract | Native task instantiated multiobjective metrics but not source-destination multigraph routing. |
| ECC-061 | Decision-focused learning | C0 | C1 | Human majority credited only scalar scoring; SCOPE-Routing found explicit feasible-decoding evidence in the excerpt. |
| Policy family | Type | Evidentiary role |
|---|---|---|
| Dynamic exact oracle | upper bound | Provides oracle quality and feasibility reference. |
| Static exact anchor | comparator | Strong nonlearned comparator for regret and route-change checks. |
| Weighted shortest path | heuristic | Classical baseline with explicit weights. |
| -shortest-path heuristic | heuristic | Tests nonlearned multigraph candidate generation. |
| Simple-graph collapse | negative control | Tests C2 multigraph identity gate. |
| Learned full-exact scorer | learned scoring | Tests feasibility-audited controlled learned-scoring claims. |
| Policy family | Regret | Fallback | Feas. | HV | Licensed interpretation |
| Dynamic exact oracle | 0.000000 | 0.000 | 1.000 | 0.892 | Upper bound; not ranked as learned policy. |
| Fallback-Heavy Safe | 0.011091 | 0.861 | 0.997 | 0.669 | Standalone C5 not licensed; hybrid safety-control credit retained in . |
| GNN Pruned Top- | 0.012508 | 0.505 | 0.955 | 0.439 | C0; fallback and coverage gates fail. |
| Static exact anchor | 0.014874 | 0.000 | 1.000 | 0.716 | Comparator anchor, not learned routing. |
| Shock-Safe Online Repair | 0.016189 | 0.012 | 1.000 | 0.819 | C5-eligible under the declared claim-conditioned rule. |
| Adaptive Pruning Repair | 0.016761 | 0.016 | 1.000 | 0.799 | C4; repair improves pruning claim level. |
| Selection rule | Selected row | Reason |
| Regret only | Fallback-Heavy Safe | Lowest aggregate regret, without an attribution eligibility requirement. |
| Target C5, then regret | Shock-Safe Online Repair | Passes inherited C5 gates; eligible rows are then ordered by regret. |
| Target C5, then normalized cost | Shock-Safe Online Repair | Remains selected after the declared latency and cost normalization. |
| Coverage-first diagnostic | Preference Hypernetwork | Highest learned non-oracle HV, but incomplete C5 dynamics. |
| Quantity used in claim | Point estimate in main table | Required uncertainty check |
| Regret order between FallbackHeavySafe and ShockSafeOnlineRepair | 0.011091 vs. 0.016189 | Request-paired bootstrap interval for regret difference; used to characterize the regret-only ranking before claim-conditioned eligibility. |
| Fallback attribution | 0.861 vs. 0.012 | Seed- and request-level fallback intervals plus primary-only/end-to-end decomposition; narrowing depends on the aggregate gain remaining rescue-dominated. |
| Feasibility boundary | 0.997 vs. 1.000 | Aggregate interval and exact invalid-route counts conditional on nonempty feasible sets; the exact conditional counts determine C1. |
| Coverage comparison | 0.669 vs. 0.819 HV | Hypervolume interval; C3+ support uses nondegraded preference-space evidence alongside scalar regret. |
| Route-change and dynamic evidence | artifact-reported W/T/L | Route-change win/tie/loss and non-static-equivalence checks; C5 support uses observed dynamic behavior beyond inherited static-anchor routes. |
| Latency and cost-conditioned selection | artifact-reported p50/p95/p99 | Latency intervals, hardware record, normalization, weights, and tie handling; operational selection depends on these records. |
| Requested claim | Observed evidence gap | Evidence-supported effect |
| Identity-preserving multigraph routing | Evaluation collapses parallel edges to a simple graph | Multigraph claim blocked; lower structural or scalar credit may remain. |
| Preference-conditioned routing | Scalar quality reported without preference-space coverage | blocked; scalar and inherited lower-scope credit retained. |
| Dynamic robust routing | No route-change or drift evidence | blocked; supported static or attribution-audited credit retained. |
| Quantity | Value | Interpretation |
| Claims / papers / domains | 120 / 40 / 6 | Public corpus outside the internal SCOPE ledger. |
| Annotators | 4 blinded non-author experts | ML4CO/routing, evaluation/benchmarking, OR/optimization, and ML systems/reproducibility perspectives. |
| Mean pairwise exact agreement | 0.761 | Experts often agree exactly on nontrivial claim-scope labels. |
| Mean pairwise within-one-level agreement | 0.931 | Most disagreements occur at adjacent scope levels. |
| Weighted Cohen’s (mean pairwise) | 0.703 | Substantial ordinal agreement. |
| Krippendorff’s (ordinal) | 0.686 | Conservative multi-annotator reliability estimate. |
| Majority group | C0/bridge | C1–C2 | C3–C4 | C5–C6 | Off/insufficient |
| C0/bridge | 25 | 2 | 0 | 0 | 1 |
| C1–C2 | 3 | 21 | 2 | 0 | 1 |
| C3–C4 | 1 | 2 | 32 | 2 | 0 |
| C5–C6 | 0 | 0 | 3 | 15 | 2 |
| Off/insufficient | 0 | 0 | 0 | 0 | 8 |
| Disagreement type | Count | Typical pattern |
| Ambiguous deployment wording | 6 | The paper uses realistic or practical language without C6-scale evidence. |
| Source-scope strictness | 5 | Humans credit bridge or native-task evidence one level higher than the source-scope rule allows. |
| Artifact availability | 4 | Humans infer reproducibility from prose, while SCOPE-Routing requires manifest-linked objects. |
| Preference coverage boundary | 2 | Scalar preference results are strong but coverage evidence is partial. |
| Latency/online-effect boundary | 2 | Dynamic or online language appears, but route-change or latency telemetry is incomplete. |
| Public source | Native/source-supported credit | SCOPE-transfer scope | Boundary reason |
| PyEPO ( Tang and Khalil, 2022 ) | Predict-then-optimize library evidence across shortest path, knapsack, TSP, and LP/IP tasks. | Bridge/native decision-quality evidence only. | Does not instantiate the SCOPE dynamic multigraph route object or its telemetry contract. |
| DataSP ( Lahoud et al., 2024 ) | Contextual cost learning and differentiable all-to-all shortest-path/path-prediction evidence. | Shortest-path bridge evidence with coverage caveat. | Parallel-edge identity, SCOPE fallback taxonomy, preference-shift coverage, and dynamic route-change telemetry are absent. |
| FrontierCO ( Feng et al., 2026 ) | Native large-scale CO solver benchmark/readiness evidence; local TSP diagnostics are executable. | Native diagnostic/readiness, not same-contract SCOPE routing. | No SCOPE source-destination route-producing policy and no SCOPE route telemetry. |
| CO-Bench ( Sun et al., 2026 ) | Native LLM-agent algorithm-search benchmark; RCSP native solver-interface audit solves or certifies local source-scope review instances and derived records. | Solver-interface/source-scope evidence; no learned multigraph-routing escalation. | Evaluated object is algorithm search or native RCSP solving, not a SCOPE route-producing policy. |
| PredOpt-SP | Reproduced shortest-path decision-focused bridge comparators. | C0 shortest-path bridge audit. | Not strict same-walltime evidence and not dynamic preference-conditioned multigraph routing. |
| PredOpt-Knapsack | Native constrained-optimization closure over local selector and official comparator rows. | Off-thesis positive evidence, not routing evidence. | Knapsack is not a source-destination route contract. |
| Evidence block | Coverage | Claim boundary |
| PredOpt-SP | Latest local official-code extension plus preference-function bridge rows; reproduced SPO+ bridge remains stronger than the local selector row on the audited bridge slice. | Shortest-path bridge audit; blocks all-baseline or state-of-the-art routing claims. |
| PredOpt-Knapsack | Local selector rows and official comparator rows across method files; the selector-family mean is better than the best comparator-family mean on the native constrained-optimization audit. | Full native PredOpt knapsack closure, but out-of-contract for routing. |
| DataSP | Three-seed SCOPE bridge over fixed synthetic episodes plus SPO+ bridge baseline. | C0 shortest-path bridge with coverage caveat, not native full DataSP routing validation. |
| FrontierCO TSP | Native TSP/LKH diagnostics and seed portfolio; no SCOPE-native tour-producing policy. | Native diagnostic/readiness evidence only under unmatched task contract. |
| CO-Bench RCSP bridge | Subset of raw RCSP instances converted through a single-resource SCOPE bridge; the subset is saturated by static/exact baselines. | RCSP bridge coverage audit; no learned-routing win claimed. |
| CO-Bench RCSP native audit | All local source-scope review instances are audited; each instance is solved optimally or certified infeasible under the native solver-interface check, without redistributing upstream raw assets whose reuse terms are unclear. | Native solver-interface/source-scope evidence; C2–C6 SCOPE routing promotion remains blocked. |
| Claim family | Regret-only interpretation | Status | Deciding evidence and supported scope |
| Aggressive pruning | Top- appears attractive because it reaches regret 0.009552. | unsupported | Fallback rises to 40.94% and hypervolume falls to 227.985; no positive pruning or operational routing claim. |
| Simple-graph collapse | Feasible-only regret could suggest that collapse is acceptable. | unsupported | Feasibility falls to 90.6% on the controlled slice and 50.0% on the diagnostic slice; C2+ requires edge identity. |
| Realistic diagnostic learned win | Equality with a strong static anchor could be read as realistic stability. | unsupported | Equality is mediated by 100% fallback and route-static behavior; supported claim is diagnostic failure analysis. |
| DataSP bridge | Scalar regret improves from 0.21546 to 0.03352. | narrowed | Hypervolume decreases from 0.32903 to 0.30314 and the source is not multigraph; supported claim is learned-scoring bridge evidence with a coverage caveat. |
| FrontierCO diagnostics | External compatibility could be read as native routing success. | unsupported | The requested routing branch is empty; retains compatibility and solver-interface efficiency diagnostics. |
| Public Scope Firewall | Real public benchmark evidence could be overread as same-contract routing validation. | narrowed | Public rows retain native/source-supported credit while SCOPE-transfer is blocked when the dynamic multigraph route contract is absent. |
| Row | Regret | vs static | Feas. | Fallback | Adjudicated claim |
| Static exact anchor | 0.014154 | – | 100% | 0.000000 | Static exact comparator. |
| Learned full-exact scorer | 0.009878 | -30.2% | 100% | 0.000000 | supported to C0–C4: feasibility-audited controlled learned scoring. |
| Top- retained decode | 0.009931 | -29.9% | 100% | 0.000000 | narrowed to top- -specific retained-decode evidence. |
| Collapsed exact control | n/a | – | 90.6% | 0.000000 | unsupported for C2+ multigraph claims. |
| Evidence block | Highest licensed level | Missing C6 evidence | Interpretation |
| Controlled internal audit | C4 | Multi-domain realistic support, dynamic robustness, non-author usability. | Valid controlled learned-scoring evidence only. |
| RouteScope-Lite semisynthetic suite | C5 | Deployed topology, live workload, non-author reproduction, intended-use validation. | Strong same-contract stress evidence, not deployment. |
| Diagnostic realistic slice | diagnostic only | Sample size, fallback-free learned win, multi-domain realism. | Mechanism evidence for fallback/static failures. |
| External public rows | Scoped public credit | Same routing contract, telemetry completeness, route identity. | External scope firewall evidence. |
| Public Scope Firewall | Scoped source credit | Dynamic preference-conditioned multigraph routing contract and telemetry. | Real public sources receive native, bridge, solver-interface, or out-of-contract credit, not deployment or multigraph-routing escalation. |
| Artifact dry run | Record closure | Live external use, public non-author adapters, and deployment evidence. | Reviewer-runnable analysis, not community adoption. |
| Gate removed or stressed | Overcredit events | False-support mode blocked by full rule |
| Fallback | 4 | Fallback-heavy pruning, fallback-heavy same-contract winners, or diagnostic equality would be credited as learned routing. |
| Multigraph identity / feasibility | 3 | Simple-graph collapse would be credited as valid multigraph routing. |
| Preference coverage | 3 | Scalar regret would be credited as preference-conditioned coverage despite HV loss. |
| Route change | 3 | Static-equivalent or no-effect online rows would be credited as dynamic behavior. |
| Latency | 2 | Operational claims would ignore slower or impractical paths. |
| Source scope | 5 | Bridge, native, or out-of-contract evidence would be silently upgraded to full routing. |
| Policy | Missing-evidence treatment | Boundary results expected to remain unchanged |
| Conservative SCOPE-Routing | Any inherited missing gate blocks support at the affected level; lower supported levels may remain. | Main paper rule. |
| Intermediate SCOPE-Routing | Record, source-scope, feasibility, fallback-attribution, and multigraph-identity requirements remain hard; latency and partial coverage gaps usually narrow. | Rescue-dominated rows still cannot become standalone learned-routing evidence; semisynthetic rows still cannot become C6. |
| Permissive SCOPE-Routing | Missing non-core telemetry narrows by one or more levels when source scope is compatible and no hard requirement fails. | Bridge/out-of-contract evidence still cannot become full multigraph routing; unlocked records still cannot support adjudicated claims. |
| Evidence object | Core snapshot | Extension snapshot | Execution mode |
| Adjudicator, schema, Claim Cards, registry | yes | unchanged/core | Fresh validation and table/card generation. |
| ClaimScope-120 packets and four-expert labels | yes | extended reporting | Fresh reanalysis from released labels and tie logs. |
| RouteScope-Lite request and policy rows | yes | unchanged/core | Cached reanalysis from locked rows; not model retraining. |
| Public source-scope audits | yes | extended summaries | Fresh local checks or derived source-scoped summaries as documented. |
| 40-claim reviewer comparison | no | added | Fresh analysis from anonymized raw judgments, assignments, and times. |
| 24-object construction audit | no | added | Fresh analysis from independent encodings and resolution logs. |
| Asset | Asset type | Source-provided status | Restricted use in this submission |
| PyEPO | Code/library | MIT License | Credited to the original repository and paper; used for source-scope and bridge/native decision-quality comparisons. Upstream notices are preserved when the package is installed or referenced. |
| PredOpt-family benchmark assets | Code/benchmark repository | MIT License for the public benchmark repository checked in the audit | Used for shortest-path and knapsack source-scope comparisons; upstream notices are preserved and no stronger SCOPE-routing claim is inferred from native PredOpt evidence. |
| FrontierCO dataset card | Dataset/benchmark card | Dataset-card license marker recorded in the manifest | Used as source-scoped external benchmark evidence; no raw dataset is relicensed by this submission and no deployment-style or same-contract SCOPE routing claim is inferred. |
| FrontierCO code repository | Code repository | No clear reusable software license identified in the repository front matter during the audit | Not vendored, mirrored, or required by the default reviewer path; only paper/repository citation and independently authored diagnostic wrappers or aggregate audit summaries are included. |
| CO-Bench dataset card | Dataset card | License marked unknown on the dataset card during the audit | Treated conservatively as source-bound; raw upstream data are not redistributed, automatically downloaded, or relicensed. The artifact includes only source citations, Claim Cards, and derived aggregate source-scope records. |
| CO-Bench code repository | Code repository | No clear reusable software license identified in repository front matter during the audit | Not vendored, mirrored, or required by the default reviewer path; solver-interface checks use independently authored review code and derived summaries rather than copied upstream code. |
| Check | Snapshot / mode | Output artifact |
| Manifest and snapshot validation | Both / fresh | Hash and chronology report. |
| Registry and main-table generation | Reviewed / fresh | Registry, exclusions, tables, and CSV mirrors. |
| Claim Cards, schema, adapters | Reviewed / fresh | Cards plus schema and 15-adapter compliance reports. |
| RouteScope-Lite analysis | Reviewed / cached reanalysis | 22,500 locked policy-row records, full 15-row table, uncertainty and selection reports. |
| ClaimScope-120 analysis | Reviewed / fresh | 120 cards, raw-label matrix, tie policies, clustered intervals, and disagreement taxonomy. |
| Reviewer comparison | Extension / fresh | 240 judgments, 36 resolved claim summaries, timing and clustered analyses. |