KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration
Organizations: The University of Tokyo
Abstract
KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.
Figures & tables
| Test part | Content/person gate | Prompt | Tips |
|---|---|---|---|
| US, even headlines | Pass | (fail) | (fail) |
| 15 countries, odd | Pass (15/15) | (fail) | (fail) |
| 15 countries, even | Pass (15/15) | (fail) | (fail) |
| Matched-subset comparison | Effect [95% CI] | Effect [95% CI] | |
| Jev | [ , 0.011] | 0.015 [ , 0.053] | |
| Astra | 0.052 [0.010, 0.094] | 0.045 [0.002, 0.088] |
| System | Raw effect MAE [95% CI] | Calibrated MAE | Model-only RMSE |
|---|---|---|---|
| Kernel | 0.0305 [0.0242, 0.0412] | 0.0376 | 0.033 |
| Hybrid | 0.0180 [0.0132, 0.0289] | 0.0183 | 0.014 |
| Flagship direct audit | 0.0200 [0.0153, 0.0314] | 0.0326 | 0.020 |
| Engine | Agents | Seconds for 20 steps | Agent-steps/s |
|---|---|---|---|
| Vectorized | 0.007 | ||
| Vectorized | 0.075 | ||
| Vectorized | 0.90 | ||
| Scalar | 0.95 |
| Property | Evidence and permitted interpretation |
|---|---|
| Content ranking | Kernel passes Arechar test gates; screening within the evaluated domain is supported. |
| Known intervention direction | Sparse flagship calibration improves Epstein effect error and SocSci210 decision gain; retain endpoint and domain qualifications. |
| Human effect intervals | D2 improves retrospective SocSci210 coverage; transport to the hybrid or another domain is not established. |
| Cost–quality comparison | Measured calls, tokens, table execution, and declared human metrics support comparison on these workloads. |
| Individual/country effects | Individual prediction and cross-country susceptibility comparisons are unsupported. |
| Repeated exposure/network | Conditional scenarios only; cascade sizes and endogenous network responses are unvalidated. |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Record | Frozen source | Main requirement and status |
|---|---|---|
| G1 | frozen.yaml | p3/choice selected on development, commit 389b984 ; marginal and condition gates pass. Published-model comparisons require metric alignment. |
| G-A | decision_value.yaml | Sign: permutation and above null p95; gain: and bootstrap interval above zero. Both pass; stronger two-thirds-of-pilot criteria do not. Earlier test sensitivity was known. |
| C2/C3 | arechar_split.yaml , arechar_criteria.yaml | Split fixed before download; criteria committed at d826501 after US/odd calibration. Content/person gates pass; kernel intervention gates fail. |
| C3b/A3b | flagship_baseline.yaml | Flagship comparison frozen at ffcb4a7 before predictions; existing criteria reused unchanged. |
| D1 | epstein_criteria.yaml , epstein_screens.yaml | Commit 7b5a5a3 ; policies and prediction hashes locked before scoring. Positive policy-value difference with interval excluding zero required. Not shown. Effect MAE remains secondary. |
| D1 A2 | epstein_headline_criteria.yaml | Commit cdc7864 , after arm aggregates but before headline outcomes were examined; positive within-arm correlation with interval above zero required. Not shown. |
| Wave | Human best | Kernel choice (human value) | Hybrid choice (human value) | Audit choice |
|---|---|---|---|---|
| 1 | Control | Control (0.059) | Control (0.059) | Control |
| 2 | LE | LE (0.128) | LE (0.128) | LE |
| 3 | Tips | GN (0.057) | E (0.089) | GN |
| 4 | TN | TN (0.113) | TN (0.113) | TN |
| 5 | IN | IN (0.101) | I (0.083) | I |
| Wave | Intervention | Human | SE | 95% interval | Kernel | Hybrid | Audit | |
|---|---|---|---|---|---|---|---|---|
| 2 | Evaluation | 0.0175 | 395/387 | 0.023 | 0.008 | |||
| 2 | Long evaluation | 0.0172 | 410/387 | 0.025 | 0.013 | |||
| 3 | Evaluation | 0.0143 | 540/533 | 0.011 | 0.015 | |||
| 3 | Generic norms | 0.0139 | 510/533 | 0.008 | 0.017 | |||
| 3 | Tips | 0.0149 | 498/533 | 0.006 | 0.016 | |||
| 4 | Partisan norms | 0.0128 | 949/487 | 0.024 | 0.027 |
| System | Policy value | Chance mean/p95 | One-sided | Same judge |
| Kernel | 0.0915 | 0.0808/0.0944 | 0.105 | 0.0912 |
| Hybrid | 0.0942 | 0.0809/0.0944 | 0.062 | 0.0938 |
| Direct audit | 0.0878 | 0.0807/0.0942 | 0.197 | 0.0875 |
| Human half-sample | — | — | — | 0.0941 |
| Hybrid minus kernel (primary) | 0.0027 [ , 0.0169] | |||
| Hybrid minus direct audit | 0.0063 [ , 0.0211] | |||
| Personas per wave | Kernel MAE | Hybrid MAE |
|---|---|---|
| 10 | ||
| 20 | ||
| 30 | ||
| 50 | ||
| 100 | ||
| 200 | 0.0305 | 0.0180 |
| System | Within-task | Sign | Gain | Raw MAE | Magnitude ratio | |
|---|---|---|---|---|---|---|
| Kernel | 0.397 | 0.671 | 0.268 | 0.052 | 0.68 | 0.608 |
| 0.489 | 0.671 | 0.336 | 0.078 | 1.64 | 1.000 | |
| 0.461 | 0.697 | 0.339 | 0.086 | 1.87 | 0.927 | |
| (694 items) | 0.433 | 0.717 | 0.382 | 0.086 | 1.85 | 0.701 |
| (1,388 items) | 0.417 | 0.742 | 0.386 | 0.084 | 1.80 | 0.721 |
| (2,082 items) | 0.423 | 0.709 | 0.386 | 0.083 | 1.75 | 0.709 |
| System | Within-task | Sign accuracy | Captured gain | Effect MAE | |
|---|---|---|---|---|---|
| Kernel | [0.061, 0.613] | [0.495, 0.801] | [0.051, 0.446] | [0.041, 0.065] | [0.323, 0.780] |
| [0.188, 0.675] | [0.511, 0.788] | [0.117, 0.519] | [0.064, 0.094] | [1.000, 1.000] | |
| [0.182, 0.639] | [0.556, 0.806] | [0.105, 0.532] | [0.073, 0.101] | [0.884, 0.958] | |
| [0.201, 0.616] | [0.564, 0.833] | [0.170, 0.560] | [0.069, 0.109] | [0.540, 0.822] | |
| [0.174, 0.611] | [0.607, 0.840] | [0.162, 0.573] | [0.065, 0.108] | [0.565, 0.838] | |
| [0.189, 0.607] | [0.548, 0.823] | [0.164, 0.566] | [0.066, 0.103] | [0.550, 0.832] |
| Comparison | Sign accuracy | Captured gain | Within-task |
|---|---|---|---|
| 0.038 [ , 0.122] | 0.119 [0.014, 0.224] | 0.026 [ , 0.213] | |
| 0.071 [ , 0.155] | 0.118 [0.006, 0.224] | 0.020 [ , 0.204] | |
| 0.046 [ , 0.121] | 0.114 [0.014, 0.215] | 0.035 [ , 0.217] | |
| 0.038 [ , 0.092] | 0.050 [ , 0.117] | [ , 0.017] | |
| 0.071 [0.022, 0.128] | 0.049 [ , 0.115] | [ , 0.012] | |
| 0.000 [ , 0.060] | 0.069 [ , 0.168] | 0.092 [ , 0.253] |
| System | Within-task | Sign | Gain | MAE | Magnitude ratio | |
|---|---|---|---|---|---|---|
| Kernel | 0.552 | 0.895 | 0.420 | 0.067 | 0.61 | 0.863 |
| 0.607 | 0.939 | 0.529 | 0.079 | 1.15 | 1.000 | |
| 0.592 | 0.937 | 0.504 | 0.084 | 1.21 | 0.979 | |
| 0.533 | 0.927 | 0.409 | 0.080 | 1.14 | 0.897 | |
| 0.551 | 0.927 | 0.446 | 0.081 | 1.15 | 0.911 | |
| 0.561 | 0.933 | 0.496 | 0.079 | 1.13 | 0.917 |
| Fit | |||
|---|---|---|---|
| Jev, selected seen | 0.925 [0.607, 1.173] | 0.037 [0.021, 0.046] | 0.080 [0.051, 0.128] |
| Astra, selected seen | 0.579 [0.317, 0.660] | 0.035 [0.020, 0.046] | 0.074 [0.047, 0.125] |
| Jev, unselected dev | 0.661 | 0.020 | 0.020 |
| Evaluation | 50% | 80% | 90% | Width (90%) | Score (90%) |
|---|---|---|---|---|---|
| Jev, split half | 0.581 | 0.808 | 0.882 | — | — |
| Astra, split half | 0.599 | 0.839 | 0.902 | — | — |
| Jev, test | 0.669 | 0.932 | 0.959 | 0.324 | 0.422 |
| Astra, test | 0.625 | 0.859 | 0.937 | 0.295 | 0.394 |
| Jev, no discrepancy | 0.140 | 0.288 | 0.362 | 0.064 | 0.746 |
| Astra, no discrepancy | 0.204 | 0.334 | 0.397 | 0.064 | 0.754 |
| Respondents per cell | 5 | 10 | 20 | 50 | 100 | 200 |
|---|---|---|---|---|---|---|
| Sign accuracy | 0.655 | 0.674 | 0.632 | 0.645 | 0.661 | 0.671 |
| Captured gain | 0.272 | 0.358 | 0.264 | 0.255 | 0.278 | 0.268 |
| Cost (USD) | 0.0007 | 0.0015 | 0.0029 | 0.0074 | 0.0147 | 0.0295 |
| Model | Condition | Magnitude ratio | Sign | Gain | |
|---|---|---|---|---|---|
| Jev p3 | 0.1360 | 0.323 | 0.89 | 0.645 | 0.272 |
| GPT-5.6 Luna (low) | 0.1320 | 0.506 | 1.61 | 0.697 | 0.305 |
| GPT-6 Astra (high) | 0.1013 | 0.489 | 1.58 | 0.671 | 0.336 |
| Comparable-block : Jev 0.347; Luna 0.495; Astra 0.489. | |||||
| Astra–Jev difference: 0.142 [0.014, 0.289] (95%). | |||||
| Astra–Jev sign difference: 0.026 [ , 0.083] (95%). | |||||
| Study | Independent magnitude | Corrected | Independent structure | Corrected | Distant/human |
|---|---|---|---|---|---|
| m52pd | 0.175 | 0.289 | 0.018 | 0.518 | 0.274 |
| 3xy9j | 0.548 | 0.645 | 0.005 | 0.263 | 0.611 |
| gx6hp | 0.456 | 0.786 | 0.001 | 0.514 | 0.791 |
| ef6my | 0.305 | 0.345 | 0.089 | 0.410 | 0.326 |
| 5mt6r | 0.338 | 0.523 | 0.132 | 0.711 | 0.471 |
| ac9jm | 0.122 | 0.281 | 0.718 | 0.258 |