Period ending 2026-09-21
7 new papers
A weekly snapshot of new work published in Reproducibility.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this topic, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Reproducibility.
Period ending 2026-09-14
A weekly snapshot of new work published in Reproducibility.
Period ending 2026-09-07
A weekly snapshot of new work published in Reproducibility.
229 papers
unknown'' when the reasoner-entailed answer is no'' under \emph{FunctionalProperty} closure or class \emph{disjointness}. Using 180 reasoner-audited queries from a procedural expansion of the observed pattern plus 18 hand-authored held-out queries in two unrelated domains (insurance and clinical), we compare four interaction modes under matched query budget: single-shot, three rounds of generic ``you-are-wrong'' retry, three rounds of reasoner-verdict repair with an open-world-assumption (OWA) hint, and the same repair without the hint. Direct faithfulness is 43.9,% (Wilson 95,% CI ); generic retry reaches 81.7,% (); the verdict-with-hint variant is \emph{worse} at 67.2,% (); the verdict-only variant reaches 97.8,% (). All pairwise comparisons remain significant under McNemar's exact test with Bonferroni correction (; all ). The same fingerprint accounts for 4/4 errors on the held-out queries. Our interpretation is bounded: prompt framing can matter more than corrective content, and reasoner-guided wrappers should be ablated explicitly.