cs.AIOct 2, 2026

Benchmarking candidate coverage and rejection policy transfer in typed decision models

Authors: Jiawen Lu, Tongtong Wu

Organizations: Monash University

Abstract

Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference labels, testing Laya/Jev policy transfer at equal calibration budgets and all three models on artificial omission, natural retrieval misses and public out-of-scope queries. Source calibration often fails to preserve the target operating point. A Jev policy calibrated on DBpedia rejects 69.3% of covered Emotion test inputs, while an Emotion policy loses detection entirely. Retrieval exposes a different tradeoff: with ten intent candidates, Laya detects 99.0% of out-of-scope queries but rejects 48.8% of covered queries. Separating missing-answer sources reveals these costs alongside retrieval coverage. The benchmark provides shared inputs, explicit decision and failure categories, and reproducible scoring to assess rejection policies under the conditions in which they are reused. Code and benchmark artifacts are available at https://github.com/luckykevvv/Decision_Model_Benchmark.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev

    Sep 30, 2026Jike Zhong, Ming Li, Yuxiang LaiNumerical Reasoning in Language ModelsSelective Prediction

  2. Benchmarking System One decision models against trained classifiers and language models for automated decision gates

    Sep 29, 2026Amir Rafe, Subasish DasLLM Decision-MakingLLM Evaluation