Period ending 2026-09-21
3 new papers
A weekly snapshot of new work published in Inference-Time.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this topic, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Inference-Time.
Period ending 2026-09-14
A weekly snapshot of new work published in Inference-Time.
Period ending 2026-09-07
A weekly snapshot of new work published in Inference-Time.
197 papers
evil'' and impolite'' scorers grade more harshly, while good'' and optimistic'' scorers grade more leniently. ELA tasks are 2.5-3 more susceptible to scorer personalization than science tasks, and the mixture-of-experts model shows roughly 6 larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring. Our findings highlight the need for task- and architecture-aware calibration when deploying personalized models in educational settings.