HygieneRoboBench: Benchmarking Hygiene-Aware Planning for Household Robots
Authors: Yurun Chen, Josh Qixuan Sun, Jason Qin, Chengtai Li, Tianyi Wang, Mark Crowley, Wentao Zhu
Organizations: Shanghai Jiao Tong University · Eastern Institute of Technology, Ningbo · University of Waterloo · Stony Brook University · University of Science and Technology of China · Ningbo Institute of Digital Twin · Chengdu Institute of Computer Applications, Chinese Academy of Sciences
Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not jointly assess how planners identify hygiene risks from contact history and plan safe continuations after new contact events. Planners must do so within time and resource limits while respecting user priorities. We introduce HygieneRoboBench, with 624 instances across 134 task families, to evaluate safe resolution of household tasks from a given execution history. Tasks capture contamination through two grippers and shared objects, treatment costs, and user priorities. We combine controlled history, profile, and event comparisons with independent plan evaluation. These assess safe resolution, cost efficiency under user priorities, and responses to contact events. Evaluation of LLM-based and symbolic planners shows that safely completing a task does not guarantee the lowest execution costs under the user's priorities. To address this problem, we introduce Hygiene-NSP. It combines LLM-based grounding, contact-history reconstruction, and CP-SAT to jointly plan hygiene treatment and task execution under user priorities. Hygiene-NSP achieves safe resolution and optimal safe resolution rates of 94.4% and 90.4%, respectively. Both rates are higher than those of the evaluated baseline planners on the full dataset. Project page: https://euron-zc.github.io/HygieneRoboBench/.
Figures & tables
Work
Safety checks
Contact hygiene
History input
Plan updates
Task limits
User preferences
SafeAgentBench [ 9 ]
✓
×
✓
✓
T
×
IS-Bench [ 10 ]
✓
✓
✓
✓
T
×
VestaBench [ 11 ]
—
×
✓
✓
T/R
×
SafeManip [ 12 ]
✓
✓
×
×
T/R
×
SIMMER [ 6 ]
✓
✓
×
×
T/R
×
PARTNR [ 13 ]
×
×
✓
✓
T/R
×
TABLE I: Household planning capabilities covered by related benchmarks and studies.
Fig. 2: Dataset construction. (a) Household activities and hygiene guidance supply construction inputs. (b) Five shared registries define task, hygiene, event, resource, and user-priority specifications. (c) Episode Builder combines a base task with contact history, user priorities, event information, and applicable shared definitions to construct controlled instances within a task family. The Verifier groups matched-condition checks, Reference validation, independent replay, and human review for dataset admission.
Fig. 3: Dataset composition. The 134 task families comprise 624 instances. From outer to inner, the rings show activity groups, construction designs, and instance difficulty. Each ring independently partitions the same 624 instances, with labels indicating instance counts.
Overall ( N=624 )
Easy ( n=200 )
Moderate ( n=238 )
Hard ( n=186 )
Planner
SR
OSR
SR
OSR
SR
OSR
SR
OSR
GPT-5.6-Sol
78.5
76.9
87.5
85.5
78.6
77.7
68.8
66.7
GPT-5.6-Luna
46.6
37.0
55.0
45.5
52.1
41.2
30.6
22.6
Gemini 3.6 Flash
89.7
86.1
95.0
91.5
90.8
88.7
82.8
76.9
DeepSeek-V4.1-Flash
90.7
75.0
96.0
83.5
92.4
71.8
82.8
69.9
MiniMax-M3
67.0
57.2
74.5
66.0
70.2
58.8
54.8
45.7
TABLE II: Overall performance and breakdowns by instance difficulty and activity.
Other no plan
Incorrect rejection
Invalid continuation
Safe, nonoptimal
GPT-5.6-Sol
1
25
108
10
GPT-5.6-Luna
27
112
194
60
Gemini 3.6 Flash
8
12
44
23
DeepSeek-V4.1-Flash
3
7
48
98
MiniMax-M3
33
10
163
61
Rule-SAT
415
12
4
6
TABLE III: Failure and cost outcomes across 624 instances per planner. Counts are mutually exclusive. The first three columns fail SR. The last passes SR but not OSR.
Planner
Relevant ( n=180 )
Irrelevant ( n=92 )
GPT-5.6-Sol
126 / 123
69 / 66
GPT-5.6-Luna
57 / 38
27 / 16
Gemini 3.6 Flash
157 / 141
81 / 76
DeepSeek-V4.1-Flash
162 / 118
86 / 60
MiniMax-M3
94 / 65
51 / 33
Rule-SAT
74 / 71
50 / 50
TABLE IV: Controlled history and priority pairs. Entries give HC/HC ∗ or PS counts; n counts eligible pairs from the construction groups in Fig. 3 . Evaluation sets may share instances.