HygieneRoboBench: Benchmarking Hygiene-Aware Planning for Household Robots
Authors: Yurun Chen, Josh Qixuan Sun, Jason Qin, Chengtai Li, Tianyi Wang, Mark Crowley, Wentao Zhu
Organizations: Shanghai Jiao Tong University · Eastern Institute of Technology, Ningbo · University of Waterloo · Stony Brook University · University of Science and Technology of China · Ningbo Institute of Digital Twin · Chengdu Institute of Computer Applications, Chinese Academy of Sciences
Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not jointly assess how planners identify hygiene risks from contact history and plan safe continuations after new contact events. Planners must do so within time and resource limits while respecting user priorities. We introduce HygieneRoboBench, with 624 instances across 134 task families, to evaluate safe resolution of household tasks from a given execution history. Tasks capture contamination through two grippers and shared objects, treatment costs, and user priorities. We combine controlled history, profile, and event comparisons with independent plan evaluation. These assess safe resolution, cost efficiency under user priorities, and responses to contact events. Evaluation of LLM-based and symbolic planners shows that safely completing a task does not guarantee the lowest execution costs under the user's priorities. To address this problem, we introduce Hygiene-NSP. It combines LLM-based grounding, contact-history reconstruction, and CP-SAT to jointly plan hygiene treatment and task execution under user priorities. Hygiene-NSP achieves safe resolution and optimal safe resolution rates of 94.4% and 90.4%, respectively. Both rates are higher than those of the evaluated baseline planners on the full dataset. Project page: https://euron-zc.github.io/HygieneRoboBench/.
Figures & tables
Work
Safety checks
Contact hygiene
History input
Plan updates
Task limits
User preferences
SafeAgentBench [ 9 ]
✓
×
✓
✓
T
×
IS-Bench [ 10 ]
✓
✓
✓
✓
T
×
VestaBench [ 11 ]
—
×
✓
✓
T/R
×
SafeManip [ 12 ]
✓
✓
×
×
T/R
×
SIMMER [ 6 ]
✓
✓
×
×
T/R
×
PARTNR [ 13 ]
×
×
✓
✓
T/R
×
TABLE I: Household planning capabilities covered by related benchmarks and studies.
Fig. 2: Dataset construction. (a) Household activities and hygiene guidance supply construction inputs. (b) Five shared registries define task, hygiene, event, resource, and user-priority specifications. (c) Episode Builder combines a base task with contact history, user priorities, event information, and applicable shared definitions to construct controlled instances within a task family. The Verifier groups matched-condition checks, Reference validation, independent replay, and human review for dataset admission.
Fig. 3: Dataset composition. The 134 task families comprise 624 instances. From outer to inner, the rings show activity groups, construction designs, and instance difficulty. Each ring independently partitions the same 624 instances, with labels indicating instance counts.
Overall ( N=624 )
Easy ( n=200 )
Moderate ( n=238 )
Hard ( n=186 )
Planner
SR
OSR
SR
OSR
SR
OSR
SR
OSR
GPT-5.6-Sol
78.5
76.9
87.5
85.5
78.6
77.7
68.8
66.7
GPT-5.6-Luna
46.6
37.0
55.0
45.5
52.1
41.2
30.6
22.6
Gemini 3.6 Flash
89.7
86.1
95.0
91.5
90.8
88.7
82.8
76.9
DeepSeek-V4.1-Flash
90.7
75.0
96.0
83.5
92.4
71.8
82.8
69.9
MiniMax-M3
67.0
57.2
74.5
66.0
70.2
58.8
54.8
45.7
TABLE II: Overall performance and breakdowns by instance difficulty and activity.
Other no plan
Incorrect rejection
Invalid continuation
Safe, nonoptimal
GPT-5.6-Sol
1
25
108
10
GPT-5.6-Luna
27
112
194
60
Gemini 3.6 Flash
8
12
44
23
DeepSeek-V4.1-Flash
3
7
48
98
MiniMax-M3
33
10
163
61
Rule-SAT
415
12
4
6
TABLE III: Failure and cost outcomes across 624 instances per planner. Counts are mutually exclusive. The first three columns fail SR. The last passes SR but not OSR.
Planner
Relevant ( n=180 )
Irrelevant ( n=92 )
GPT-5.6-Sol
126 / 123
69 / 66
GPT-5.6-Luna
57 / 38
27 / 16
Gemini 3.6 Flash
157 / 141
81 / 76
DeepSeek-V4.1-Flash
162 / 118
86 / 60
MiniMax-M3
94 / 65
51 / 33
Rule-SAT
74 / 71
50 / 50
TABLE IV: Controlled history and priority pairs. Entries give HC/HC ∗ or PS counts; n counts eligible pairs from the construction groups in Fig. 3 . Evaluation sets may share instances.
While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations where robots are expected to choose actions that prioritize diverse values such as human autonomy, efficiency, or social appropriateness. Yet, there are no benchmarks for evaluating robots' value preferences in such scenarios. We introduce RobotValues, a benchmark to evaluate household robot planners in 8K value-conflict scenarios. Each instance consists of a realistic, synthetically generated household image with multiple plausible robot actions that prioritize different human values. We construct ROBOTVALUES through LLM-assisted scenario generation, stakeholder-grounded value extraction, image generation and automatic quality control. We evaluate 10 VLMs used in robotics and find that models show default value preferences, including safety and accommodation, while underselecting privacy-prioritizing actions. When models are prompted to prioritize values that conflict with their preferences, models often fail to override the default actions, choosing incorrect actions 80% of the time on average across models. These findings highlight the need to go beyond task completion or safety evaluations and assess robots' decision-making capability when human values conflict.
Jongwook Han, Hyeongjin Kim, Yohan Jo
Graduate School of Data Science, Seoul National University
Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introduce DESPITE, a benchmark of 12,279 tasks spanning physical and normative dangers with fully deterministic validation. Across 23 models, even near-perfect planning ability does not ensure safety: the best-planning model fails to produce a valid plan on only 0.4% of tasks but produces dangerous plans on 28.3%. Among 18 open-source models from 3B to 671B parameters, planning ability improves substantially with scale (0.4-99.3%) while safety awareness remains relatively flat (38-57%). We identify a multiplicative relationship between these two capacities, showing that larger models complete more tasks safely primarily through improved planning, not through better danger avoidance. Three proprietary reasoning models reach notably higher safety awareness (71-81%), while non-reasoning proprietary models and open-source reasoning models remain below 57%. As planning ability approaches saturation for frontier models, improving safety awareness becomes a central challenge for deploying language-model planners in robotic systems.
Tao Zhang, Kaixian Qu, Zhibin Li +4
ETH Zurich, Zurich, Switzerland. · University College London, London, United Kingdom. · Stanford University, Stanford, California, United States. +2
Large Language Models (LLMs) can reason over complex instructions but often fail to satisfy the physical and spatial constraints required for robotic task planning. Recent LLM-based planners directly translate text into action sequences, yet they lack structured reasoning about feasibility, reachability, and logical order, resulting in invalid or incomplete plans. We present a heterogeneous multi-LLM framework that decomposes instructions into atomic reasoning tasks and allocates them to role-specialized expert agents under a token budget for real-world computational and communicational constraints. By combining role-oriented reasoning from heterogeneous agents followed by constraint-driven plan synthesis, HEART validates capability, reachability, and constraint conditions before planning and helps produce physically executable plans while maintaining efficiency. Experiments across different household benchmarks show that HEART consistently improves plan success compared to single-LLM and rule-based planners, demonstrating that heterogeneous LLM collaboration enables robust and scalable robotic task planning under resource constraints.
Junho Lee, Seabin Lee, Wonjong Lee +3
Dept. of Electronic Engineering, Sogang University, Seoul, Korea