Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation
Organizations: The Chinese University of Hong Kong · Knowin AI · The Hong Kong University of Science and Technology (Guangzhou)
Abstract
Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.
Figures & tables
| Object | Goal | Spatial | Average | ||||
| Method | Position | Task | Position | Task | Position | Task | |
| VLA Methods | |||||||
| OpenVLA | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 17.0 | 1.0 | 38.0 | 0.0 | 20.0 | 1.0 | 12.8 | |
| MolmoAct | 6.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 1.0 |
| References ( Zhang et al., 2026a ) | K1+ Gemini (ours) | ||||
| Task/Methods or Arms | CaP-Agent0 | + RATs skills | Panda | UR5e | IIWA |
| Cube lifting | 68.0 | 84.0 | 100.0 | 100.0 | 100.0 |
| Cube restacking | 34.0 | 46.0 | 100.0 | 100.0 | 100.0 |
| Cube stacking | 46.0 | 60.0 | 100.0 | 100.0 | 100.0 |
| Nut assembly | 0.0 | 0.0 | 60.0 | 55.0 | 45.0 |
| Spill wiping | 100.0 | 100.0 | 65.0 | – | – |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool | Contract and boundary |
|---|---|
| find_regions | Text-prompted SAM 3 masks from a current original RGB image, with RGB-D summaries and stable S identifiers. Multiple candidates require model-side disambiguation. |
| inspect_region | Box-prompted segmentation or explicit refresh of a region. The supplied label does not guide semantic segmentation. Returns visible geometry, not a full-object pose. |
| measure_depth | Back-projects a selected valid visible pixel to world coordinates, with TCP-relative displacement. Creates a trackable point reference, not a grasp target. |
| fit_geometry | Fits local visible geometry in an image region. A fitted plane or unsigned principal direction is not an articulation model. |
| region_relation | Measures region-to-robot or region-to-region displacement and available plane relations. |
| check_region_view | Projects an existing measured region into the other available camera, with a depth consistency check. This is not a newly rendered viewpoint. |
| Arm | Trainable components | Learning rate | Examples/epoch |
|---|---|---|---|
| Qwen K1 RUA | Language LoRA , , dropout .05 | 2,482 | |
| Qwen RUA | Same language LoRA, native text output head | 1,670 | |
| Qwen VLA (RGB-only / RGB-D) | Same LoRA plus numeric-state/action head, optional depth encoder | 10,241 | |
| Language LoRA , , plus full action expert/projections | 10,241 | ||
| OpenVLA | All-linear LoRA , , dropout .05 | 40,808 |
| Property | Original RGB interface | K1 |
|---|---|---|
| Scene information | Two RGB views | Same two views plus depth-backed measurements |
| Text history | Original complete history | Last eight exchanges plus retrieval |
| Image history | Current and previous decision | Current and previous decision, optional retrieved images |
| Motion | Original interpolated pose tool | Local corrections, persistent poses, full-target transit |
| Perception | Model interprets RGB | Explicit region, depth, tracking, grasp tools |
| Progress memory | Dialogue | Dialogue plus model-maintained milestones |
| Suite | Trained condition indices (A/B) | Held-out (C) |
|---|---|---|
| goal_swap | 0, 1, 4, 7, 8, 9 | 2, 6 |
| goal_task | 3, 4, 5, 7, 8, 9 | 2, 6 |
| object_swap | 2, 3, 4, 5, 6, 7, 8, 9 | 0, 1 |
| object_task | 2, 3, 4, 5, 6, 7, 8, 9 | 0, 1 |
| spatial_swap | 1, 2, 5, 6, 7, 8, 9 | 0, 3 |
| spatial_task | 1, 2, 3, 4, 6, 7, 8, 9 | 0, 5 |
| System | Spatial | Object | Goal | Average |
|---|---|---|---|---|
| RGB + GPT-6 Astra | 33.3 | 83.3 | 66.7 | 61.1 |
| K1 + Gemini 3.7 Flash | 66.7 | 100.0 | 66.7 | 77.8 |
| K1 + GPT-6 Astra | 83.3 | 100.0 | 83.3 | 88.9 |
| Model | Epoch | A | B | C |
|---|---|---|---|---|
| Qwen K1 RUA | 1 | 7/43 (16.3%) | 10/43 (23.3%) | 3/36 (8.3%) |
| Qwen K1 RUA | 3 | 14/43 (32.6%) | 20/43 (46.5%) | 5/36 (13.9%) |
| Qwen K1 RUA | 5 | 22/43 (51.2%) | 19/43 (44.2%) | 5/36 (13.9%) |
| Qwen RUA | 1 | 0/43 (0.0%) | 0/43 (0.0%) | 0/36 (0.0%) |
| Qwen RUA | 3 | 0/43 (0.0%) | 0/43 (0.0%) | 0/36 (0.0%) |
| Qwen RUA | 5 | 3/43 (7.0%) | 2/43 (4.7%) | 0/36 (0.0%) |
| Model | Group | Spatial | Object | Goal |
|---|---|---|---|---|
| Qwen K1 RUA | A | 6/15 (40.0%) | 11/16 (68.8%) | 5/12 (41.7%) |
| Qwen K1 RUA | B | 7/15 (46.7%) | 7/16 (43.8%) | 5/12 (41.7%) |
| Qwen K1 RUA | C | 3/12 (25.0%) | 1/12 (8.3%) | 1/12 (8.3%) |
| Qwen RUA | A | 0/15 (0.0%) | 1/16 (6.2%) | 2/12 (16.7%) |
| Qwen RUA | B | 0/15 (0.0%) | 1/16 (6.2%) | 1/12 (8.3%) |
| Qwen RUA | C | 0/12 (0.0%) | 0/12 (0.0%) | 0/12 (0.0%) |
| Method | Robot | Task | Successes / trials | Accuracy (%) |
|---|---|---|---|---|
| K1 + Gemini | Panda | Cube lifting | 20/20 | 100.0 |
| K1 + Gemini | Panda | Cube restacking | 20/20 | 100.0 |
| K1 + Gemini | Panda | Cube stacking | 20/20 | 100.0 |
| K1 + Gemini | Panda | Nut assembly | 12/20 | 60.0 |
| K1 + Gemini | Panda | Spill wiping | 13/20 | 65.0 |
| K1 + Gemini | Panda | Two-arm handover | 1/20 | 5.0 |
| Task | Easy | Hard | Excluded | Action limit |
| Grab Roller | 8/10 (80.0%) | 9/10 (90.0%) | 0/0 | 400 |
| Handover Mic | 2/10 (20.0%) | 0/10 (0.0%) | 0/0 | 600 |
| Lift Pot | 2/10 (20.0%) | 2/10 (20.0%) | 0/0 | 400 |
| Move Can Pot | 4/8 (50.0%) | 5/8 (62.5%) | 2/2 | 400 |
| Place Dual Shoes | 0/8 (0.0%) | 0/6 (0.0%) | 2/4 | 600 |
| Place Phone Stand | 2/10 (20.0%) | 0/8 (0.0%) | 0/2 | 400 |
| Task | ACT | DP | DP3 | RDT | |
|---|---|---|---|---|---|
| Grab Roller | 66/6 | 98 /1 | 77/1 | 74/43 | 96/ 80 |
| Handover Mic | 9/0 | 53/0 | 93/7 | 98/ 41 | 100 /13 |
| Lift Pot | 7/2 | 37/0 | 85 /0 | 82/19 | 84/ 36 |
| Move Can Pot | 0/0 | 42/0 | 28/0 | 47/23 | 74 / 32 |
| Place Dual Shoes | 0/0 | 7/0 | 1/0 | 4/ 4 | 15 /0 |
| Place Phone Stand | 0/0 | 17/0 | 8/1 | 15/6 | 35 / 7 |
| Category | Tool | Requests |
| Motion | move_toward | 1,835 |
| move_relative | 654 | |
| move_to_pose | 547 | |
| rotate_toward | 271 | |
| align_axis | 48 | |
| set_gripper | 883 |
| Successes | Accuracy (%) | Mean calls | ||
|---|---|---|---|---|
| 0 | 8 | 15/18 | 83.3 | 28.3 |
| 1 | 8 | 14/18 | 77.8 | 33.0 |
| 2 | 8 | 14/18 | 77.8 | 41.2 |
| 3 | 8 | 11/18 | 61.1 | 47.4 |
| 4 | 8 | 12/18 | 66.7 | 45.9 |
| 1 | 1 | 13/18 | 72.2 | 70.0 |
| Call | Operation | Observable role |
|---|---|---|
| 0 | find_regions | Search for black bowls. S2 is selected from the existing external view based on the ramekin relation. |
| 2 | grasp_candidates | Request nominal grasp hypotheses for S2. No motion is executed by this call. |
| 3–9 | Transit and rotation | Approach with open jaws, rotate in bounded steps, and begin the persistent pose approach. |
| 10, 17 | inspect_region | Refresh S2 using a wrist-image bounding box while approaching the rim. |
| 20 | set_gripper | Close at frame 159 and allow native settling. |
| 21–23 | Inspect, lift, update | Reinspect S2, request a 25 mm lift, and update the carry milestone from the resulting observation. |
| Call | Operation | Observable role |
|---|---|---|
| 0–3 | Ground and inspect | Find the nut, refine its handle as S2, measure depth, and request grasp hypotheses. |
| 4–9 | Approach and close | Record progress, approach using wrist depth, close at frame 79, and request a 3 cm lift. |
| 10–14 | Measure and revisit | Measure the scene and move toward the peg. A new search at frame 171 finds the nut still near its original position. |
| 15–23 | Second attempt | Return with open jaws, remeasure, close at frame 235, and repeat lift and transport. |
| 24–29 | Reground and regrasp | Find the nut again at frame 326, use wrist depth to refine approach, and close at frame 384. |
| 30–31 | Lift and inspect | Lift with closed jaws. The frame-428 wrist image shows the nut near the gripper; a depth query follows. |
| Call | Operation | Observable role |
|---|---|---|
| 0–2 | Ground and propose | Find roller region S1 in the head view, then request grasp hypotheses separately for the left and right arms. |
| 3–5 | move_effectors | Approach, lower, and close both grippers using paired arm targets. |
| 6–7 | Adjust right grasp | Reopen and reposition the right gripper, then close it while retaining the left grasp command. |
| 8–9 | Correct a failed target | A negative left-arm height produces planner failures. The next call restores a positive height and refines the right grasp. |
| 10 | done | Native success is false at command 30. The completion claim is rejected and the episode remains active. |
| 11 | move_effectors | Lift both closed grippers by approximately 6 cm in achieved motion. Native success follows at command 31. |