Task Success Rate

Momentum

8 papers in the last four weeks, up 60% on the four weeks before. 0.1% of all new papers.

Jul 6Week of Sep 21

Latest papers 35

All topics
CardsList
  1. Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

    Oct 1, 2026Ruiyang Si, Jianxin Bi, Shunyu Yang +9Icr-RlTask Success Rate

  2. SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

    Sep 29, 2026Yu Cheng, Yongkang Hu, Shuaijie Ma +12SaferLarge Language Model Agents

  3. AGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic Debridement

    Sep 28, 2026Shutong Jin, Ziyang Chen, Preethi Satish +4Surgical RobotsSurgery

  4. Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

    Sep 28, 2026Dong Xu, Zhangfan Yang, Jiantao Wu +5Task Success RateLarge Language Model Agents

  5. Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics

    Sep 22, 2026Eshika Pathak, Leela KrishnaRobot SystemsTask Success Rate

  6. SAGE: Governed Artifact Generation from Enterprise Guidelines

    Sep 15, 2026Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli +2Task Success Rate

  7. Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    Sep 15, 2026Caiqi Zhang, Xiaochen Zhu, Chengzu Li +3Confidence EstimationConfidence Calibration

  8. WholeBodyWAM: Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination

    Sep 15, 2026Zhuo Li, Yiming Yao, Jim Tan +3Humanoid Loco-ManipulationWhole-Body Control

  9. Signed Sensitivity of Expected Hitting Time to Mutation Rate in the (1+1) EA: Per-State Sign Theorems and Verifiable Certificates for Non-Lumpable Families

    Sep 14, 2026RenKai WangCertificatesTask Success Rate

  10. Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

    Sep 3, 2026Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7Graphical User Interface AgentsGraphical User Interface

  11. AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit

    Aug 17, 2026Yuchen Yuan, Zhenghuang Wu, Yuangan Li +2PilotLarge Language Model Agents

  12. Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    Aug 12, 2026Gen Dong, Yanjie Gao, Liqun Li +3Large Language Model AgentsTask Success Rate

  13. VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

    Aug 11, 2026Haiyu Wu, Randall Balestriero, Morgan LevineLatent World ModelsWorld Model Planning

  14. Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems

    Aug 7, 2026Zhongchao Zhou, Yixuan Xie, Wenwei Yu +5Robust ControlTask Success Rate

  15. A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

    Jul 9, 2026Liuyi Wang, Kai Sheng, Zongtao He +8Vision-Language NavigationVisual Perception

  16. Smooth Operator: A Real-Time Sampling-Based Algorithm for Kinematic Hand Retargeting

    Jul 8, 2026Robert Jomar Malate, Erik Bauer, Norica Bacuieti +4Cross-Embodiment RetargetingTeleoperation

  17. TREK: Distill to Explore, Reinforce to Refine

    Jul 6, 2026Yuanda Xu, Zhengze Zhou, Kayhan Behdin +10Repair-Based Group-Relative Policy OptimizationReasoning Trajectory

  18. Robot Critics that Sweat the Small Stuff

    Jun 19, 2026Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan +3Robot PoliciesTask Success Rate

  19. AutoACSL: Synthesizing ACSL Specifications by Integrating LLMs with CPG-Based Static Analysis

    Jun 18, 2026Han Zhou, Yu Luo, Dianxiang XuProgram AnalysisLlm-Driven Code Synthesis

  20. Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

    Jun 17, 2026Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10Vision-Language Foundation ModelsTask Success Rate

  21. WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

    Jun 11, 2026Arnav Kumar Jain, Yilin Wu, Jesse Farebrother +2Efficient World-Action ModelRobotic Manipulation

  22. Blind Dexterous Grasping via Real2Sim2Real Tactile Policy Learning

    Jun 10, 2026Shengcheng Luo, Xiyan Huang, Zhe Xu +3Robotic GraspingTactile

  23. Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search

    Jun 4, 2026Shing Yin Wong, Shaocheng Liu, Linqi Song +2Theorem ProvingSmall Large Language Models

  24. What are the Right Symmetries for Formal Theorem Proving?

    May 21, 2026Krzysztof Olejniczak, Radoslav Dimitrov, Xingyue Huang +3Theorem ProvingEquivalence

  25. Rollout Cards: A Reproducibility Standard for Agent Research

    May 12, 2026Charlie Masters, Ziyuan Liu, Stefano V. AlbrechtReproducibilityRollout

  26. Autonomous LLM Agents & CTFs: A Second Look

    Apr 29, 2026Youness Bouchari, Matteo Boffa, Marco Mellia +3Attacker Large Language ModelLarge Language Model Agents

  27. Beyond Success Rates: Trainability and Extractability for Offline GCRL

    Feb 5, 2026Jan Malte Töpperwien, Aditya Mohan, Marius LindauerGoal-Conditioned Reinforcement LearningTask Success Rate

  28. First Plan Then Evaluate: Multi-Target Planning with Post-Planning Success Evaluation Improves Learning-Based Grasping Pipelines

    Sep 8, 2025Martin Matak, Mohanraj Devendran Shanthi, Karl Van Wyk +1Robotic GraspingClassical Planning