Synthetic Task

Momentum

10 papers in the last four weeks, against 1 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 69

All topics
CardsList
  1. GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

    Sep 30, 2026Qisheng Su, Hanchen Wang, Guanru Zhu +10Synthetic Task

  2. Can a Cacheable Decision Model Follow Rules?

    Sep 29, 2026Dushyant Rajput, Nirdesh Chauhan, Siddharth KosarajuSynthetic Task

  3. Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

    Sep 28, 2026Dongsheng Ma, Sizhe Wang, Xinyi Huang +5Coding AgentsReuse

  4. From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction

    Sep 28, 2026Zixiao Dong, Wei Yang, Zihao Liu +4Synthetic TaskExtraction

  5. Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

    Sep 21, 2026Haocheng Xia, Eugene Wu, Yongjoo ParkPull RequestsMulti-Llm Agents

  6. Limits of Confidence in Diffusion

    Sep 17, 2026Russ Webb, Amitis Shidani, Alice Bizeul +1Synthetic Task

  7. Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

    Sep 14, 2026Mykhailo Kozyrev, Andrei Kozyrev, Anton PodkopaevRepository-Level Code UnderstandingCoding Agents

  8. CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

    Sep 7, 2026Yonghong Zhang, Ricardo Correia, Isabel M. Parra +1Causal InferencesCode Generation

  9. FrogNano: Training a 4B Coding Agent via Online Task Synthesis

    Sep 7, 2026Minseon Kim, Zhengyan Shi, Emiliano Penaloza +14Coding AgentsSynthetic Task

  10. SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

    Aug 12, 2026Ruitao Wang, Yuwen Hao, Menglin YangWeb AgentsSynthetic Task

  11. CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

    Aug 6, 2026Fanzhe Meng, Guoxin Chen, Jiale Zhao +6Agentic BenchmarksSynthetic Task

  12. Recursive Synthesis for Long-Horizon Terminal Tasks

    Aug 5, 2026Zhongzhi Li, Yucheng Shi, Zongxia Li +8Synthetic TaskLong-Horizon Agents

  13. AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

    Aug 3, 2026Qiushi Lin, Chaojie Zhang, Íñigo Goiri +3Artificial Intelligence InfrastructureData Centers

  14. SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

    Aug 3, 2026Zelin Tan, Yiqun Zhang, Hao Li +11Synthetic TaskLanguage-Model Agents

  15. Lossless Tensor Compression as Program Synthesis

    Aug 3, 2026Jieke Shi, Junda He, Wenjia Jiang +11Data Compression MethodsIntermediate Checkpoints

  16. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates

    Jul 31, 2026Bohan Chen, Shivam N. Patel, Richard Hoffmann +2Symbol-Theorem Proving

  17. Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models

    Jul 30, 2026Yang Li, Ping Hou, Nobuko YoshidaInvariant SynthesisProtocol

  18. Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training

    Jul 30, 2026Zhihong Pan, Jiyuan He, Kai Zhang +5Synthetic TaskAgentic Benchmarks

  19. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

    Jul 26, 2026Weihuang Zheng, Tianyuan Zou, Eileen Ye +5Agentic BenchmarksEvaluation Agent

  20. SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

    Jul 25, 2026Lang Mei, Xiaohan Yu, Chong Chen +27Synthetic TaskSynthetic Data

  21. Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation

    Jul 20, 2026Siddharth Mishra-SharmaModel SelectionSynthetic Task

  22. NexForge: Scaling Executable Agent Tasks via Requirement-First Synthesis

    Jul 15, 2026Jiarong Zhao, Zhikai Lei, Zhiheng Xi +5Synthetic TaskForge

  23. When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

    Jul 11, 2026Cheng-Ting Chou, Duc Binh HoangClass ImbalanceSpurious Correlations

  24. Generative Skill Composition for LLM Agents

    Jun 30, 2026Xinyu Zhao, Zhen Tan, Vaishnav Tadiparthi +5SkillsLarge Language Model Agents

  25. Axon: A Synthesizing Superoptimizer for Tensor Programs

    Jun 24, 2026Akash Kothari, Shaowei Zhu, Daniel Kroening +1Synthetic Task

  26. AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs

    Jun 24, 2026Dhruv Sharma, Gautam ShroffAlgorithmic TradingEvolutionary Search

  27. CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents

    Jun 22, 2026Zhanbo Hua, Yifan Yao, Weihao Xie +14Synthetic TaskSynthesis

  28. Explaining Attention with Program Synthesis

    Jun 17, 2026Amiri Hayes, Belinda Z Li, Jacob AndreasTransformer ArchitecturesSynthetic Task

  29. Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier

    Jun 10, 2026Lorenz Wolf, Connor Watts, Roger Creus Castanyer +4Functional Code SolversSynthetic Task

  30. A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions

    Jun 6, 2026Lukas Fesser*, Mozes Jacobs*, Thomas Fel* +2PerspectiveStreaming

  31. WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

    Jun 1, 2026Hengrui Gu, Xiaotian Han, Kaixiong ZhouSynthetic TaskSynthesis

  32. SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

    Jun 1, 2026Hao Cheng, Changtao Miao, Tianle Song +21Security EvaluationAutonomous Agents

  33. BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

    May 31, 2026Yangzhen Wu, Aaron J. Li, Wenjie Ma +10Large Language Model BenchmarksSynthetic Task

  34. Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

    May 29, 2026Jiasheng Zheng, Boxi Cao, Boxi Yu +6Reinforcement Learning With Verifiable RewardSynthetic Task

  35. unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning

    May 27, 2026Geoffrey Bradway, Roger Creus Castanyer, Lorenz Wolf +3RobotwinExploitation

  36. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    May 27, 2026Tomer Keren, Nitay Calderon, Asaf Yehudai +3Agentic BenchmarksSynthetic Task

  37. A Systematic Study of Behavioral Cloning for Scientific Data Annotation

    May 26, 2026Ishaan Singh Chandok, Core Francisco ParkBehavior CloningHuman Annotations

  38. QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

    May 22, 2026Jian Xie, Tianhe Lin, Zilu Wang +16Deep ResearchSynthetic Task

  39. Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

    May 20, 2026Zihao Cheng, Hongru Wang, Zeming Liu +6Synthetic TaskSingle-Agent Baselines

  40. Mechanisms of Misgeneralization in Physical Sequence Modeling

    May 19, 2026Kento Nishi, Raphael Tang, Karun Kumar +2Physics-Aware ModelsGenerative Models

  41. Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback

    May 17, 2026Guijin Son, Jehyun Park, Seyeon Park +2EngineeringFinite Element Method

  42. Large language models reorganize representational geometry during in-context learning

    May 16, 2026Hua-Dong Xiong, Li Ji-An, Robert C. Wilson +2In-Context LearningRepresentational Capacity

  43. Property-Guided LLM Program Synthesis for Planning

    May 15, 2026André G. Pereira, Augusto B. Corrêa, Jendrik SeippLlm-Driven Code SynthesisLarge Language Model Planning

  44. Orchard: An Open-Source Agentic Modeling Framework

    May 14, 2026Baolin Peng, Wenlin Yao, Qianhui Wu +11Agentic FrameworkSynthetic Task

  45. The two clocks and the innovation window: When and how generative models learn rules

    May 11, 2026Binxu Wang, Emma Lucia Byrnes Finn, Bingbin LiuGenerative ModelsClocks

  46. Towards Robust Sequential Decomposition for Complex Image Editing

    May 10, 2026Zilai Zeng, Mingdeng Cao, Zijie Li +5Image EditingMulti-Reference Image Generation

  47. ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis

    May 6, 2026Atharva Naik, Yash Mathur, Prakam +2Llm-Driven Code SynthesisSynthetic Task

  48. Skill Neologisms: Towards Skill-based Continual Learning

    May 6, 2026Antonin Berthon, Nicolas Astorga, Mihaela van der SchaarUnsupervised Skill DiscoverySkills

  49. Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

    Apr 30, 2026Zainab Rehan, Christian Medeiros Adriano, Sona Ghahremani +1Neuro-Symbolic FrameworkSafety-Critical Scenarios

  50. Beyond the Training Distribution: Mapping Generalization Boundaries in Neural Program Synthesis

    Apr 30, 2026Henrik Voigt, Michael Habeck, Joachim GiesenImproved GeneralizationSynthetic Task

  51. Toward Scalable Terminal Task Synthesis via Skill Graphs

    Apr 28, 2026Zhiyuan Fan, Tinghao Yu, Yuanjun Cai +8Synthetic TaskAgentic

  52. Why Search When You Can Transfer? Amortized Agentic Workflow Design from Structural Priors

    Apr 27, 2026Shiyi Du, Jiayuan Liu, Weihua Du +6Agentic WorkflowsAgentic Workflow Design