SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
Organizations: Tencent Hy AI Data · Beijing Zhongguancun Academy
Abstract
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
Figures & tables
| Persona | Turn/Session | Proc. F2P | Final F2P | Final P2P | Average | Cost/Session ($) |
|---|---|---|---|---|---|---|
| Non-coder | 13.56 | 20.22 | 23.00 | 84.62 | 42.61 | 12.81 |
| Product Manager | 7.83 | 26.50 | 30.48 | 85.79 | 47.59 | 6.85 |
| New Developer | 25.84 | 28.62 | 35.57 | 88.15 | 50.78 | 7.07 |
| Software Architect | 25.92 | 78.07 | 78.50 | 97.50 | 84.69 | 13.90 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Knowledge domain | Verified | Pro | EVO | Mara. | Ours | |
| Core Task | Software testing and quality assurance | 75.6 | 86.2 | 91.7 | 60.0 | 93.3 |
| Source-code reading, program analysis, and code search | 61.8 | 59.4 | 22.9 | 35.0 | 73.3 | |
| Debugging, troubleshooting, and performance diagnosis | 63.0 | 27.1 | 22.9 | 25.0 | 30.0 | |
| Version control, dependency compatibility, and release management | 1.4 | 5.1 | 56.2 | 0.0 | 0.0 | |
| Command-line, shell, and CLI tool development | 1.6 | 5.5 | 27.1 | 40.0 | 6.7 | |
| Build systems and package management | 0.2 | 2.3 | 27.1 | 40.0 | 0.0 |
| New task angle | Verified | Pro | EVO | Marathon | Ours |
|---|---|---|---|---|---|
| Bug fixing and semantic-correctness alignment | 90.0 | 53.1 | 54.2 | 5.0 | 70.0 |
| Feature addition and capability expansion | 6.4 | 27.2 | 25.0 | 65.0 | 56.7 |
| Support-scope expansion, compatibility adaptation, and technology migration | 3.2 | 9.6 | 18.8 | 35.0 | 13.3 |
| Architectural refactoring, abstraction decoupling, and code-structure cleanup | 2.0 | 13.7 | 20.8 | 0.0 | 30.0 |
| Authentication, authorization, security, and compliance governance | 1.2 | 9.2 | 2.1 | 15.0 | 23.3 |
| User interface, pages, and interaction development | 0.0 | 6.4 | 0.0 | 20.0 | 3.3 |
| Dimension | Category | Definition |
|---|---|---|
| User Expertise | Non-coder | Lacks programming or repository knowledge; can describe symptoms, errors, or desired effects, but struggles to judge implementation correctness independently. |
| Product Manager | Understands business goals, expected behavior, or problem symptoms, but usually does not know implementation details, file structure, or test constraints. | |
| New Developer | Has general engineering ability; can read and write relevant code, run local commands or tests, and make evidence-based correctness judgments, but is still not fully familiar with the current repository. | |
| Software Architect | Knows the repository structure, module boundaries, historical conventions, and compatibility constraints, and constrains the coding assistant’s changes with expert-level standards. | |
| Detail Preference | Brief | Wants the coding assistant to give conclusions, commands, changes, or next steps directly, avoiding long background and lengthy explanations. |
| Structured | Wants the coding assistant to separate causes, fixes, risks, and next steps clearly, making the response easy to read, execute, and check. |
| Persona | User Expertise | Detail Preference | Control Style | Goal Stability | Verification Style |
|---|---|---|---|---|---|
| Non-coder | Non-coder | Structured | Hands-off | Pivoting | No User Verification |
| Product Manager | Product Manager | Brief | Hands-off | Stable | No User Verification |
| New Developer | New Developer | Brief | Hands-on | Pivoting | Single Check |
| Software Architect | Software Architect | Structured | Align-first | Evolving | Systematic Verification |
| Non-coder ( ) | ||
|---|---|---|
| Dimension | Category | Avg |
| Detail Preference | Brief | 0.46 |
| Detailed | 0.00 | |
| Structured | 2.91 | |
| Control Style | Hands-off | 3.80 |
| Align-first | 0.00 | |
| Benchmark | #Inst. | Files | +Lines | Lines |
|---|---|---|---|---|
| SWE-bench Verified ( Jimenez et al., 2024 ) | 500 | 1.25 | 9.9 | 4.4 |
| SWE-bench ( Jimenez et al., 2024 ) | 2294 | 1.66 | 26.7 | 11.0 |
| SWE-bench Multilingual | 300 | 1.74 | 28.9 | 18.9 |
| SWE-Gym ( Pan et al., 2024 ) | 2438 | 2.48 | 55.6 | 14.2 |
| SWE-bench Pro ( Deng et al., 2025b ) | 731 | 5.07 | 120.3 | 49.2 |
| SWE-EVO ( Thai et al., 2025 ) | 48 | 21.15 | 357.9 | 252.4 |
| Persona | Conf | Ask | Find | Fix | |||
|---|---|---|---|---|---|---|---|
| Ask% | Rel% | Miss% | OC% | Skip% | Crit% | ||
| Non-coder | 58.92 | 29.89 | 19.94 | 72.18 | 23.14 | 14.10 | 11.43 |
| Product Manager | 60.14 | 19.59 | 16.67 | 55.56 | 20.24 | 14.75 | 12.69 |
| New Developer | 67.94 | 26.41 | 23.90 | 40.68 | 2.81 | 6.77 | 4.64 |
| Software Architect | 42.26 | 40.27 | 37.58 | 22.52 | 6.47 | 8.30 | 5.37 |
| Retr. turns/session | Reach% | Succ% | Succ/Reach% | |
| By retrieval method | ||||
| Local git history | 0.75 | 1.60 | 0.20 | 12.33 |
| Repository clone / fetch | 0.50 | 2.08 | 0.22 | 10.56 |
| GitHub code search | 0.38 | 1.69 | 0.16 | 9.44 |
| Package upgrade | 0.03 | 0.02 | 0.00 | 20.00 |
| By user persona | ||||