Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
Figures & tables
Figure 1: Coding-assistant benchmarks by development horizon (horizontal: engineering scope and continuity) and interaction dynamics (vertical: adaptation to agent outputs and user states). SWE-Journey (Ours) occupies the upper-right region, near the highest level on both axes: project-scale development involving building, replicating, optimizing, or evolving an entire system, together with user-grounded interaction shaped by user states, preferences, personas, and evolving goals.
Figure 2: Overview of SWE-Journey . (A) The task-generation loop iteratively synthesizes a long-horizon task, realizing each subtask via step-request, patch-and-F2P, and P2P generation gated by a rubric and execution checks. (B) The persona taxonomy is mined from real logs, yielding five persona dimensions. (C) Four representative personas driving multi-turn interaction.
Table 1: Main results by model on SWE-Journey ; each model is the arithmetic mean of its four personas. Average is (Proc. F2P+Final F2P+Final P2P)/3 . Rows are sorted by the Average point estimate; bold marks the best point estimate in each column.
Persona
Turn/Session
Proc. F2P
Final F2P
Final P2P
Average
Cost/Session ($)
Non-coder
13.56
20.22
23.00
84.62
42.61
12.81
Product Manager
7.83
26.50
30.48
85.79
47.59
6.85
New Developer
25.84
28.62
35.57
88.15
50.78
7.07
Software Architect
25.92
78.07
78.50
97.50
84.69
13.90
Table 2: Main results by persona on SWE-Journey ; each persona is the mean of the thirteen models.
Table 3: Fine-grained capability breakdown. Green indicates better performance; bold indicates the best value in each column. − indicates that Fix is unavailable for closed-source models because their internal reasoning traces are inaccessible, precluding analysis of this dimension.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Knowledge domain
Verified
Pro
EVO
Mara.
Ours
Core Task
Software testing and quality assurance
75.6
86.2
91.7
60.0
93.3
Source-code reading, program analysis, and code search
61.8
59.4
22.9
35.0
73.3
Debugging, troubleshooting, and performance diagnosis
63.0
27.1
22.9
25.0
30.0
Version control, dependency compatibility, and release management
1.4
5.1
56.2
0.0
0.0
Command-line, shell, and CLI tool development
1.6
5.5
27.1
40.0
6.7
Build systems and package management
0.2
2.3
27.1
40.0
0.0
Appendix
Table 4: Knowledge-domain coverage by canonical dimension. Percentages use the latest classification batch; the Ours column is computed over the 30 synthesized cases. We analyze four typical benchmarks: SWE-bench Verified ( Jimenez et al., 2024 ) , SWE-bench Pro ( Deng et al., 2025b ) , SWE-EVO ( Thai et al., 2025 ) , and SWE-Marathon ( Desai et al., 2026 ) .
New task angle
Verified
Pro
EVO
Marathon
Ours
Bug fixing and semantic-correctness alignment
90.0
53.1
54.2
5.0
70.0
Feature addition and capability expansion
6.4
27.2
25.0
65.0
56.7
Support-scope expansion, compatibility adaptation, and technology migration
3.2
9.6
18.8
35.0
13.3
Architectural refactoring, abstraction decoupling, and code-structure cleanup
2.0
13.7
20.8
0.0
30.0
Authentication, authorization, security, and compliance governance
1.2
9.2
2.1
15.0
23.3
User interface, pages, and interaction development
0.0
6.4
0.0
20.0
3.3
Appendix
Table 5: Distribution over the 20-angle synthesis pool across SWE-bench Verified ( Jimenez et al., 2024 ) , SWE-bench Pro ( Deng et al., 2025b ) , SWE-EVO ( Thai et al., 2025 ) , SWE-Marathon ( Desai et al., 2026 ) , and ours; the Ours column is computed over the 30 synthesized cases using the canonical 20-angle taxonomy.
Dimension
Category
Definition
User Expertise
Non-coder
Lacks programming or repository knowledge; can describe symptoms, errors, or desired effects, but struggles to judge implementation correctness independently.
Product Manager
Understands business goals, expected behavior, or problem symptoms, but usually does not know implementation details, file structure, or test constraints.
New Developer
Has general engineering ability; can read and write relevant code, run local commands or tests, and make evidence-based correctness judgments, but is still not fully familiar with the current repository.
Software Architect
Knows the repository structure, module boundaries, historical conventions, and compatibility constraints, and constrains the coding assistant’s changes with expert-level standards.
Detail Preference
Brief
Wants the coding assistant to give conclusions, commands, changes, or next steps directly, avoiding long background and lengthy explanations.
Structured
Wants the coding assistant to separate causes, fixes, risks, and next steps clearly, making the response easy to read, execute, and check.
Appendix
Table 6: User persona dimensions and category definitions.
Persona
User Expertise
Detail Preference
Control Style
Goal Stability
Verification Style
Non-coder
Non-coder
Structured
Hands-off
Pivoting
No User Verification
Product Manager
Product Manager
Brief
Hands-off
Stable
No User Verification
New Developer
New Developer
Brief
Hands-on
Pivoting
Single Check
Software Architect
Software Architect
Structured
Align-first
Evolving
Systematic Verification
Appendix
Table 7: Representative simulated user configurations.
Non-coder ( n=27 )
Dimension
Category
× Avg
Detail Preference
Brief
× 0.46
Detailed
× 0.00
Structured
× 2.91
Control Style
Hands-off
× 3.80
Align-first
× 0.00
Appendix
Table 8: Per-persona behavior profile on SWE-Chat. Instances are grouped by user expertise (D1) into the four representative personas, and each task reports how often a D2–D5 category occurs within that group as a multiple of its overall Log frequency. Values >×1 occur more often than the population average; values <×1 less often.
Benchmark
#Inst.
Files
+Lines
− Lines
SWE-bench Verified ( Jimenez et al., 2024 )
500
1.25
9.9
4.4
SWE-bench ( Jimenez et al., 2024 )
2294
1.66
26.7
11.0
SWE-bench Multilingual
300
1.74
28.9
18.9
SWE-Gym ( Pan et al., 2024 )
2438
2.48
55.6
14.2
SWE-bench Pro ( Deng et al., 2025b )
731
5.07
120.3
49.2
SWE-EVO ( Thai et al., 2025 )
48
21.15
357.9
252.4
Appendix
Table 9: Reference-patch scale, recomputed uniformly: files count diff --git headers; lines count additions plus deletions, excluding file headers. Rows are sorted by total changed lines. SWE-Journey aggregates all subtask patches per instance; others use one reference commit.
Figure 3: Coverage and composition of SWE-Journey : a five-axis radar characterizing each dataset by task coverage and composition.
Figure 4: Authenticity of SWE-Journey (1–5 mean scores; higher is better). (a) Task-authenticity over five dimensions; (b) interaction-authenticity over five dimensions, on real logs (Log) and our synthesized SWE-Journey interactions (Ours).
Table 10: Single-shot long-horizon test-node pass rates (one mega-prompt, no interaction, no personas). Final F2P/P2P pool passed and recorded node counts across instances, using the same post-processing as the interactive results. Average is (Final F2P+Final P2P)/2 . Rows are sorted by the Average point estimate; bold marks column-best point estimates.
Table 11: Main results for the Non-coder persona.
Table 12: Main results for the Product Manager persona.
Table 13: Main results for the New Developer persona.
Table 14: Main results for the Software Architect persona.
Persona
Conf
Ask
Find
Fix
Ask% ↑
Rel% ↑
Miss% ↓
OC% ↓
Skip% ↓
Crit% ↓
Non-coder
58.92
29.89
19.94
72.18
23.14
14.10
11.43
Product Manager
60.14
19.59
16.67
55.56
20.24
14.75
12.69
New Developer
67.94
26.41
23.90
40.68
2.81
6.77
4.64
Software Architect
42.26
40.27
37.58
22.52
6.47
8.30
5.37
Appendix
Table 15: Per-persona capability breakdown. Columns and conventions as in Table 3 ; Fixing Right columns average the eleven models with available reasoning.
Table 20: Answer-seeking retrieval behavior. Retr. turns/session is the mean retrieval turns per completed session; Reach% and Succ% use all coding-assistant turns. The final columns partition process F2P by provenance: searchable original SEED F2P , synthesized follow-up Synth. F2P , and Δ= SEED − Synth. The four most frequent searchers gain 23 – 40 points on seed tasks, whereas infrequent retrievers show no advantage or a reversal.
Retr. turns/session
Reach%
Succ%
Succ/Reach%
By retrieval method
Local git history
0.75
1.60
0.20
12.33
Repository clone / fetch
0.50
2.08
0.22
10.56
GitHub code search
0.38
1.69
0.16
9.44
Package upgrade
0.03
0.02
0.00
20.00
By user persona
Appendix
Table 21: Answer-seeking retrieval behavior by retrieval method (top) and user persona (bottom), over completed cells. Columns follow Table 20 ; method rows use all sessions, and persona rows use that persona’s own sessions/turns.