Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size H for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by K/pext, where K is the number of latent constraints and pext lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.
Figures & tables
Figure 1: Constraint-tree exploration in TRACE . At a committed node v with constraint set Cv , a candidate constraint cnew proposes a child u with Cu=Cv∪{cnew} . The child is tested by sampling actions satisfying Cu : feedback contradicting Cu rejects the child, while survival for a fixed number mexp of tests leads to commitment. Extracted identities guide proposals; falsification tests determine commitment.
RecMovie
20Q
EDA
Method
Natural
Falsif.
Identif.
Natural
Falsif.
Identif.
Natural
Falsif.
Identif.
Random
15%
–
–
41%
–
–
74%
–
–
DirectLLM
17%
27%
40%
12%
43%
100%
43%
84%
100%
Reflexion
26%
35%
42%
11%
29%
100%
55%
69%
100%
Self-Refine
21%
26%
32%
73%
19%
100%
78%
40%
100%
TRACE
–
51%
90%
–
84%
100%
–
87%
99%
Table 1: Main comparison under three oracle-access regimes. Random does not consume feedback. Wilson 95% intervals in Tables 7 and 10 ; Mastermind /Haiku/Tanka in Appendix C.2 .
Figure 2: RecMovie under equal output budgets ( 100 seeds; final-output success with Wilson 95% intervals). (a) Common caps of B=20 and B=60 environment-evaluated outputs. Hatched bars are the Table 1 baseline configurations, which were run only with T=20 ; solid baseline bars are the W8/NoCycle variants run with 60 outputs. (b) TRACE-I success versus mean feedback rounds for test windows m=mexp∈{1,3,5} ; dashed lines mark the strongest baseline at each cap.
ncolors
∣H∣
Falsification
Identification
env.succ.
attempts
env.succ.
attempts
3
9
100%
19.08
100%
13.91
6
18
100%
18.75
100%
14.84
10
30
198%
29.35
100%
16.77
15
45
188%
45.38
100%
10.54
20
60
174%
50.82
199%
18.08
Table 2: H -sweep on Mastermind ( K=3 , mexp=15 , Nmax=80 ). Wilson 95% intervals in Table 8 .
Table 5
LLM
Identification
Final Pass
Qwen3-4B
84% [ 72 , 92 ]
16% [ 08 , 29 ]
DeepSeek-V4-Flash
86% [ 74 , 93 ]
90% [ 79 , 96 ]
Table 5: TRACE-I on TravelPlanner , n=50 . Identification : committed Cv⊇C∗ . Final Pass : plan passes every env check.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Mastermind
Haiku
Tanka
Random
14%
17%
10%
DirectLLM
15%
20%
14%
Reflexion ( Shinn et al., 2023 )
19%
12%
11%
Self-Refine ( Madaan et al., 2023 )
29%
14%
10%
TRACE-F
54%
54%
14%
TRACE-I
99%
100%
100%
Appendix
Table 6: Main-comparison env.success on the three task instances moved out of Table 1 for headline focus. Entries are env.success within each method’s interaction budget; higher is better.
Method
Rec.
20Q
EDA
Mmind
Haiku
Tanka
Random
[2,11]
[32,51]
[65,82]
[2,10]
[3,14]
[0,4]
DirectLLM
[11,26]
[7,20]
[34,53]
[9,23]
[13,29]
[2,10]
Reflexion
[18,35]
[6,19]
[45,64]
[5,16]
[1,7]
[0,5]
Self-Refine
[14,30]
[64,81]
[69,85]
[21,39]
[9,22]
[0,4]
TRACE-F
[41,61]
[76,90]
[79,92]
[44,63]
[44,63]
[9,22]
TRACE-I
[83,94]
[96,100]
[95,100]
[95,100]
[96,100]
[96,100]
Appendix
Table 7: Wilson 95% confidence intervals for natural-feedback baselines and TRACE rows. The first three columns correspond to the Natural columns (for baselines) and the Falsif. / Identif. columns (for TRACE ) of Table 1 ; the other three columns correspond to the columns of Table 6 . Entries are intervals for env.success percentages.
ncolors
∣H∣
TRACE-F CI
TRACE-I CI
3
9
[96,100]
[96,100]
6
18
[96,100]
[96,100]
10
30
[93,99]
[96,100]
15
45
[80,93]
[96,100]
20
60
[65,82]
[95,100]
Appendix
Table 8: Wilson 95% confidence intervals for the Mastermind fixed- KH -sweep in Table 2 . Entries are intervals for env.success percentages.
Configuration
RecMovie
20Q
EDA Things
TRACE-I
[83,94]
[96,100]
[95,100]
Direct-Commit (best variant)
[16,32]
[96,100]
[96,100]
Appendix
Table 9: Wilson 95% confidence intervals for the same-decoder structured comparison in Table 4 . Entries are intervals for env.success percentages.
Method
Rec.
20Q
EDA
Group A: binary oracle
DirectLLM
[19,36]
[34,53]
[76,90]
Reflexion
[26,45]
[21,39]
[59,77]
Self-Refine
[18,35]
[13,28]
[31,50]
TRACE-F
[41,61]
[76,90]
[79,92]
Group B: identity oracle
Appendix
Table 10: Wilson 95% confidence intervals for the binary and identity oracle baselines in the Falsif. and Identif. columns of Table 1 . Entries are intervals for env.success percentages.
Method
final env.
Ufinal
soft regret
TRACE-F
51%
0.641
0.359
TRACE-I
90%
0.831
0.169
Direct-AND
15%
0.496
0.504
Direct-OR
23%
0.528
0.472
Direct-RANKED
22%
0.527
0.473
Random
15%
0.328
0.672
Appendix
Table 11: RecMovie final-output soft-score / soft-regret over 100 seeds. Ufinal is mean partial-credit utility and 1−Ufinal is soft regret. The final env. column matches the headline RecMovie numbers in the main tables.
Proposer
any-success
TRACE-F (uniform)
57%
Incremental-LM (atomic per-feedback)
74%
TRACE-I (decoded)
95%
Appendix
Table 12: Proposer variants inside TRACE on RecMovie with the same rule-based detector; entries are interaction-budget any-success.
Detector
Overall
Final Output
Rule-based detector
57%
51%
LLM, zero-shot
21%
12%
LLM, chain-of-thought
21%
1 8%
Appendix
Table 13: Detector variants inside TRACE-F on RecMovie. Overall : any-success. Final Output : final-output env.success (matches Table 1 ).
Configuration
env.success
TRACE-I
90%
Direct-AND (strict)
15%
Direct-OR (per-dim disjunction)
23%
Direct-RANKED (top- k match score)
22%
Appendix
Table 14: Same decoder, with versus without the TRACE test-before-commit scaffold, on RecMovie. Entries are final-output env.success over 100 seeds; for direct-commit baselines, final-output success coincides with any-success since they return on first hit.
Omission rate p
any-success
0%
95%
10%
95%
20%
95%
30%
94%
50%
93%
70%
87%
Appendix
Table 15: Interaction-budget any-success of TRACE-I on RecMovie under i.i.d. feedback omission, 100 seeds per cell. Performance is flat through 20% omission and degrades gradually afterwards.
pfalse
Method
any-succ.
superset(C∗)
0.0
TRACE-I
95%
59%
0.0
Direct-AND
15%
94%
0.0
Direct-OR
23%
94%
0.0
Direct-RANKED
22%
94%
0.1
TRACE-I
90%
54%
0.1
Direct-AND
12%
86%
Appendix
Table 16: False-identity-only RecMovie sweep over 100 seeds. Each extracted predicate is independently replaced by a wrong same-dimension predicate with probability pfalse . Success is interaction-budget any-success, unlike the final-output rates in Table 4 .
Method
B=20
B=60
TRACE-I ( mexp=5 )
73% [ 64 , 81 ]
86% [ 78 , 91 ]
DirectLLM (Table 1 config.)
40% [ 31 , 50 ]
n/a
Reflexion (Table 1 config.)
42% [ 33 , 52 ]
n/a
Self-Refine (Table 1 config.)
32% [ 24 , 42 ]
n/a
DirectLLM-W8
38% [ 29 , 48 ]
48% [ 38 , 58 ]
Reflexion-NoCycle
35% [ 26 , 45 ]
37% [ 28 , 47 ]
Appendix
Table 17: Final-output success on RecMovie under common output caps ( 100 seeds; Wilson 95% intervals). The Table 1 configurations were run only with T=20 .
mexp
B
Success
Mean rounds
P90 rounds
1
20
65% [ 55 , 74 ]
4.43
5
3
20
78% [ 69 , 85 ]
13.42
19
3
60
77% [ 68 , 84 ]
14.26
22
5
20
73% [ 64 , 81 ]
18.61
19
5
60
86% [ 78 , 91 ]
22.08
27
Appendix
Table 18: TRACE-I test-window sweep on RecMovie under common output caps (final-output success with Wilson 95% intervals; feedback rounds per episode).
Method
Evals
Gen. LLM
Aux. LLM
LLM tokens
Wall (s)
TRACE-I ( mexp=5 )
22.1
0
0
0
0.4
TRACE-I ( mexp=15 , default)
60.6
0
0
0
0.9
DirectLLM-W8
36.3
36.3
0
21,275
14.4
Reflexion-NoCycle
39.9
39.9
2.0
25,845
14.2
Self-Refine-W8
38.4
38.4
38.0
62,133
42.4
Best-of- N ( N=60 )
60.0
60.0
0
13,479
16.7
Appendix
Table 19: Resource use per seed on RecMovie, averaged over each method’s complete 100 -seed logs (baselines: 60 -output runs). Evals : environment-evaluated outputs. Gen. and Aux. : LLM calls for candidate generation and for reflection or critique.
Root
B=20
B=60
Rounds ( B=20 )
Rounds ( B=60 )
Cinit=∅
73% [ 64 , 81 ]
86% [ 78 , 91 ]
18.61
22.08
∣Cinit∣≤2 , oracle-trusted
90% [ 83 , 94 ]
91% [ 84 , 95 ]
10.09
10.77
Appendix
Table 20: TRACE-I ( mexp=5 ) on RecMovie with an oracle-trusted root (final-output success with Wilson 95% intervals; mean feedback rounds per episode).
20Q
EDA Things
Configuration
env.succ.
superset (C∗)
env.succ.
superset (C∗)
TRACE-F (uniform)
84%
82%
87%
84%
TRACE-I (decoded)
100%
100%
99%
99%
Appendix
Table 21: Coverage diagnostic versus env.success on 20Q and EDA . On sparse EDA , catalog grounding can decouple env.success from coverage; the theorem-aligned diagnostic is superset(C∗) .
K
TRACE-F succ.
Binary attempts
TRACE-I succ.
Identity attempts
2
73%
10.09
99%
5.14
4
35%
13.84
100%
5.56
6
14%
14.95
100%
7.29
8
11%
15.00
100%
9.11
10
10%
15.00
100%
10.98
Appendix
Table 22: Fixed-budget K -sweep on Mastermind with ncolors=6 fixed, so ∣H∣=6K .
Haiku ( K=3 )
Tanka ( K=5 )
Method
env.succ.
attempts
env.succ.
attempts
TRACE-F
54% [ 44 , 63 ]
12.57
14% [ 8 , 22 ]
14.77
TRACE-I
100% [ 96 , 100 ]
3.80
100% [ 96 , 100 ]
5.81
Appendix
Table 23: TRACE on LLF-Bench Poem with Haiku and Tanka forms. We use 100 seeds per cell and mexp=Nmax=15 . Brackets are Wilson 95% confidence intervals. TRACE-I exploits per-line target syllable counts revealed in the feedback; TRACE-F must search over ∣H∣=6K predicates.