AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model weights, enable rapid iteration but expose a trade-off: large models adapt zero-shot yet are too slow, whereas compact models meet latency targets but overfit to fixed Harness configurations. We propose Harness-Aware Training (HAT), which trains compact models to adapt to changing Harnesses. Its key component, Harness-State Augmentation (HSA), applies task-preserving transformations to Skill identifiers and content, tool schemas, prompt structures, and Hook functions. Training proceeds in three stages: HSA-SFT learns reasoning and tool use from strong-model trajectories across diverse environments; General On-Policy Distillation restores generalization lost during SFT; and HSA-RL improves robustness to changing Harnesses through reinforcement learning in augmented environments. Across four evaluation sets, HAT achieves 94.8 on Live-Stream QA (base: 80.3; strongest general LLM: 93.0) and 94.6 on Harness-Variant QA (base: 75.4). Unlike Fixed-Harness SFT, which lowers IFEval by 7.7 points from the base model, HAT avoids this regression and reaches 83.5. On one NVIDIA H20 GPU, the optimized system delivers P50 and P95 latencies of 3.4 s and 8.1 s. Deployed in Taobao Live's digital-avatar service, it also yields positive online A/B test results for item-page views.
Figures & tables
Figure 1: Architecture of the live-streaming digital-avatar system and its Harness Agent runtime. (a) Viewer requests and live-room context are processed by the interaction agent. The resulting text response is converted to speech, rendered by the avatar generator, and broadcast to the live room. (b) The expanded runtime operates under an evolvable Harness state h=(S,T,P,K) , comprising active Skills, the tool registry, the assembled system prompt, and lifecycle Hooks, respectively. In each agent-loop round, Hooks mediate input processing, model inference, tool execution, and stopping, while the policy may load Skills, invoke externally mapped tools, or propose a final response. A successful stop check emits the response; a failed check returns control to context assembly for the next round.
Figure 2: Harness Evolution with a fixed DeepSeek-V4-flash on dev-set ( n=482 ). Optimizing the agent without training model. Evolution 2 is the selected development checkpoint Harness, while later long-tail edits introduce regressions.
Scenario
Proportion
Product Q&A
∼ 46%
Casual chat and engagement
∼ 19%
Clarification follow-up
∼ 16%
After-sales handling
∼ 7%
Discount and promotion inquiry
∼ 4%
Presentation-order adjustment
∼ 2%
Table 1: Approximate intent mix in the available internal traffic summary.
Figure 3: Overview of harness-state augmentation and judging. Left: Examples of augmenting tools, skills, hooks, and prompts. The augmented harness is shared across training stages: a teacher model generates candidate trajectories for SFT, while the policy model produces rollouts for RL. Right: Representative RL trajectories. From top to bottom, Trajectory 1 invokes Tool A, which is unavailable after augmentation, and is penalized by Tool Rationality ; Trajectory 2 selects a nonexistent Skill α and is penalized by Skill Selection ; Trajectory 3 violates the modified prompt and is penalized by Accuracy ; and Trajectory 4 correctly adapts to the augmented harness and is rewarded. During SFT, the same judges guide rejection sampling. Together, augmentation and judging teach the model to understand the current harness rather than overfit to a fixed configuration.
Figure 4: Overview of Harness-Aware Training (HAT). Harness-State Augmentation (HSA) transforms the original Harness into diverse, task-preserving states. A strong teacher generates candidate trajectories under these states, which are filtered by a rejection-sampling judge to provide diverse, high-quality supervision for HSA-SFT. Starting from the base model, HAT then follows three stages: HSA-SFT on the filtered trajectories; General OPD, using the base model as the teacher and the HSA-SFT checkpoint as the student to mitigate general-capability degradation; and HSA-RL in changing HSA environments. The resulting policy is the final HAT-trained model.
Set
Size
Source
Evidentiary role
T1 Live-Stream QA
978
Real live-room w/ fixed Harness
Primary industrial quality (live-stream reply).
T2 Harness-Variant QA
978
Real live-room w/ augmented Harness
In-family robustness to new Harness contexts.
T3 Synthetic Live-Stream QA
2023
Synthetic live-room scenarios
In-domain transfer under broader tools/prompts.
T4 IFEval
541
Public Benchmark
General instruction following (official evaluator).
Cjudge
482
Real live-room, human-labeled
Judge calibration, not a policy test set. Also dev-set for harness evolving.
Dperf
110
Deployment replay
End-to-end deployment latency.
Table 2: Evaluation sets and sources. The sets cover real live-stream reply quality, in-family Harness variation, synthetic tool and prompt robustness, general instruction following, and actual deployment latency.
Harness Configuration
T1
T2
T3 Tool Robustness
T3 Prompt Robustness
T4 IFE-P
T4 IFE-I
Non-augmented
Base
80.3
75.4
69.5
72.8
81.5
87.7
+SFT
89.5 ↑ 9.2
88.2 ↑ 12.8
82.0 ↑ 12.5
68.2 ↓ 4.6
73.8 ↓ 7.7
82.4 ↓ 5.3
+General OPD
89.5
89.1 ↑ 0.9
84.3 ↑ 2.3
72.6 ↑ 4.4
82.3 ↑ 8.5
87.9 ↑ 5.5
+RL
95.1 ↑ 5.6
94.4 ↑ 5.3
83.7 ↓ 0.6
66.7 ↓ 5.9
82.7 ↑ 0.4
87.8 ↓ 0.1
Augmented
Table 4: Training-pipeline ablation. Each row adds one stage to the previous checkpoint; ↑ / ↓ denotes a descriptive score change from the row above and does not include retraining variance.
Figure 5: Checkpoint trajectories across four training configurations formed by crossing HSA at the SFT and RL stages, evaluated on Accuracy, Effectiveness, Skill Selection, and Tool Rationality. In these trajectories, HSA-SFT is associated with higher reply-quality rewards, while HSA-RL is associated with stronger agentic-behavior rewards, especially Skill Selection and Tool Rationality; combining HSA-SFT and HSA-RL gives the most balanced late-stage reward profile.
Figure 6: Single-run CoT-length diagnostics for the selected HSA-RL trajectory. Faint lines show recorded step values and bold lines show 25-step trailing means. Panel (a) separates tool-call and final-reply CoT lengths and marks the no-penalty ( L≤100 ), linear ( 100<L<200 ), and saturated ( L≥200 ) regions. Panel (b) shows the logged score q(L)∈[0,1] and the equivalent raw reward contribution −0.1q(L) . The figure is a training diagnostic, not a causal ablation of latency or quality.
Configuration
MTP
C
Wall P50 (s)
Wall P95 (s)
TTFT P95 (s)
Decode (tokens/s)
≤ 15s
Vendor API
DeepSeek-V4-flash
–
1
11.210
21.191
2.067
103.44
71%
DeepSeek-V4-flash
–
2
11.976
26.278
1.291
97.01
71%
DeepSeek-V4-pro
–
1
14.312
28.882
1.312
53.53
61%
DeepSeek-V4-pro
–
2
13.752
23.393
1.248
55.04
60%
Qwen3.6-35B-A3B
Table 5: Controlled low-concurrency deployment replay. Latencies are seconds; Decode is the arithmetic mean of client-observed completion tokens/s over model calls. It is length-sensitive and is not an intrinsic cross-checkpoint kernel-speed measure. TTFT is the first call’s client-observed P95 and includes network and queueing. Each row contains 100 measured Agent requests after 10 warm-up cases. API rows are operational references, not same-hardware model comparisons.
Figure 7: Single-run standalone MTP-loss trajectories for a task-trained policy with a transplanted and subsequently post-trained NextN head, and for the base model with its native factory-trained head. Faint lines show step-level loss; bold lines show a 25-step trailing mean. The runs have different lengths and provide adaptation diagnostics only: they do not compare convergence speed, training stability, sample efficiency, or compute efficiency.
Serving mode
Accuracy
Effectiveness
AVG
MTP Off
95.8
93.8
94.8
MTP On
96.2
94.2
95.2
Table 6: Same-checkpoint T1 quality check with MTP Off and On ( n=978 ). Both rows use the same Final Evaluation Judge. AVG is the arithmetic mean of Accuracy and Effectiveness.
Checkpoint
C
Decode TPS
Speedup
Wall P95 (s)
TTFT P95 (s)
≤ 15s
Qwen3.6-35B-A3B
1
195.42 → 215.95
1.11 ×
10.172 → 9.805
0.508 → 0.502
100 → 99%
2
138.81 → 165.74
1.19 ×
11.504 → 10.251
0.623 → 0.698
97 → 97%
4
105.04 → 113.25
1.08 ×
13.408 → 16.266
0.643 → 1.088
97.3 → 92.3%
8
70.28 → 70.19
1.00 ×
17.848 → 26.136
0.953 → 1.519
90 → 79%
HAT (Ours)
1
160.30 → 271.40
1.69 ×
8.98 → 8.114
0.499 → 0.553
100 → 100%
2
129.22 → 196.15
1.52 ×
10.074 → 9.047
0.692 → 0.754
100 → 100%
Table 7: MTP trade-off across concurrency. Arrows show Off → On. Decode TPS is client-observed mean per call; Wall and TTFT are P95 seconds. C=1, 2, and 8 are single runs. C=4 values are means of three run-level statistics, not pooled percentiles.
Preference
Count
Share
Harness better
35
35.0%
Tie
64
64.0%
ReAct better
1
1.0%
Table 8: Human blind-test preference distribution on 100 paired real live-streaming requests.
Attribution category
Count
Share
More accurate input understanding
12
34.3%
More reliable output
8
22.9%
More appropriate scenario behavior
8
22.9%
Higher response quality
5
14.3%
More appropriate tool use
2
5.7%
Table 9: Attribution of the 35 Harness-preferred examples. Categories are assigned from annotator rationales after the blind decision.
Metric per participating UV
Relative uplift
Platform assessment
Item Page View (IPV)
+0.9107%
Significantly positive
Table 10: UV-normalized item-page-view result from the online A/B test in Taobao Live’s production digital-avatar business. The relative uplift compares the Harness treatment with the ReAct control; the significance assessment is reproduced from the experiment platform.
Long-tail rules interact and regress both metrics.
Evolution 4
Relax length, tool-trigger, parameter, transaction-intent, and attribution restrictions.
91.51
89.96
Partial rollback does not recover Effectiveness.
Appendix
Table 13: Fixed-policy Harness Evolution record. Values come from six direct evaluation summaries on the same 482-item dev-set. “Selected” is an engineering early-stop decision.
Figure 8: Qualitative prompt-edit cases under the same Harness revision. The edit prohibits disclosure-process language when the available evidence does not support a definite conclusion. In both examples, the Fixed-Harness SFT checkpoint continues to expose internal uncertainty after the edit, whereas HAT follows the revised instruction and gives a direct response or routes the viewer to product details or customer service. Red text marks the disclosure phrases targeted by the edit. These examples illustrate compliance behavior and are not additional quantitative evidence.
Figure 9: Judge alignment to a human-labeled calibration cohort ( n=482 ). Accuracy and Effectiveness agreement improve through evidence-tool, rubric, and voting revisions; the cohort evaluates the scoring instrument rather than the policy.
Contrast / dimension
Δ [95% CI]
Pboot(Δ≤0)
IFEval prompt-level accuracy
Fixed-Harness SFT − Base
−6.47[−10.72,−2.22]
0.9992
Ours − Base
+2.03[−1.11,+5.18]
0.1134
Ours − Fixed-Harness SFT
+8.50[+4.44,+12.38]
0.0000
Prompt Robustness
Ours − Fixed-Harness SFT
+8.46[+5.27,+11.66]
<0.0001
Appendix
Table 14: Independent paired-bootstrap stability check. Differences are reported in percentage points; intervals are 95% percentile bootstrap intervals over evaluation items.
Route / checkpoint
MTP
Calls/req
Input/call
Output/call
DeepSeek-V4-flash API
–
4.04
6,052
146
DeepSeek-V4-pro API
–
3.67
5,537
119
Qwen3.6-35B-A3B
Off
3.23
5,622
115
Qwen3.6-35B-A3B
On
3.08
5,571
109
HAT (Ours)
Off
3.04
5,679
106
HAT (Ours)
On
3.12
5,786
105
Appendix
Table 15: C=1 workload shape. Values are means over successful complete-Agent requests or their constituent model calls.
Setting
Engine argument
Value
Speculative algorithm
--speculative-algo
NEXTN
Speculative steps
--speculative-num-steps
3
EAGLE top- k
--speculative-eagle-topk
1
Draft tokens
--speculative-num-draft-tokens
4
Single-token acceptance threshold
--speculative-accept-threshold-single
0.5
Accumulated acceptance threshold
--speculative-accept-threshold-acc
0.7
Appendix
Table 16: MTP serving configuration used for the reported quality check and deployment measurements.
Checkpoint
Accept len
Accept rate
Qwen3.6-35B-A3B
2.94/2.92/3.58
.735/.730/.890
HAT (Ours)
3.12/3.12/3.62
.780/.780/.910
Appendix
Table 17: SGLang log-sample diagnostics for MTP-On serving. Triples are mean/P50/P95.
Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infrastructure from execution feedback, but they are typically evaluated on fixed offline benchmarks. Real deployments instead present open-ended task streams: histories grow without a fixed endpoint, heterogeneous tasks require different harnesses, and problem distributions shift over time. These challenges make a single repeatedly and densely updated harness brittle, causing performance degradation as accuracy peaks early and then declines. This motivates sustained harness construction with task-wise adaptation. We introduce Adaptive Auto-Harness, a framework and system for such streams. The framework decomposes the gap to an oracle harness into evolution loss and adaptation loss. The system addresses these losses with a stateful multi-agent evolver, a harness tree with solve-time routing, and human-steering hooks for cases where history lacks the needed signal. Across prediction-market, security-competition, and event-forecasting streams, Adaptive Auto-Harness outperforms five existing auto-harness baselines and ablations attribute gains to better construction, routing, or targeted human steering. Code is available in Link.
Zewen Liu, Zhan Shi, Yisi Sang +7
Emory University · Amazon · The Pennsylvania State University +2
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses remain largely hand-crafted and static: each new model or task still demands bespoke scaffolding, and the rich traces produced during execution are rarely distilled back into systematic improvement. We introduce HarnessX, a foundry for composable, adaptive, and evolvable agent harnesses. HarnessX assembles typed harness primitives via a substitution algebra, adapts them through AEGIS, a trace-driven multi-agent evolution engine grounded in an operational mirror between symbolic adaptation and reinforcement learning, and closes the harness-model loop by turning trajectories into both harness updates and model training signal. Across five benchmarks (ALFWorld, GAIA, WebShop, tau^3-Bench, and SWE-bench Verified), HarnessX yields an average gain of +14.5% (up to +44.0%), with gains largest where baselines are lowest. These results suggest that agent progress need not come from model scaling alone: composing and evolving runtime interfaces from execution feedback is an actionable and complementary lever. The complete codebase will be open-sourced in a future release.
Streaming video understanding demands more than watching longer videos: assistants must decide when to speak in real time, balancing responsiveness against verbosity. Yet most video-language models (VideoLLMs) are trained for offline inference, and existing streaming benchmarks externalize this timing decision to the evaluator. We address this gap with RealStreamEval, a frame-level multi-turn evaluation protocol that exposes models to sequential observations and penalizes unnecessary responses. Under this protocol, we observed that strong offline VideoLLMs retain useful visual understanding but lack an interaction policy for deciding when to respond. Motivated by this observation, we propose EvoStreaming, a self-evolved streaming adaptation framework in which the base model itself acts as data generator, relevance annotator, and roll-out policy to synthesize streaming trajectories without external supervision. With only 1,000 self-generated samples (139× less than the leading streaming instruction-tuning approach) and no architectural changes, EvoStreaming consistently improves the overall RealStreamEval score by up to 10.8 points across five open VideoLLM backbones (Qwen2/2.5/3-VL, InternVL-3.5, MiniCPM-V4.5) while largely preserving offline video performance. These results suggest that data-efficient interaction tuning is a practical path for adapting existing VideoLLMs to streaming assistants.
Zichen Wen, Boxue Yang, Junlong Ke +5
EPIC Lab, Shanghai Jiao Tong University · Tsinghua University · The Hong Kong University of Science and Technology (Guangzhou) +1