Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
Organizations: TaoLive AIGC LLM Team
Abstract
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model weights, enable rapid iteration but expose a trade-off: large models adapt zero-shot yet are too slow, whereas compact models meet latency targets but overfit to fixed Harness configurations. We propose Harness-Aware Training (HAT), which trains compact models to adapt to changing Harnesses. Its key component, Harness-State Augmentation (HSA), applies task-preserving transformations to Skill identifiers and content, tool schemas, prompt structures, and Hook functions. Training proceeds in three stages: HSA-SFT learns reasoning and tool use from strong-model trajectories across diverse environments; General On-Policy Distillation restores generalization lost during SFT; and HSA-RL improves robustness to changing Harnesses through reinforcement learning in augmented environments. Across four evaluation sets, HAT achieves 94.8 on Live-Stream QA (base: 80.3; strongest general LLM: 93.0) and 94.6 on Harness-Variant QA (base: 75.4). Unlike Fixed-Harness SFT, which lowers IFEval by 7.7 points from the base model, HAT avoids this regression and reaches 83.5. On one NVIDIA H20 GPU, the optimized system delivers P50 and P95 latencies of 3.4 s and 8.1 s. Deployed in Taobao Live's digital-avatar service, it also yields positive online A/B test results for item-page views.
Figures & tables
| Scenario | Proportion |
|---|---|
| Product Q&A | 46% |
| Casual chat and engagement | 19% |
| Clarification follow-up | 16% |
| After-sales handling | 7% |
| Discount and promotion inquiry | 4% |
| Presentation-order adjustment | 2% |
| Set | Size | Source | Evidentiary role |
|---|---|---|---|
| Live-Stream QA | 978 | Real live-room w/ fixed Harness | Primary industrial quality (live-stream reply). |
| Harness-Variant QA | 978 | Real live-room w/ augmented Harness | In-family robustness to new Harness contexts. |
| Synthetic Live-Stream QA | 2023 | Synthetic live-room scenarios | In-domain transfer under broader tools/prompts. |
| IFEval | 541 | Public Benchmark | General instruction following (official evaluator). |
| 482 | Real live-room, human-labeled | Judge calibration, not a policy test set. Also dev-set for harness evolving. | |
| 110 | Deployment replay | End-to-end deployment latency. |
| Harness Configuration | Tool Robustness | Prompt Robustness | IFE-P | IFE-I | ||
|---|---|---|---|---|---|---|
| Non-augmented | ||||||
| Base | 80.3 | 75.4 | 69.5 | 72.8 | 81.5 | 87.7 |
| +SFT | 89.5 9.2 | 88.2 12.8 | 82.0 12.5 | 68.2 4.6 | 73.8 7.7 | 82.4 5.3 |
| +General OPD | 89.5 | 89.1 0.9 | 84.3 2.3 | 72.6 4.4 | 82.3 8.5 | 87.9 5.5 |
| +RL | 95.1 5.6 | 94.4 5.3 | 83.7 0.6 | 66.7 5.9 | 82.7 0.4 | 87.8 0.1 |
| Augmented | ||||||
| Configuration | MTP | Wall P50 (s) | Wall P95 (s) | TTFT P95 (s) | Decode (tokens/s) | 15s | |
|---|---|---|---|---|---|---|---|
| Vendor API | |||||||
| DeepSeek-V4-flash | – | 1 | 11.210 | 21.191 | 2.067 | 103.44 | 71% |
| DeepSeek-V4-flash | – | 2 | 11.976 | 26.278 | 1.291 | 97.01 | 71% |
| DeepSeek-V4-pro | – | 1 | 14.312 | 28.882 | 1.312 | 53.53 | 61% |
| DeepSeek-V4-pro | – | 2 | 13.752 | 23.393 | 1.248 | 55.04 | 60% |
| Qwen3.6-35B-A3B | |||||||
| Serving mode | Accuracy | Effectiveness | AVG |
|---|---|---|---|
| MTP Off | 95.8 | 93.8 | 94.8 |
| MTP On | 96.2 | 94.2 | 95.2 |
| Checkpoint | Decode TPS | Speedup | Wall P95 (s) | TTFT P95 (s) | 15s | |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | 1 | 195.42 215.95 | 1.11 | 10.172 9.805 | 0.508 0.502 | 100 99% |
| 2 | 138.81 165.74 | 1.19 | 11.504 10.251 | 0.623 0.698 | 97 97% | |
| 4 | 105.04 113.25 | 1.08 | 13.408 16.266 | 0.643 1.088 | 97.3 92.3% | |
| 8 | 70.28 70.19 | 1.00 | 17.848 26.136 | 0.953 1.519 | 90 79% | |
| HAT (Ours) | 1 | 160.30 271.40 | 1.69 | 8.98 8.114 | 0.499 0.553 | 100 100% |
| 2 | 129.22 196.15 | 1.52 | 10.074 9.047 | 0.692 0.754 | 100 100% |
| Preference | Count | Share |
|---|---|---|
| Harness better | 35 | 35.0% |
| Tie | 64 | 64.0% |
| ReAct better | 1 | 1.0% |
| Attribution category | Count | Share |
|---|---|---|
| More accurate input understanding | 12 | 34.3% |
| More reliable output | 8 | 22.9% |
| More appropriate scenario behavior | 8 | 22.9% |
| Higher response quality | 5 | 14.3% |
| More appropriate tool use | 2 | 5.7% |
| Metric per participating UV | Relative uplift | Platform assessment |
|---|---|---|
| Item Page View (IPV) | +0.9107% | Significantly positive |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Responsibility | Interfaces | Role |
|---|---|---|
| Product retrieval | search_product_by_keyword , get_current_product_info , get_product_info_by_link_id , get_product_extra_info_by_link_id_and_keywords , search_preset_faq | Find the active or referenced product and retrieve catalog, detail-page, knowledge-base, or FAQ evidence. |
| Pricing and promotion | get_price_info_by_link_id , get_promotion_info | Retrieve SKU-level prices, entitlements, coupons, and room-level promotions. |
| Action and flow control | change_explain_order , load_skill , get_current_time , refuse_to_reply | Change presentation order, load behavioral instructions, resolve time-sensitive rules, or silently discard meaningless input. |
| MCP lookup | query_promotion , query_buyer_resource , query_fund_asset , query_coupon_detail | Retrieve campaign conditions and buyer-side resources from the external marketing service. |
| MCP calculation | calculate_optimal_promotion , calculate_promotion | Compute eligible promotional combinations and resulting prices. |
| Reply Skill | Response responsibility | Strategy Skill | Tool-chain responsibility |
|---|---|---|---|
| item_qa | Product attributes, specifications, comparison, price, and recommendation. | tool_strategy_benefit | Entitlement, price, and promotion retrieval. |
| chillchat | Non-product conversation and engagement. | tool_strategy_compare | Multi-product comparison and retrieval. |
| aftersale | Returns, exchanges, complaints, and other after-sales requests. | tool_strategy_multi | Aggregation and orchestration for multiple comments. |
| clarification | Follow-up questions when available context is insufficient. | tool_strategy_switch | Presentation-order adjustment. |
| faq_reply | Responses grounded in configured FAQ entries. | general_discount | Room-level promotion retrieval and response composition. |
| refusal | Declines for inappropriate or unsupported requests. |
| Stage | Harness changes | Acc. | Eff. | Engineering diagnosis |
|---|---|---|---|---|
| ReAct | Conventional reasoning–action loop without the modular Harness. | 80.33 | 84.58 | System baseline. |
| Harness base | Initial manually maintained Skills and modular runtime. | 82.40 | 87.16 | Establishes the editable Harness. |
| Evolution 1 | Add refusal tool and stop-loop Hook; rewrite refusal, chat, and after-sales Skills; add four global constraints. | 92.13 | 84.16 | Accuracy rises sharply; internal diagnosis attributes Effectiveness loss to over-triggered refusal. |
| Evolution 2 | Add seven whitelist exclusions before refusal; map explanation triggers and require the relevant tool call. | 92.55 | 92.75 | Selected ; restores Effectiveness while retaining Accuracy. |
| Evolution 3 | Add attribution, factuality, parameter-completeness, tool-use, and system-message rules; expand refusal exclusions. | 91.51 | 90.89 | Long-tail rules interact and regress both metrics. |
| Evolution 4 | Relax length, tool-trigger, parameter, transaction-intent, and attribution restrictions. | 91.51 | 89.96 | Partial rollback does not recover Effectiveness. |
| Contrast / dimension | [95% CI] | |
|---|---|---|
| IFEval prompt-level accuracy | ||
| Fixed-Harness SFT Base | ||
| Ours Base | ||
| Ours Fixed-Harness SFT | ||
| Prompt Robustness | ||
| Ours Fixed-Harness SFT | ||
| Route / checkpoint | MTP | Calls/req | Input/call | Output/call |
|---|---|---|---|---|
| DeepSeek-V4-flash API | – | 4.04 | 6,052 | 146 |
| DeepSeek-V4-pro API | – | 3.67 | 5,537 | 119 |
| Qwen3.6-35B-A3B | Off | 3.23 | 5,622 | 115 |
| Qwen3.6-35B-A3B | On | 3.08 | 5,571 | 109 |
| HAT (Ours) | Off | 3.04 | 5,679 | 106 |
| HAT (Ours) | On | 3.12 | 5,786 | 105 |
| Setting | Engine argument | Value |
|---|---|---|
| Speculative algorithm | --speculative-algo | NEXTN |
| Speculative steps | --speculative-num-steps | 3 |
| EAGLE top- | --speculative-eagle-topk | 1 |
| Draft tokens | --speculative-num-draft-tokens | 4 |
| Single-token acceptance threshold | --speculative-accept-threshold-single | 0.5 |
| Accumulated acceptance threshold | --speculative-accept-threshold-acc | 0.7 |
| Checkpoint | Accept len | Accept rate |
|---|---|---|
| Qwen3.6-35B-A3B | 2.94/2.92/3.58 | .735/.730/.890 |
| HAT (Ours) | 3.12/3.12/3.62 | .780/.780/.910 |