Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, an agentic plug-in framework that replaces source-specific content while preserving task-relevant relations and the required input modality. ReCast locally converts inputs into a shared textual evidence-query record, jointly rewrites entities and topics with a distilled 4B model, and substitutes values through a locally invertible, role-aware numerical map. A reconstruction agent generates and validates the required media from the protected record. The remote solver returns a program whose protected operands are restored locally before execution. On 4,000 held-out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of unprotected remote accuracy, while a model-based audit flags source-content leakage in 7.95% of solver-bound requests. It outperforms all evaluated local baselines, preserving the benefit of remote reasoning while reducing source-content exposure under existing media interfaces.
Figures & tables
Figure 1: ReCast : replacing source content while preserving task structure. Top: The illustrated alternatives address different targets: text rewriting does not directly provide the required media, dataset-level image synthesis is not paired with a specific query, and speaker anonymization retains spoken content. Bottom: ReCast jointly replaces topics, entities, and values across evidence and query while retaining task-relevant relations and the required modality. The remote solver returns a program over protected operands; the client restores the original operands locally before executing the program to recover the answer. The chart example illustrates a unified protection framework for text, charts, and speech.
Figure 2: Overview of ReCast . The pipeline first converts multimodal inputs into a unified textual representation (Sec. 3.2 ), then applies topic rewriting and local numerical remapping to protect entities and numbers (Sec. 3.3 ). The protected text is subsequently reconstructed into a surrogate in original modality, which regenerates charts through ChartSpec-based rendering and validation, passes through text directly, and synthesizes audio as proxy speech (Sec. 4 ). An example in the figure shows this pipeline.
Figure 3: Joint rewriting and numerical protection. A distilled local rewriter changes entities and topics consistently across evidence and query while retaining numerical literals. A separate request-specific map replaces numerical values under the supported relational constraints. Its inverse remains local for operand recovery.
Figure 4: Modality reconstruction and release validation. The interface selects the reconstruction plug-in. Charts use specification generation, local rendering, and review-guided repair; speech uses synthesis followed by transcription-based checks; text is serialized directly. Release requires both protected-record consistency and a local source-consistency check.
(a) Controlled Rewriter comparison
Rewriter
Accuracy ↑
Source-content leakage ↓
ChartQA
NMSQA
Overall
ChartQA
NMSQA
Overall
Qwen3-4B (untuned, local)
66.40
71.00
68.70
37.90
41.60
39.75
ReCast (distilled local 4B)
70.20
80.00
75.10
8.00
7.90
7.95
GPT-5.6-Sol (offline teacher)
70.30
81.00
75.65
5.20
1.20
3.20
Δ vs. untuned (pp)
+3.80
+9.00
+6.40
−29.90
−33.70
−31.80
Table 1: Stage 2 Rewriter comparison (a) and native exposure diagnostics of protection references (b). Values are percentages; Δ : ReCast minus the reference (pp); “–”: unavailable; ∗ : single-dataset value.
Table 6Table 7
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Item
ChartQA
NMSQA
ChartX
Spoken-MQA
Evaluation setting
ID
ID
OOD
OOD
Evaluated examples
2,000
2,000
2,000
2,000
Source pool
Held-out evaluation partition
Held-out evaluation partition
Official benchmark pool
Official benchmark pool
Evidence modality
Image
Audio
Image
Audio
Query / instruction modality
Text
Text
Text
Audio problem with textual instruction
Eligibility
Single numeric answer; readable PNG; duplicate questions removed
Single consistent numeric reference; passage length ≤350 words
QA task; parseable single numeric answer; readable chart image
Valid audio; parseable numeric reference
Appendix
Table 6: Construction of the evaluation subsets. ID denotes in-distribution evaluation, while OOD denotes datasets not used for Rewriter distillation or checkpoint selection. All evaluation sample IDs are fixed before model execution.
Component
Model or implementation
Location
Input
Text parser
deterministic parser with field and schema validation
Local
Original text evidence and question
Chart-to-text
Qwen3-VL-8B-Instruct, question-aware prompt
Local
Original chart and question
Speech-to-text
Qwen3-ASR-1.7B
Local
Evidence audio
Numeral normalization
text2num , isolated number words included
Local
Evidence, question, choices
Rewriter
Qwen3-4B + distilled QLoRA adapter
Local
Normalized five-field record
Numerical map
role-aware banded map (Appendix A.3 )
Local
Rewritten evidence and question
Appendix
Table 7: Component configuration.
Item
Total
NMSQA
ChartQA
Source records
16,370
7,000
9,370
Accepted pairs
15,795
6,742
9,053
Rejected (changed numerical literals)
575
258
317
Training pairs
15,480
6,604
8,876
Validation pairs
315
138
177
Appendix
Table 8: Distillation corpus.
Figure 5: Chart types of the image records (a; 32 records with no identifiable type are omitted) and topics of the audio records (b) in the distillation corpus.
Figure 6: Question operations within each chart type of the image records. Each row sums to 100%; grey numbers give the share of each operation over all image records.
Figure 7: Topics of the image (a) and audio (b) records. Circle area is proportional to the share of records; circle position carries no meaning.
Figure 8: Most frequent content words in the image (a, b) and audio (c, d) records before and after teacher rewriting. Word size and shade increase with frequency; chart-description and function words are removed.
Model
ChartQA
NMSQA
Overall
Gemma-4-E4B
49.50
44.60
47.05
Qwen2.5-Omni-3B
23.70
46.30
35.00
Qwen3-VL-8B-Instruct image-to-code
62.80
–
–
Appendix
Table 9: Accuracy of evaluated local references.
Method
Accuracy ↑
Leakage ↓
ChartQA
NMSQA
Overall
ChartQA
NMSQA
Overall
Local, original input
Qwen2.5-Omni-3B
23.70
46.30
35.00
0
0
0
Gemma-4-E4B
49.50
44.60
47.05
0
0
0
Qwen3-VL-8B
62.80
–
–
0
0
0
ReCast (Ours)
70.20
80.00
75.10
8.00
7.90
7.95
Appendix
Table 10: Accuracy and leakage (%) of local models, ReCast , and the remote solver.
Rewriter
ChartQA coverage
ChartQA leakage
NMSQA coverage
NMSQA leakage
RLAA-Qwen3-4B
100.00
86.80
100.00
73.70
Eternis-Anonymizer-4B
100.00
96.70
100.00
93.20
Appendix
Table 11: Per-modality coverage and leakage.
View
Judge-human agreement (%)
Cohen’s κ (judge vs. human)
Inter-annotator κ
Semantic
92.0
0.82
0.81
Lexical
89.5
0.74
0.72
Correspondence
93.8
0.85
0.84
Linkability
88.3
0.70
0.69
Any-positive (reported metric)
91.5
0.83
0.78
Appendix
Table 12: Agreement between the judge and human labels.
Parameter
Value
Student
Qwen3-4B, QLoRA (NF4, bfloat16 compute)
LoRA rank / α / dropout
16 / 32 / 0.05, all linear layers
Learning rate / schedule
1e-4, cosine, 0.03 warm-up ratio
Epochs / gradient accumulation
2 / 8
Maximum sequence length
4,096
Loss
assistant tokens only
Appendix
Table 13: Rewriter training.
Component
Setting
Chart-to-text
greedy decoding, bfloat16, max 2,200 new tokens
Speech-to-text
bfloat16, batch size 1, max 512 new tokens
Rewriter
thinking disabled, temperature 0.2, top- p 0.9, max 16,384 new tokens