Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.
Figures & tables
Figure 1: Conceptual comparison between a dataset-first annotation workflow and Goldsmith’s definition-first workflow.
Figure 2: Two-stage Goldsmith workflow: gold-loss-guided definition training followed by retrieval- and review-assisted batch annotation.
Figure 3: Goldsmith definition-training loop: textual-gradient edits are evaluated on gold examples and accepted only when they reduce gold loss.
Dataset
Task
Input
Output
OpenNER
Typed spans
Text
PER / ORG / LOC
CrossRE
Relations
Text + HEAD/TAIL pair
Relation label
PHEE V2
Event arguments
Text + trigger
Subject / Treatment / Effect
Table 1: Task interfaces used in the OpenNER, CrossRE, and PHEE V2 experiments.
Setting
Supervision / feedback
P
R
F1
Macro
Exact
Goldsmith score-routed review
15 gold + 193 review
0.8092
0.8257
0.8174
0.8919
0.8140
Goldsmith without review
15 gold
0.7293
0.7505
0.7398
0.8669
0.7749
Goldsmith zero-shot
15 gold
0.7133
0.7440
0.7283
0.8556
0.7613
Initial-prompt zero-shot
0 gold
0.7040
0.7441
0.7235
0.8546
0.7565
mBERT
208 ex.
0.4867
0.5595
0.5206
0.5413
0.5630
XLM-R
208 ex.
0.3281
0.3799
0.3521
0.3735
0.4210
Table 2: OpenNER task-stage results on a shared held-out subset, comparing Goldsmith review settings, zero-shot prompting, and low-budget PLM baselines.
Optimizer
Pass
Loss ↓
Best rnd./gen.
LLM rewrite-only
7/15
22.0303
7
OPRO
9/15
14.5909
16
APE
10/15
11.5606
N/A
PromptBreeder
12/15
8.9091
23 gen.
Goldsmith
14/15
1.8030
18
Table 3: OpenNER prompt-optimization results under a shared initial definition and 15-example gold calibration set. Pass denotes the number of gold examples solved, and lower gold loss is better.
Figure 4: OpenNER gold-loss and token-use trajectories for LLM rewrite-only, OPRO, and Goldsmith. Thick lines show gold loss and thin lines show cumulative token usage over optimization rounds.
Setting
Acc.
Pos. Macro
Pos. R
Pos. F1
Goldsmith
0.8650
0.4273
0.7273
0.5439
Zero-shot LLM
0.8720
0.1640
0.1980
0.2470
BERT
0.3402
0.0883
0.4845
0.1307
RoBERTa
0.3998
0.1108
0.5231
0.1511
DeBERTa
0.4174
0.1088
0.4834
0.1445
Table 4: CrossRE pair-level relation-classification results on the shared 29,171-pair evaluation slice.
Setting
P
R
F1
Δ F1
Ours
0.8092
0.8257
0.8174
–
Initial def.
0.6510
0.6820
0.6661
-0.1513
Random-10 ref.
0.7832
0.8001
0.7915
-0.0259
Table 5: OpenNER ablations of definition optimization and reference selection, reported with precision, recall, and micro-F1.
Backend
Provider
w/o review
w/ review
Exact
GPT-5.5
OpenAI
0.7398
0.8174
0.8140
GPT-5.4
OpenAI
0.7222
0.8012
0.7989
Kimi K2.5
Moonshot
0.7160
0.7890
0.7820
DSv4p
DeepSeek
0.7660
0.7947
0.7863
Table 6: OpenNER backend ablation using the same optimized definition, with and without score-routed review.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Use
Provider family
Run-facing identifier
Access
Temp.
Main runtime settings
Default annotation / judge
OpenAI
gpt-5.5
May 2026
0.0 / 0.0
max tokens 4096; retry bound 20; concurrency 30
OpenNER model ablation
OpenAI
gpt-5.4
May 2026
0.0 / 0.0
same score-routed review protocol as default
OpenNER model ablation
Moonshot AI
kimi-k2.5
May 2026
0.0 / 0.0
same first-pass protocol; pass1 completed-item probe
OpenNER model ablation
DeepSeek
deepseek-v4-pro
May 2026
0.0 / 0.0
same score-routed review protocol as default
Prompt optimization
OpenAI
gpt-5.5
May 2026
0.3 / 0.0
30 rounds; 5 candidates; patience 5
Appendix
Table 7: LLM backends and runtime configurations used for annotation, judging, model ablations, and prompt optimization.
k
Micro P
Micro R
Micro F1
Macro F1
Exact
0
0.7173
0.7296
0.7233
0.8459
0.7360
1
0.7542
0.7652
0.7597
0.8611
0.7480
2
0.7569
0.7599
0.7584
0.8627
0.7600
3
0.7644
0.7704
0.7674
0.8660
0.7620
4
0.7894
0.7863
0.7878
0.8709
0.7720
5
0.7971
0.7876
0.7923
0.8752
0.7720
Appendix
Table 8: OpenNER GT-only diagnostic sweep over the number of retrieved gold references k .
Round
Rewrite loss
OPRO loss
Goldsmith loss
Goldsmith best
Def. tok.
Round tok.
0
25.5152
25.5152
25.5152
25.5152
433
0
1
24.5303
19.5303
16.9697
16.9697
450
95,848
2
22.3636
23.0152
12.5152
12.5152
521
98,151
3
23.1667
17.1970
12.1364
12.1364
578
104,963
4
23.1212
20.6667
12.1364
12.1364
671
112,011
5
25.2727
19.1515
9.5758
9.5758
740
120,960
Appendix
Table 9: Per-round OpenNER gold loss and token usage for LLM rewrite-only, OPRO, and Goldsmith. Definition tokens estimate the accepted Goldsmith definition length; round tokens estimate optimization and evaluation usage.
Parameter
Value
Role
ρ
100
Reporting scale applied to the normalized gold loss
αmiss
1.2
Weight on missed boundaries/relations
αextra
1.0
Weight on extra boundaries/relations
λB
0.75
Boundary component of span loss
λT
0.25
Type component of span loss
λS
0.5
Span component of span–relation loss
Appendix
Table 10: Gold-loss parameters used by the Rosetta implementation. The denominator safeguards max(1,n) prevent division by zero for empty gold or prediction sets. For relation-only classification, label disagreement is the task loss and is not counted again as a separate exact-mismatch penalty.
Round
Gold loss ↓
Pass
Initial
46.6667
8/15
1
46.6667
8/15
2
46.6667
8/15
3
46.6667
8/15
4
33.3333
10/15
5
26.6667
11/15
Appendix
Table 11: CrossRE Goldsmith training trace on the 15-example gold calibration set, showing gold loss and pass count by optimization round.
The Graduate University for Advanced Studies, SOKENDAI · National Institute of Informatics (NII) · BioData Science Initiative (BSI), National Institute of Genetics (NIG)