QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations
Authors: Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel, Yelena Mejova, Mariano G. Beiró, Kyriaki Kalimeri
Organizations: Universidad de San Andrés, Buenos Aires, Argentina · ISI Foundation, Turin, Italy · CONICET, Buenos Aires, Argentina · UNICEF, New York, NY, USA
Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.
Figures & tables
System
Two-stream
Field-level
Char-level
IAA /
Decision
reconcile
hybrid
align
campaign
audit log
Label Studio
partial
–
–
partial
partial
INCEpTION
partial
partial
–
✓
partial
Prodigy
partial
–
–
–
partial
Argilla
partial
–
–
partial
✓
QuanReview
✓
✓
✓
✓
✓
Table 1: Feature comparison across annotation and adjudication tools. QuanReview is the only system combining two-stream reconciliation, field-level hybrids, and character-level alignment. Partial denotes support for a related capability but not the complete post-hoc two-stream workflow defined here.
Figure 1: System architecture. Reference (HumSet) and model (LLM) spans are aligned at character level; the policy engine auto-resolves one-sided records (reference-only kept, model-only deferred) and routes conflicting pairs to the adjudication UI. Reviewed documents are merged across annotators, unanimous ones automatically, into the corrected layer, which replaces the original files directly; flagged pairs and deferred model-only records are queued for a second annotation round.
Figure 2: The adjudication interface, showing one conflict pair. The two candidates appear under neutral, order-randomised labels; in this pair they agree on quantity, unit, and modifier but differ on event type ( EventO vs. EventP ), the disagreement the reviewer resolves below.
Rev.
Docs
Pairs
Anns
Flags
Flag
Changed
rate
vs ref.
R1
141
745
996
166
22.3%
42.3%
R2
138
761
609
69
9.1%
20.9%
R3
133
703
978
14
2.0%
71.3%
R4
118
580
539
85
14.7%
12.4%
R5
65
352
406
53
15.1%
52.8%
Table 2: Per-reviewer activity over the eight-reviewer campaign. Docs and Pairs count the documents and discrepant pairs each reviewer adjudicated, Anns the annotations they saved, Flag rate the share of their pairs flagged, and Changed vs ref. the share of pairs whose decision differs from the reference. Rates are descriptive because reviewers were assigned different sets of conflict pairs. Reviewers are pseudonymous.
Outcome
Anns
%
Reference accepted unchanged
522
12.7
Model accepted unchanged
24
0.6
Field-level hybrid
2,604
63.2
Both streams agreed
971
23.6
Total
4,121
100.0
Table 3: Distribution of final annotation outcomes recorded by the campaign. Both streams agreed denotes automatically retained matched records that were not routed to reviewers.
Field
κ (all pairs)
κ (unflagged)
Event type
0.225
0.473
Quantity
0.688
0.717
Unit
0.337
0.529
Modifier
0.683
1.000
Overall
0.624
0.720
Table 4: Fleiss’ κ over the five-reviewer overlap subset (10 shared documents), on all 78 shared pairs and on the 13 pairs no reviewer flagged. Agreement is higher on the unflagged pairs, a descriptive result given their small number; event type is the schema’s hardest field.