A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. Re:Cognize evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set retrieval to a cast the model must build and grow itself. The surprise is where models fail. Recognising is close to solved: one reference image per character already ranks as well as a gallery built in advance. Knowing what to believe is not: a model that adds its own matches makes its cast worse, while the same growth with correct labels would gain over twenty points of top-1 accuracy. The bottleneck is acceptance, not vision, and one comparison decides it: an addition pays exactly when it is right more often than the cast already was on the queries it takes over. The comparison has nothing to fit, and measured on half of a new corpus it calls the other half correctly. ReCast puts it to work with nothing fitted on data: a cast sheet of one running average per character, grown only where the page itself vouches for a crop. It recovers a third to two thirds of what perfect labels would, depending on whether the cast starts from random examples or from first appearances. Re:Cognize measures whether a model can read along; ReCast is a cast that does. Our claims are on identity maintenance, recognising characters already met; the emergence of new ones is measured as a diagnostic under a fixed reference rule, and we propose no method for it.
Figures & tables
Figure 2 : The four Re:Cognize protocols. All four answer one stream of query crops in reading order and differ in the gallery and whether it may change. k is the number of seed crops per character, τnov the novelty threshold of P3 and Bmax the cap on crops P4 adds per character; frames are coloured by character and dashed where the model added the crop. P3 is a diagnostic of emergence, and its fixed threshold is a reference rule rather than a clustering algorithm we propose.
Method
Closed Set
Open Set
Sequential
Online Gallery
Cross Corpus
Memory
TransReID [ 13 ]
✓
–
–
–
✓
–
OSNet [ 55 ]
✓
–
–
–
✓
–
Instruct-ReID [ 14 ]
✓
–
–
–
✓
–
Zhang et al. [ 52 ]
✓
–
✓
–
–
–
Zhang & Chu [ 51 ]
✓
–
–
–
–
–
Soykan et al. [ 37 ]
✓
–
–
–
–
–
Table 1 : Method positioning. Prior comic and person Re-ID against Re:Cognize across the six capabilities its protocols exercise; Memory is the memory block baseline Re:Cognize evaluates. Cross corpus means evaluation on a second annotated corpus of the same medium.
P1: Closed-Set
P2: Seeded k=1 , mAP
P4: Seq-R k=1 , id. R-1
P4: Seq-T k=1 , id. R-1
Backbone
Config
mAP
R-1
Seq-R
Seq-T
Static
Pred.
Oracle
Static
Pred.
Oracle
Chance
Random ranking
33.1
30.1
33.9
33.9
12.9
–
–
12.9
–
–
TransReID
Finetuned
37.4
40.6
38.6
33.2
17.4
15.0
41.7
12.8
12.6
42.2
FT + Mem + LoRA †
38.1
39.9
39.0
34.2
18.1
16.1
42.3
13.9
13.3
42.4
MagiV2
Finetuned
51.2
57.2
54.4
52.9
36.8
36.3
59.1
36.6
37.9
59.0
FT + Mem †
51.7
56.3
54.6
51.3
37.2
35.9
59.8
35.3
36.9
60.4
Table 2 : Cross-protocol summary on POPCharacters , 8 held-out series, k=1 seed per character, mean over three training runs. Chance is a random ranking of the same galleries. Per backbone the top row, Finetuned , is the frozen backbone with a trained BNNeck and no LoRA, and the bottom row ( † ) the best memory-block configuration by P1 mAP. Seeds are random (Seq-R) or first appearances (Seq-T). P4 is identity Rank-1 at Bmax=50 under three policies: static , the P2 gallery; predicted , adding each query under its top-1 match; oracle , adding it under its true character.
Figure 3 : What the protocols measure, and where the headroom is. (a) One seed per character recovers the closed-set mAP but only part of its Rank-1. (b) Growth by the model’s own top-1 matches stays below a static gallery, while correct labels would add over twenty points (labels: oracle minus static). (a, b): memory-block configuration, one random seed per character, three training runs. (c) Encoder-side adaptation helps most on the backbones weakest in this domain. (d) Each point is one change to the gallery on one backbone and corpus: adding under the page constraint (Section 6 ), adding by top-1 match, or restricting the candidate characters. The stronger the gallery already is, the less any change adds, and Section 5 predicts which points fall below zero.
Backbone
Rule
#clusters
Purity
Hung. Acc
ARI
NMI
TransReID
fixed τnov=0.55
95.8
60.2
15.1
2.0
21.8
MagiV2
fixed τnov=0.55
38.9
69.6
46.6
24.0
34.1
TransReID
fixed τnov=0.80 (loose)
431.1
93.0
4.8
0.2
37.0
MagiV2
fixed τnov=0.80 (loose)
190.9
84.3
24.6
9.8
38.5
Table 3 : P3 identity emergence on full per-series streams ( POPCharacters test series, finetuned backbones from the first training run, macro over 8 series). The fixed rule at τnov=0.55 is the reference instantiation. The loose rows show why Purity cannot be read alone: raising the threshold multiplies clusters four- to fivefold and Purity rises with them, while Hungarian accuracy and ARI collapse. Alternative decision rules are in Appendix Table 15 and the comparison with the memory block in Appendix Table 16 ; neither is a baseline of the framework.
Corpus
Backbone
a
c
peff
a+
c(peff−a+)
measured
POPCharacters
TransReID
23.82
34.96%
31.532
23.791
+2.71
+2.71
POPCharacters
InstructReID
25.38
33.15%
30.851
23.437
+2.46
+2.46
POPCharacters
ReID5o
28.68
35.14%
36.682
28.106
+3.01
+3.01
POPCharacters
MagiV3
32.88
35.67%
39.119
31.918
+2.57
+2.57
Manga109
MagiV3
34.03
43.35%
39.513
33.519
+2.60
+2.60
POPCharacters
MagiV2
44.35
34.67%
47.832
47.960
-0.04
-0.04
Table 4 : The commit condition reproduces every measured change. Commitment under the page constraint (Section 6 ) on the bag of k=5 random seeds per character, without the cast sheet, over three seed draws. a is the static gallery’s identity Rank-1 and c , peff and a+ are the terms of Equation 1 , in percent. Weighting peff and a+ by each record’s capture rate, ∑icixi/∑ici , makes c(peff−a+) exactly the mean measured change, so the last two columns agree.
Figure 4 : Re:Cast on Re:Zero crops from Re:Verse [ 5 ] , with Rom as identity A. Top: the three changes of Section 6 . A page names a character when one of its crops is already committed to it, and commitment then adds that crop’s page-group sibling. Bottom: over six pages the cast sheet is updated only on the four that name Rom; elsewhere the rule abstains.
Backbone
Static
Cast sheet
+ commitment
+ expansion
Oracle
POPCharacters , 8 test series, Seq-R, k=1
TransReID
17.17
+0.00
+4.71
—
MagiV2
35.95
+0.00
+0.62
—
MagiV3
21.79
+0.00
+4.84
—
InstructReID
16.83
+0.00
+6.84
—
ReID5o
21.03
+0.00
+5.66
—
Table 5 : Re:Cast , one change at a time : P4 identity Rank-1 with random seeds at Bmax=50 , on the 8 held-out POPCharacters series and, zero-shot, on 27 held-out Manga109 volumes. Each column after Static adds one change and gives the cumulative gain over Static , except + expansion , which is scored against its own static reference with the expanded crops removed from the queries; Oracle adds every query under its true label. Cells are one training run and three seed draws, and bold marks p<0.05 on a paired t -test over series. At k=1 the first two changes coincide, the oracle is in Table 2 , and seed expansion was not run on Manga109.
Figure 5 : Binding, and pricing the binder. Top, on Re:Verse’s Re:Zero annotations [ 5 ] : (a) a page’s crops are grouped above one similarity threshold, (b) page groups merge in reading order into the best earlier match above a second, or start new ones, and (c) a group’s first seed names it. Bottom: any frozen encoder can bind, with threshold τ ; its crops are committed when Equation 1 predicts a positive Δ , with peff the additions’ precision on the queries they capture and a+ the static gallery’s accuracy there, and τ is chosen by that prediction, never the measured gain.
Figure 6 : Where the supervision sits. Labelled share of the last twenty crops read, by quarter of the stream, over the 8 held-out POPCharacters series at k=5 . Both seedings place the same number of labels in the volume and distribute them differently.
Shared binder
Self
Controls
Backbone
a
POPCharacters
Manga109
POPCharacters
Oracle
Shuffled
Seq-R
TransReID
20.7
+14.32
+10.32
+4.96
+20.94
-5.56
+3.93
InstructReID
20.4
+15.45
+10.34
+4.60
+22.46
-5.93
+3.86
ReID5o
20.9
+16.94
+10.35
+2.52
+24.63
-6.07
+2.70
MagiV3
26.5
+16.45
+8.00
+1.73
+23.56
-10.71
+1.89
MagiV2
37.4
+12.51
-7.44
+8.46
+21.25
-20.44
-4.92
Table 6 : Two-stage binding under chronological seeding : change in P4 identity Rank-1 over the static gallery ( a on POPCharacters ) at k=5 , averaged over series and seed draws; bold is a gain at p<0.05 paired over series. Columns are POPCharacters unless marked Manga109 (27 volumes, zero-shot). Shared is the binder of Section 7 for every backbone and Self each backbone’s own. Oracle is the same representation’s ceiling ( +20.2 to +26.4 on Manga109), Shuffled permutes the added identities ( −8.1 to −43.5 on Manga109), and Seq-R the same binding with random seeds.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : The memory block. A frozen backbone produces L2-normalised features via BNNeck. Working Memory (cross-attention over recent observations) and Episodic Memory (cross-attention over identity prototypes) produce residuals δtwm,δtem , combined by Gated Fusion through a small-init projection. Optional LoRA adapts the last L backbone attention layers. Trainable: MLPs and attention projections; frozen: backbone B ; runtime state: prototypes Pc and FIFO buffers Bc .
Configuration
Backbone
BNNeck
Memory
LoRA
Pretrained
frozen
–
–
–
Finetuned
frozen
trained
–
–
Finetuned + Memory
frozen
trained
trained
–
Finetuned + LoRA
adapted
trained
–
trained
Finetuned + Memory + LoRA
adapted
trained
trained
trained
Appendix
Table 7 : Five configurations analysed per backbone: BNNeck fine-tuning, the memory block, and LoRA toggled along three axes.
Backbone
Architecture
Pre-training
Input
d
TransReID [ 13 ]
ViT-B/16
ImageNet + Re-ID
256×128
768
MagiV2
ViT-B
Manga embeddings
224×224
768
MagiV3
Florence-2 [ 50 ]
Manga comprehension
384×384
1024
InstructReID [ 14 ]
ViT-B/16
Multi-modal Re-ID
256×128
768
ReID5o
CLIP ViT-B/16 [ 29 ]
CLIP + Re-ID
384×128
512
Appendix
Table 8 : Backbone architectures evaluated by Re:Cognize . d is the native feature width, at which each backbone’s BNNeck and memory block operate.
Figure 8 : Two-pass open-set inference schematic. Pass 1 produces a coarse identity hypothesis. Pass 2 uses the hypothesis to route Working Memory and emits the final descriptor. The FIFO buffer and prototype bank update only after Pass 2.
Latency (ms/crop)
Parameters
Backbone
Backbone
+ Memory
Block alone
Backbone
Memory block
TransReID
5.8
17.8
4.9
85.6M
10.1M
MagiV2
6.8
20.2
5.0
85.8M
10.1M
MagiV3
24.2
59.6
4.8
360.7M
17.9M
InstructReID
7.2
20.3
4.9
85.8M
10.1M
ReID5o
13.4
34.3
4.8
79.1M
4.5M
Appendix
Table 9 : Computational overhead on one NVIDIA RTX A6000 at batch size one in FP32, median over 300 crops of one test series. + Memory is the two-pass inference as implemented, which runs the backbone once per pass, and block alone is the same two passes over cached backbone features. Backbone parameters count what the crop forward touches, since MagiV3 and ReID5o ship modules it never calls. The block’s buffers and prototypes are filled at test time and are not parameters.
Backbone
Configuration
mAP
R-1
R-5
R-10
TransReID
Pretrained
37.1
39.7
79.0
89.2
Finetuned
37.4
40.6
78.9
89.0
Finetuned + Memory
37.5
39.5
77.9
88.5
Finetuned + LoRA
37.7
41.1
78.9
89.0
Finetuned + Memory + LoRA
38.1
39.9
78.1
88.4
MagiV2
Pretrained
50.8
57.1
85.1
91.0
Appendix
Table 10 : Full P1 closed-set retrieval on the 8 held-out POPCharacters series (Appendix A.16 ). Best per backbone in bold , second-best underlined . Finetuned trains a BNNeck on a frozen backbone; only LoRA rows adapt backbone weights. Every row is the mean over three training runs.
Finetuned
FT + Memory
FT + Mem + LoRA
Manga
#C
mAP
R-1
mAP
R-1
mAP
R-1
Bakuman
7
62.9
70.2
63.5
68.7
63.5
70.2
Demon Slayer Kimetsu No Yaiba
11
47.2
52.7
48.5
52.5
48.0
51.7
Dr Stone
4
63.8
65.6
63.8
65.0
63.8
66.0
Hunter X Hunter
12
40.7
48.8
41.2
46.5
41.6
45.4
Kagurabachi
6
41.2
51.0
42.0
52.4
41.7
52.3
Appendix
Table 11 : Per-manga P1 results for MagiV2 across three configurations, mean over three training runs and five seed draws. #C is the character count. Best per series in bold .
Manga109 , 27 volumes
Re:Verse , 1 series, pretrained
P1 mAP
P4 id. R-1
P1
Backbone
Pretrained
Finetuned
+ Mem.
+ Mem.
mAP
R-1
TransReID
28.4
29.1
30.3
10.7
34.5
46.0
MagiV2
65.3
65.9
67.4
47.1
84.2
91.0
MagiV3
37.4
39.4
41.5
18.4
48.2
73.0
InstructReID
28.6
30.4
35.5
14.3
30.7
50.9
Appendix
Table 12: Transfer beyond POPCharacters , headline rows. Manga109 is 27 held-out volumes. Pretrained is the released backbone with no POPCharacters training at all; Finetuned and + Mem. are POPCharacters -trained checkpoints scored zero-shot here from one training run, the latter the memory-block configuration with the highest P1 mAP on this corpus, so the first gap is what BNNeck fine-tuning on another corpus is worth and the second is the memory block. P4 is identity Rank-1 at k=1 under random seeding at Bmax=50 , for that same configuration. Re:Verse is one series (Re:Zero) and its two columns are the released backbone with no POPCharacters training. Manga109 cells are the mean over three seed draws, Re:Verse cells over five. The full tables are Table 13 and Table 14 in the appendix. Three vision-language models on the same Re:Verse benchmark reach 1.11, 0.00 and 0.85 percent character-identification accuracy [ 5 ] .
P1: Closed-Set
P2-R k=1
Seq-R k=1 , id. R-1
Backbone
Config
mAP
R-1
mAP
P2
P4
TransReID
Pretrained
28.4
35.9
26.6
12.2
10.5
Finetuned
29.1
36.5
26.9
12.5
10.2
FT + Mem + LoRA †
30.3
37.1
27.9
13.3
10.7
MagiV2
Pretrained
65.3
76.6
63.8
50.4
45.7
Finetuned
65.9
77.4
63.6
50.5
45.7
Appendix
Table 13 : Cross-corpus transfer to Manga109 , 27 held-out volumes (Appendix A.16 ). Pretrained is the released backbone with no POPCharacters training at all. The two rows under it are POPCharacters -trained checkpoints scored zero-shot here, the no-memory baseline and then the best memory-block configuration by P1 mAP ( † ), so the first gap in each block is what BNNeck fine-tuning on another corpus is worth. P4 is identity Rank-1, with P2 given at Rank-1 beside it.
Method
Char-ID Acc (%)
mAP
Rank-1
VLMs
Qwen2.5-VL-3B [ 5 ]
1.11
–
–
InternVL3-14B [ 5 ]
0.00
–
–
Ovis2-8B [ 5 ]
0.85
–
–
Re-ID
TransReID (pre)
–
34.5
46.0
InstructReID (pre)
–
30.7
50.9
ReID5o (pre)
–
36.2
60.8
Appendix
Table 14 : Re:Verse benchmark. Character identification on Re:Zero. The VLM rows carry the character-identification accuracy published with the benchmark [ 5 ] . Every Re-ID row is our own harness (Appendix A.16 ), the mean over five seed draws, so the pretrained rows and the trained rows are directly comparable. The memory block is worth +0.1 mAP here and costs 1.6 Rank-1, on a single series.
Backbone
Rule
#clusters
Purity
Hung. Acc
ARI
NMI
TransReID
fixed τnov=0.55
95.8
60.2
15.1
2.0
21.8
variance-adaptive
111.2
62.3
13.8
1.6
22.9
density-aware
180.6
69.5
8.6
0.9
27.8
cohesion-relative
144.0
65.9
11.4
1.4
25.8
graph community detection
21.8
51.0
17.5
2.6
13.0
MagiV2
fixed τnov=0.55
38.9
69.6
46.6
24.0
34.1
Appendix
Table 15 : P3 decision rules on full per-series streams (finetuned TransReID and MagiV2 from the first training run, POPCharacters test series, macro over 8 series). Best ARI per backbone in bold . No alternative improves on the fixed rule at a comparable cluster count: on MagiV2 every rule is below it, and on TransReID graph community detection edges past it only by collapsing to a quarter of the clusters. These rules are reference instantiations of the decision, not baselines of the framework, and no claim in the paper depends on their ranking.
Finetuned
Trained with the memory block
Backbone
#cl.
Pur.
Hung.
ARI
NMI
#cl.
Pur.
Hung.
ARI
NMI
TransReID
95.8
60.2
15.1
2.0
21.8
90.1
59.3
16.3
2.2
21.3
MagiV2
38.9
69.6
46.6
24.0
34.1
39.9
69.9
47.4
24.5
34.5
MagiV3
79.9
65.2
20.7
5.9
26.0
69.0
64.1
23.0
6.7
24.6
InstructReID
278.8
81.0
10.5
2.1
33.3
260.6
78.8
11.0
2.2
32.2
ReID5o
213.8
76.0
10.6
1.8
31.2
199.1
74.3
10.5
1.8
30.4
Appendix
Table 16 : P3 on full per-series streams, all five backbones , fixed rule at τnov=0.55 , macro over the 8 held-out POPCharacters series against 8.8 identities per series, first training run. Every crop of every series is streamed in reading order. P3 has no gallery to initialise memory from, so the block is bypassed and the right-hand columns score the BNNeck trained alongside it.
Seq-R (random seeding)
Seq-T (chronological seeding)
Backbone
Config
Static
Pred.
Oracle
Wrong-app.
Contam.
Static
Pred.
Oracle
Wrong-app.
Contam.
TransReID
Finetuned
17.4
15.0
41.7
85.0
84.2
12.8
12.6
42.2
87.4
83.3
FT + Mem
17.2
15.7
41.4
84.3
82.8
13.2
13.6
41.8
86.4
84.0
MagiV2
Finetuned
36.8
36.3
59.1
63.7
67.8
36.6
37.9
59.0
62.1
69.8
FT + Mem
37.2
35.9
59.8
64.1
68.7
35.3
36.9
60.4
63.1
68.3
MagiV3
Finetuned
22.7
19.9
49.9
80.1
80.0
17.4
17.7
50.2
82.3
82.0
Appendix
Table 17 : P4 under its three update policies at k=1 and Bmax=50 on the 8 held-out POPCharacters test series, identity Rank-1, mean over three training runs and five seed draws. Static is the unchanged P2 gallery, predicted adds each query under its top-1 match (the protocol’s rule) and oracle adds each query under its true character. Wrong-app. is the fraction of added entries that carry the wrong character and contam. the mislabelled fraction of the grown gallery at stream end, both under the predicted policy. Finetuned trains a BNNeck head on the frozen backbone, with no LoRA; FT + Mem adds the memory block to it. The three policies share one seed draw per run.
Figure 9 : P4 over the stream at k=1 and Bmax=50 with the memory block, mean over three training runs, five seed draws and the 8 test series. Left and centre: identity Rank-1 in each quarter of the stream, pooled over the five backbones, for the static gallery, growth by the model’s own top-1 (predicted) and growth by the true label (oracle). Right: the share of the grown gallery that is mislabelled at the end of the stream under predicted growth.
Configuration
0 (=P2)
5
10
25
50
100
unbounded
TransReID, chronological
13.2
13.2
14.2
13.0
13.6
14.6
14.9
TransReID, random
17.2
16.1
16.4
15.9
15.7
15.4
15.4
MagiV2, chronological
35.3
35.8
33.4
35.5
36.9
37.5
37.4
MagiV2, random
37.2
35.4
34.9
35.9
35.9
36.1
36.1
Appendix
Table 18 : Sensitivity of P4 to the buffer cap Bmax (identity Rank-1 at k=1 on the 8 held-out POPCharacters series, mean over three training runs). Bmax=0 is the static P2 gallery. No cap turns growth by predicted labels into a gain over that gallery: under random seeding every cap is below it, and under chronological seeding MagiV2 recovers about two points as the cap loosens without reaching the oracle’s range. The cap is therefore not the free parameter that decides whether growth pays. Acceptance is (Section 6 ).
TransReID FT
TransReID +Mem
MagiV2 FT
MagiV2 +Mem
Box condition
IoU
P1
P2
P1
P2
P1
P2
P1
P2
clean box
1.00
37.4
38.6
37.5
38.4
51.2
54.4
51.7
54.5
shift 10%
0.83
+0.1
+0.2
+0.2
-0.1
-0.0
+0.4
+0.2
+0.0
shift 20%
0.69
-0.2
-0.3
-0.3
-0.4
-0.9
-1.4
-0.5
-1.4
shift 30%
0.58
-0.7
-0.4
-0.8
-0.2
-2.4
-2.7
-2.1
-2.7
tight 0.7×
0.49
+0.6
+1.5
+0.6
+1.6
-0.6
-0.2
-0.5
-0.7
Appendix
Table 19 : Bounding-box displacement. P1 and P2 Seq-R at k=1 , three seed draws, macro over the 8 test series. The first row is absolute mAP and every later row is the change from it in the same column. IoU is the mean overlap of the displaced box with the true one.
TransReID FT
TransReID +Mem
MagiV2 FT
MagiV2 +Mem
Pixel condition
P1
P2
P1
P2
P1
P2
P1
P2
clean pixels
37.4
38.6
37.5
38.4
51.2
54.4
51.7
54.5
jitter 10%
+0.0
+0.3
+0.1
-0.3
-0.1
+0.1
-0.0
+0.0
jitter 20%
-0.2
+0.1
-0.2
-0.5
-0.5
-0.3
-0.5
-0.4
blur σ=2
+0.2
+0.6
+0.1
+0.3
-1.2
-1.2
-1.2
-1.3
blur σ=4
+0.2
+0.4
-0.1
-0.7
-4.9
-5.4
-5.0
-6.3
Appendix
Table 20 : Pixel corruption (same setting as Table 19 ). Blur and occlusion, not box placement, are what separate the two backbones.
Backbone
Configuration
P1 mAP
P1 R-1
P2-R@1 mAP
Δ P1 mAP vs full
MagiV2
Finetuned (no memory)
51.20
57.17
54.44
−0.54∗
Full memory block
51.75
56.29
54.58
–
− working memory
51.15
57.28
54.53
−0.59∗
− episodic memory
52.00
56.26
54.62
+0.26
− ID-drop
51.91
56.25
54.62
+0.16
− memory-consistency loss
51.92
56.60
54.64
+0.18∗
Appendix
Table 21 : Per-component ablation of the memory block on the two backbones where it earns anything, macro over the 8 held-out POPCharacters test series over three training runs. Each row removes one component from the full memory block; no row carries LoRA. Δ is against the full memory block, paired over the 24 (series, training run) cells, and ∗ marks p<0.05 under a two-sided exact sign test on those pairs. Working memory carries the whole effect: removing it returns P1 mAP to the finetuned baseline on both backbones, while removing episodic memory, ID-drop or the memory-consistency loss moves P1 mAP upward .
Figure 10 : P3 novelty-threshold sensitivity on full per-series streams: predicted clusters per series (log scale; dashed, the 8.8 true identities), Purity, NMI and ARI against τnov for the five finetuned backbones, first training run, macro over the 8 test series. The dotted line is the reference τnov=0.55 . Purity and NMI rise as the stream fragments, ARI falls or peaks early, and MagiV2 leads on ARI at every threshold.
Dataset
Split
#Series
#Chars
#Crops
Avg C/Ch
POPCharacters
Train
13
198
7,668
38.7
Development
2
10
873
87.3
Test
8
70
4,058
58.0
Manga109
Test
27
784
29,315
37.4
Re:Verse
Test
1
12
1,825
152.1
Appendix
Table 22 : Dataset statistics. Splits are series-disjoint, and volume-disjoint on Manga109. Avg C/Ch is the average crops per character. No model is trained on Manga109 or Re:Verse: every number on them is zero-shot.
Figure 11 : Crops per series in POPCharacters , coloured by split: 13 training series, 2 development series and the 8 test series. Character counts above bars.
Figure 12 : Crop count distribution per series. Series with large casts (left) exhibit heavier tails in per-character crop counts.
A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark that scores character voice directly instead of speaker voice, built from 85 human-audited identities in a dubbed anime corpus. It poses two challenging conditions: one that swaps the performer under a fixed character, and one that fixes the performer under changing characters. Listening studies with 78 participants provide human reference scores for cross-performer verification and same-actor discrimination. Speaker-verification baselines show increased errors under these character-specific conditions. We then train KyaraEmbed, a compact encoder using multilingual character supervision, same-actor negatives, and a language-alignment term. The model achieves the best performance on all character-specific conditions in the main comparison. We release the benchmark, protocol, and encoder publicly.
Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at. We act on the complementary relation, equivariance: edit a figure's data and the correct answer must change by a computable amount. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap. We instantiate the theory as the Equivariance-Consistency Score, a label-free, training-free detector, and release REND-EQUIV, pairing matched invariance and equivariance sets over identical data. The predicted ordering holds across three models and a hand-labeled population immune to the one circularity in how it is selected; a second invariance-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample. The same characterization explains a reported inversion of this ordering in the classifier metamorphic-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone.
Rasul Khanbayov, Hasan Kurban
College of Science and Engineering, Hamad Bin Khalifa University, Doha, Qatar
Practical text-recognition pipelines for historical documents typically decompose layout analysis into line detection followed by a separate reading-order step, with the latter most often handled by a hand-coded geometric heuristic that struggles with marginalia, multiple columns, tables, and source-specific editorial conventions. This article introduces Orli (Ordered Regression of Lines), an end-to-end model that casts both sub-tasks as a single image-to-sequence problem: from a page image, Orli autoregressively generates text-line baselines directly in reading order. Baselines are represented in a chord-frame parameterization that anchors a line's position, orientation, and extent while encoding local geometry through perpendicular offsets; an iterative refinement head and a local visual refiner produce the final curve. Trained on a heterogeneous corpus of 196,691 pages spanning ten writing systems, Orli marginally exceeds the previously reported state of the art for cBAD line detection without dataset-specific training, reaches near perfect coverage and ordering on multiple reading-order benchmarks zero-shot, and adapts to more specialized out-of-domain layouts with limited fine-tuning. The method's source code and model weights are available under an open license at https://github.com/mittagessen/orli.