We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.
Figures & tables
Figure 1: Study flow (left: the three stages of each session) and the InterView-C corpus (right), whose three released parts share one session clock. VR interaction data : records per table (fingers tracked in 515,995 / 544,280 frames; SurveyItem: 55 interview, 26 protocol and 50 experience items); thin arrows are record links, the dashed arrow marks automatic transcription and thick arrows mark derivation. The reference transcripts are post-edited from the automatic transcript against the withheld audio and the negation annotations are sampled from it. Badges give the creating stage; 4 (post-editing) and 5 (annotation) continue the study stages.
Figure 2: The virtual interview environment with the interviewee’s 3D gaze points, colored by dwell time from blue (shorter) to red (longer).
Interviewee
Interviewer
Total
Speakers
27
3
30
Tracked time (h)
9.9
10.5
20.4
Tracking frames
654,682
695,890
1,350,572
Word tokens, automatic ( Word )
14,772
38,703
53,475
Word tokens, post-edited reference ( ReferenceWord ) ‡
15,573
40,191
55,764
Words per interview, automatic, median (range)
420 (115–1,564)
1,315 (1,182–2,083)
–
Table 1: General statistics about the data InterView-C contains.
System
WER [95% CI]
CER
SemD
IR
IE
Verb.
Del.
Ins.
Cue
Num
Hall./h
CrisperWhisper †
7.4 [6.3–8.6]
5.8
1.72
6.2
10.5
9.1
4.0
1.6
94.2
86.3
2.2
Own VAD, batched decoding
WhisperX (large-v3)
8.0 [7.4–8.7]
5.5
1.60
6.6
12.0
10.0
4.0
1.4
93.2
86.8
1.0
Shared Silero VAD segments
Qwen3-ASR 1.7B
10.4 [9.4–11.6]
6.1
1.90
8.6
15.4
12.4
3.3
2.0
90.8
84.5
1.1
Voxtral Mini 3B
10.8 [9.9–11.8]
6.7
2.16
8.5
17.3
12.6
3.5
2.1
90.4
77.6
1.3
Table 2: ASR results on all 54 recordings (%, SemDist ×100 ; lower is better except for recall). WER : WER without filled pauses; CER : character error rate; SemD : mean semantic distance per reference line; IR / IE : WER on interviewer/interviewee recordings; Verb. : WER with filled pauses kept; Del. / Ins. : deletions/insertions relative to the reference length; Cue / Num : recall of negation cues and number words (Section 5.1 , Results); Hall./h : inserted passages of at least five words per hour of audio. † The reference was post-edited from this transcript and is shown for reference only. ‡ Without conditioning on previously decoded text.
Train
Val
Test
Interviews
21
3
3
Sentences
1,106
158
158
Tokens
9,464
1,531
1,750
Sentences w/ negation
225
29
27
Negation instances
267
38
37
Cue tokens
278
40
38
Table 3: Consolidated negation annotation per split.
Training data
Cue F1
Scope F1
BioScope (abstracts)
0.58
0.49
BioScope (full)
0.26
0.34
ConanDoyle-neg
0.81
0.53
DT-Neg
0.11
0.61
SFU Review
0.82
0.34
SOCC
0.79
0.38
Table 4: Word-level F1 (%) on the test split for our and released D-Neg models (mean ± std. over three seeds). More detailed results are reported in Appendix C.3 .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Rec. (IR/IE)
Edit
Model
Anch.
Low
T1
12 (6/6)
10.3
9.7
5.4
3.9
T2
12 (6/6)
10.2
10.9
5.8
4.5
T3
10 (5/5)
12.2
9.8
5.7
3.6
T4
10 (4/6)
7.5
13.1
11.7
4.7
T5
10 (6/4)
4.7
9.1
9.7
2.6
Appendix
Table 5: Reference quality per annotator. Rec. : recordings (interviewer/interviewee); Edit : words changed relative to the automatic transcript (%); Model : WER of the five VAD-based systems (%); Anch. : tokens reproduced only by the automatic transcript, per 1,000 reference tokens; Low : lines recognized by at least one system with alignment confidence below 0.05 (%).
Cohen’s κ
Task
A1–A2
A1–A3
A2–A3
α
Exact (%)
Cue
0.90
0.85
0.85
0.87
93.8–96.3
Scope †
0.79
0.79
0.84
0.81
59.3–75.3
Appendix
Table 6: Token-level inter-annotator agreement. Exact gives the range over annotator pairs of the percentage of shared sentences (cue) or matched negation instances (scope) with identical annotations.
Cue
Scope
Training data
P
R
F1
EM
P
R
F1
EM
BioScope (abstracts)
0.94
0.42
0.58
0.89
0.60
0.42
0.49
0.19
BioScope (full)
0.75
0.16
0.26
0.85
0.51
0.25
0.34
0.08
ConanDoyle-neg
0.93
0.71
0.81
0.92
0.47
0.61
0.53
0.24
DT-Neg
0.06
0.63
0.11
0.31
0.69
0.54
0.61
0.27
SFU Review
0.93
0.74
0.82
0.92
0.50
0.26
0.34
0.11
Appendix
Table 7: Negation results on our test split for GAT models: word-level precision (P), recall (R) and F1 for the positive class, together with exact match (EM), measured as the proportion of exactly correct sentences for cue detection and exactly correct scopes for scope detection. D-Neg models are applied as released; our results are averaged over three seeds (mean ± standard deviation). Scope detection uses gold cues.
Model
Len.
Exact
Short
Long
Other
Ours GAT
3.6
42
32
14
12
D-NEG GAT, DT-Neg
4.0
27
27
24
22
Appendix
Table 8: Scope error types on the test split. Len. denotes mean predicted scope length in tokens; the mean reference scope length is 5.1 tokens.
Section
Variable
Content (English gloss)
Satisfaction, partnership, family
t514001
Life satisfaction (0–10)
t731112
Steady partnership
t731111
Living with partner
t731110
Marital status
tp1
Feelings about being single
tp2
Disagreements in the relationship
Appendix
Table 9: Main questionnaire read out in VR.
Figure 3: The four avatars used during the interview. From A on the left to D on the right.
Layer
Table
Content
Rate
Records
Session
Experiment
interview metadata
–
27
Player
speaker per interview, connection sessions
–
54
Tracking
Eye
binocular gaze poses, validity, confidence
18.4 Hz
1,350,572
Head , Body
6-DoF head and rig pose
18.4 Hz
1,350,572 each
Facial
63 blendshape weights
18.4 Hz
1,350,572
Left/RightHand
6-DoF wrist pose
18.4 Hz
1,350,572 each
Appendix
Table 10: Released data layers of the 27 interviews. Finger tables hold a row for every frame; the counts give the frames in which the hand was tracked.