We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.
Figures & tables
Figure 1: Study flow (left: the three stages of each session) and the InterView-C corpus (right), whose three released parts share one session clock. VR interaction data : records per table (fingers tracked in 515,995 / 544,280 frames; SurveyItem: 55 interview, 26 protocol and 50 experience items); thin arrows are record links, the dashed arrow marks automatic transcription and thick arrows mark derivation. The reference transcripts are post-edited from the automatic transcript against the withheld audio and the negation annotations are sampled from it. Badges give the creating stage; 4 (post-editing) and 5 (annotation) continue the study stages.
Figure 2: The virtual interview environment with the interviewee’s 3D gaze points, colored by dwell time from blue (shorter) to red (longer).
Interviewee
Interviewer
Total
Speakers
27
3
30
Tracked time (h)
9.9
10.5
20.4
Tracking frames
654,682
695,890
1,350,572
Word tokens, automatic ( Word )
14,772
38,703
53,475
Word tokens, post-edited reference ( ReferenceWord ) ‡
15,573
40,191
55,764
Words per interview, automatic, median (range)
420 (115–1,564)
1,315 (1,182–2,083)
–
Table 1: General statistics about the data InterView-C contains.
System
WER [95% CI]
CER
SemD
IR
IE
Verb.
Del.
Ins.
Cue
Num
Hall./h
CrisperWhisper †
7.4 [6.3–8.6]
5.8
1.72
6.2
10.5
9.1
4.0
1.6
94.2
86.3
2.2
Own VAD, batched decoding
WhisperX (large-v3)
8.0 [7.4–8.7]
5.5
1.60
6.6
12.0
10.0
4.0
1.4
93.2
86.8
1.0
Shared Silero VAD segments
Qwen3-ASR 1.7B
10.4 [9.4–11.6]
6.1
1.90
8.6
15.4
12.4
3.3
2.0
90.8
84.5
1.1
Voxtral Mini 3B
10.8 [9.9–11.8]
6.7
2.16
8.5
17.3
12.6
3.5
2.1
90.4
77.6
1.3
Table 2: ASR results on all 54 recordings (%, SemDist ×100 ; lower is better except for recall). WER : WER without filled pauses; CER : character error rate; SemD : mean semantic distance per reference line; IR / IE : WER on interviewer/interviewee recordings; Verb. : WER with filled pauses kept; Del. / Ins. : deletions/insertions relative to the reference length; Cue / Num : recall of negation cues and number words (Section 5.1 , Results); Hall./h : inserted passages of at least five words per hour of audio. † The reference was post-edited from this transcript and is shown for reference only. ‡ Without conditioning on previously decoded text.
Train
Val
Test
Interviews
21
3
3
Sentences
1,106
158
158
Tokens
9,464
1,531
1,750
Sentences w/ negation
225
29
27
Negation instances
267
38
37
Cue tokens
278
40
38
Table 3: Consolidated negation annotation per split.
Training data
Cue F1
Scope F1
BioScope (abstracts)
0.58
0.49
BioScope (full)
0.26
0.34
ConanDoyle-neg
0.81
0.53
DT-Neg
0.11
0.61
SFU Review
0.82
0.34
SOCC
0.79
0.38
Table 4: Word-level F1 (%) on the test split for our and released D-Neg models (mean ± std. over three seeds). More detailed results are reported in Appendix C.3 .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Rec. (IR/IE)
Edit
Model
Anch.
Low
T1
12 (6/6)
10.3
9.7
5.4
3.9
T2
12 (6/6)
10.2
10.9
5.8
4.5
T3
10 (5/5)
12.2
9.8
5.7
3.6
T4
10 (4/6)
7.5
13.1
11.7
4.7
T5
10 (6/4)
4.7
9.1
9.7
2.6
Appendix
Table 5: Reference quality per annotator. Rec. : recordings (interviewer/interviewee); Edit : words changed relative to the automatic transcript (%); Model : WER of the five VAD-based systems (%); Anch. : tokens reproduced only by the automatic transcript, per 1,000 reference tokens; Low : lines recognized by at least one system with alignment confidence below 0.05 (%).
Cohen’s κ
Task
A1–A2
A1–A3
A2–A3
α
Exact (%)
Cue
0.90
0.85
0.85
0.87
93.8–96.3
Scope †
0.79
0.79
0.84
0.81
59.3–75.3
Appendix
Table 6: Token-level inter-annotator agreement. Exact gives the range over annotator pairs of the percentage of shared sentences (cue) or matched negation instances (scope) with identical annotations.
Cue
Scope
Training data
P
R
F1
EM
P
R
F1
EM
BioScope (abstracts)
0.94
0.42
0.58
0.89
0.60
0.42
0.49
0.19
BioScope (full)
0.75
0.16
0.26
0.85
0.51
0.25
0.34
0.08
ConanDoyle-neg
0.93
0.71
0.81
0.92
0.47
0.61
0.53
0.24
DT-Neg
0.06
0.63
0.11
0.31
0.69
0.54
0.61
0.27
SFU Review
0.93
0.74
0.82
0.92
0.50
0.26
0.34
0.11
Appendix
Table 7: Negation results on our test split for GAT models: word-level precision (P), recall (R) and F1 for the positive class, together with exact match (EM), measured as the proportion of exactly correct sentences for cue detection and exactly correct scopes for scope detection. D-Neg models are applied as released; our results are averaged over three seeds (mean ± standard deviation). Scope detection uses gold cues.
Model
Len.
Exact
Short
Long
Other
Ours GAT
3.6
42
32
14
12
D-NEG GAT, DT-Neg
4.0
27
27
24
22
Appendix
Table 8: Scope error types on the test split. Len. denotes mean predicted scope length in tokens; the mean reference scope length is 5.1 tokens.
Section
Variable
Content (English gloss)
Satisfaction, partnership, family
t514001
Life satisfaction (0–10)
t731112
Steady partnership
t731111
Living with partner
t731110
Marital status
tp1
Feelings about being single
tp2
Disagreements in the relationship
Appendix
Table 9: Main questionnaire read out in VR.
Figure 3: The four avatars used during the interview. From A on the left to D on the right.
Layer
Table
Content
Rate
Records
Session
Experiment
interview metadata
–
27
Player
speaker per interview, connection sessions
–
54
Tracking
Eye
binocular gaze poses, validity, confidence
18.4 Hz
1,350,572
Head , Body
6-DoF head and rig pose
18.4 Hz
1,350,572 each
Facial
63 blendshape weights
18.4 Hz
1,350,572
Left/RightHand
6-DoF wrist pose
18.4 Hz
1,350,572 each
Appendix
Table 10: Released data layers of the 27 interviews. Finger tables hold a row for every frame; the counts give the frames in which the hand was tracked.
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.
Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate in group conversation. Yet available corpora often focus on meetings, task-oriented interaction, text-based interaction, or acted scenarios, and fewer resources support cross-linguistic comparison of spontaneous face-to-face triadic discussion. This paper presents TEIDAN, a multilingual multimodal corpus that currently consists of Japanese and English three-party conversations. TEIDAN records groups of three participants discussing open-ended topics with individual pin microphones, a microphone array, and participant-facing cameras, and provides IPU-based transcripts for both language portions. Earlier studies used subsets of the Japanese portion for task-specific benchmarks in multi-party dialogue modeling; in contrast, this paper presents TEIDAN as a corpus resource spanning both Japanese and English, with planned expansion to additional languages. We describe the collection design, participants, recording setup, transcription format, and corpus statistics, and provide preliminary analyses to illustrate how TEIDAN can support research on turn-taking, addressee recognition, and multimodal grounding in human-human and human-agent interaction.
Taiga Mori, Koji Inoue, Mikey Elmers +2
Graduate School of Informatics, Kyoto University · Japan
Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior. We ask whether contexts centered on spoken negation cues contain measurable multimodal behavioral information: whether they can be distinguished from matched control contexts without lexical or acoustic input, where this information occurs in time, which modalities carry it, and whether it extends to the dialogue partner. We study 27 human-human interviews conducted in virtual reality, comprising temporally aligned gaze, facial, head, body, hand, and finger behavior and 964 annotated negation cues. Treating classification as a predictive probe, we compare 20 time-series models while excluding lexical and acoustic information, and then systematically vary temporal context, interactional source, modality availability, and event timing. Across grouped 10-fold cross-validation, the strongest probes reach up to .75 mean held-out AUROC from speaker-side behavior. Temporal analyses show that predictive information is concentrated around cue onset but remains detectable over a broader surrounding interval, while dialogue-partner behavior carries weaker predictive information with a comparatively diffuse temporal profile. Ablation and timing perturbations further show that facial features produce the largest modality-ablation effect and that the trained probe is sensitive to the temporal organization of the observed events.
Leon Hammerla, Patrick Schrottenbacher, Alexander Mehler