cs.CVAug 24, 2026

Simulate, record, verify: A language-portable framework for muscle-grounded articulatory QA (extended version)

Authors: Seungho EumUnsang Park

Organizations: Department of Computer Science and Engineering, Sogang University, Seoul, Korea · Department of Artificial Intelligence, Sogang University, Seoul, Korea

Abstract

Articulatory corpora from real-time MRI and electromagnetic articulography capture tongue motion but carry no traceable labels for the muscle-driven process behind each configuration, and authoring such supervision by hand, separately for every language, does not scale. We present a simulator-based framework that turns controlled biomechanical inputs into verifiable, language-portable QA supervision. Each simulated configuration is stored with its generating input as a structured fact record; deterministic generators derive gold answers from records alone; and naturalization changes only surface form, with every output checked against its record. A new language therefore needs only a renderer and a lexicon, and new question types need no re-simulation. Instantiated as 3DTongueQA on the ArtiSynth Badin tongue model, 295,115 valid meshes yield 891,156 record-checked QA per language in English and Korean (87.2% and 88.6% first-pass verification); a Spanish renderer authored in about 20 minutes reaches 94.1%, and the checker detects 97--99% of injected corruptions. The generated supervision is domain-specific: zero-shot GPT-5 Pro reaches 7.2 Muscle EM, whereas a SpiralNet++--Qwen3-8B model trained on it reaches 62.9±9.262.9\pm9.2 (2.2 with shuffled meshes) and task-specific readouts reach 88.7±0.788.7\pm0.7. Code and templates: https://github.com/esh0504/muscle-grounded-qa.

Explore similar work

CardsList
  1. LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

    Jul 2, 2026Nina Hosseini-Kivanani, Marco Matassoni, Alessio BruttiLuxembourgishText Corpora