Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack
Organizations: Laboratory for Artificial Intelligence and New Forms of Education Faculty of Artificial Intelligence in Education, Central China Normal University, Wuhan, China · Wuhan United Imaging Healthcare Co., Ltd., Wuhan, China
Abstract
Science teachers frequently search for documentary excerpts not by describing what appears on screen, but by querying the abstract concepts they intend to teach. This use case exposes a limitation of existing language-based video moment retrieval methods, which typically assume that queries describe observable events, whereas instructional search requires retrieving concrete visual phenomena that instantiate an underlying scientific principle. We study this setting as concept-to-example video retrieval, an abstract-needle-in-a-haystack problem where compact curriculum concepts must be grounded in temporally sparse documentary evidence. To bridge this abstraction gap, we propose Concept-Driven Domain Adaptation (CDDA), a three-stage framework for adapting two-tower vision-language models to concept-level retrieval. CDDA treats concepts as intermediate semantic anchors: it first structures the textual embedding space with textbook and teacher-handbook example-concept pairs, then transfers this concept-aware geometry to documentary visuals under a frozen visual encoder, and finally jointly adapts both encoders with sparse visual concept supervision. From a geometric perspective, this staged alignment reduces text-concept and vision-concept angular gaps, thereby encouraging concept-level adaptation while preserving the pretrained model's concrete image description alignment. On a curated middle-school physics retrieval benchmark, CDDA achieves stronger pedagogically oriented concept retrieval than several competitive multimodal baselines, including Qwen3-VL-Embedding-2B, while maintaining concrete image-text matching after adaptation.
Figures & tables
| Collection | Episodes | Duration | Production team |
|---|---|---|---|
| When We Left Earth: The NASA Missions | 4 | 45 min. | Discovery Channel |
| Earth: The Power of the Planet | 1 | 60 min. | BBC Two |
| The Fabric of the Cosmos | 2 | 55 min. | PBS NOVA |
| 100 Greatest Discoveries | 1 | 44 min. | Discovery Channel |
| Cosmos: A Spacetime Odyssey | 1 | 45 min. | National Geographic |
| What on Earth Is Wrong with Gravity? | 1 | 44 min. | BBC Horizon |
| Force | Gravity | Newton | Newton’s | Universal | pedagogical | |
| First Law | Gravity | scores | ||||
| # of segments | (6) | (20) | (9) | (1) | (25) | |
| CDDA | 5 | 6 | 5 | almost | 5+1NtM | 24 |
| CDDA(II&III) | 3 | 6 | 4 | almost | 4+1NtM | 20 |
| CDDA(II&III) d+c | 1 | 5 | 3 | almost | 4 | 13 |
| CDDA(II&III) d | 0 | 4+1 | 1 | almost | 3+1NtM | 12 |
| Method | MAP@10 | NDCG@10 | MAP@5 | NDCG@5 |
|---|---|---|---|---|
| Gemini Embed. 2 | 0.4924 | 0.5710 | 0.5300 | 0.5968 |
| CDDA | 0.3848 | 0.5402 | 0.3807 | 0.5094 |
| Qwen3-VL-2B | 0.3723 | 0.5374 | 0.3793 | 0.4814 |
| InternVideo-Next | 0.3245 | 0.4448 | 0.4167 | 0.4941 |
| CDDA(II&III) | 0.2974 | 0.4309 | 0.3753 | 0.4601 |
| ImageBind | 0.2789 | 0.4161 | 0.2340 | 0.3470 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model Name and Version | Parameter Scales | Checkpoint File Size |
|---|---|---|
| Gemini-Embedding-2 | Not Officially Declared | Not Officially Declared |
| Qwen3-VL-Embedding-2B | 2B | 4.3 GB |
| ImageBind-huge | Not Officially Declared | 4.8 GB |
| InternVideo-Next large_p14_res224_f16) | 0.3B 1 | 622.1 MB |
| CN-CLIP ViT-B/16 | 188M [ 35 ] | 753.2 MB |
| Concept | Example / fact |
|---|---|
| Force | The tires, pedals, or handlebars of a bicycle are engraved with uneven patterns to increase friction by increasing the roughness of the contact surface. |
| A person kicks a ball; a horse pulls a cart; a person pushes a table. | |
| Gravity | The effect of gravity on an object in free fall is to make its velocity increase. |
| An astronaut experiences weightlessness in the Space Station. | |
| Newton | In 1664, force was defined as the time derivative of momentum. |
| Newton’s three laws of motion were formulated as axioms. |
| Data source | Size | Usage |
|---|---|---|
| Image–concept examples | 44 / 12 / 6 | Train / validation / test diagnostic |
| Textual example–concept pairs | 118 + 32 = 150 | Stage 1 text structuring |
| Evaluation documentary subtitles | 832 | Candidate segmentation |
| Evaluation candidate segments | 146 | Retrieval evaluation |
| Concept queries | 5 | Retrieval queries |
| Concept query | Relevant segments |
|---|---|
| Force | 6 |
| Gravity | 20 |
| Newton | 9 |
| Newton’s First Law | 1 |
| Universal Gravity | 25 |
| Variant | Difference from CDDA |
|---|---|
| CDDA(II&III) | Removes Stage 1 textual concept structuring and trains only with image–concept pairs. |
| CDDA(II&III) d | Replaces abstract concept labels with literal fact descriptions. |
| CDDA(II&III) d+c | Appends a concept sentence to each fact description, e.g., “Planets orbit the Sun in elliptical orbits. This is universal gravity.” |
| CDDA(II&III) d text | Text-only retrieval counterpart of the description-based variant. |
| Item | Value |
|---|---|
| Base model | CN-CLIP [ 35 ] |
| Text adaptation component | Chinese RoBERTa [ 36 ] |
| Optimizer | AdamW |
| Batch size, Stage 2 / Stage 3 | 4 / 4 |
| Learning rate, Stage 2 / Stage 3 | / |
| Epochs, Stage 2 / Stage 3 | 15 / 15 |