Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic, and subject-matter expertise. In particular, traditional medical systems such as Ayurveda embody centuries of nuanced textual and clinical knowledge that mainstream LLMs fail to accurately interpret or apply. We introduce AyurParam-2.9B, a domain-specialized, bilingual language model fine-tuned from Param-1-2.9B using an extensive, expertly curated Ayurveda dataset spanning classical texts and clinical guidance. AyurParam's dataset incorporates context-aware, reasoning, and objective-style Q&A in both English and Hindi, with rigorous annotation protocols for factual precision and instructional clarity. Benchmarked on BhashaBench-Ayur, AyurParam not only surpasses all open-source instruction-tuned models in its size class (1.5--3B parameters), but also demonstrates competitive or superior performance compared to much larger models. The results from AyurParam highlight the necessity for authentic domain adaptation and high-quality supervision in delivering reliable, culturally congruent AI for specialized medical knowledge.
Figures & tables
Figure 1: Data preparation pipeline
Figure 2: Language-wise distribution of the collected corpus across Sanskrit, Hindi, Marathi, and English sources
Figure 3: QA-Distribution
Similar Range Models
Model
BBA
BBA English
BBA Hindi
AyurParam-2.9B-Instruct
39.97
41.12
38.04
Llama-3.2-3B-Instruct
33.20
35.31
29.67
Qwen2.5-3B-Instruct
32.68
35.22
28.46
granite-3.1-2B
31.10
33.39
27.30
gemma-2-2B-it
28.40
29.38
26.79
Table 1: Overall performance comparison on BBA dataset. Results show accuracy (%) across different model sizes. Our AyurParam-2.9B-Instruct model achieves the best performance among similar-sized models and competitive results compared to much larger models.
Similar Range Models
Difficulty
AyurParam-2.9B
Llama-3B
Qwen-3B
Granite-2B
Gemma-2B
Llama-1B
Easy
43.93
36.42
35.55
33.90
29.96
27.44
Medium
35.95
29.66
29.57
28.06
26.83
25.23
Hard
31.21
28.51
28.23
26.81
24.96
25.39
Table 2: Performance breakdown by question difficulty. Results demonstrate that AyurParam-2.9B-Instruct maintains strong performance across all difficulty levels, with particularly notable results on easy questions.
Similar Range Models
Type
Llama-1B
Qwen-3B
Llama-3B
AyurParam-2.9B
Gemma-2B
Assert./Reason.
59.26
51.85
40.74
44.44
33.33
Fill blanks
26.97
29.21
34.83
29.78
32.02
MCQ
26.34
32.70
33.17
40.12
28.33
Match col.
26.83
29.27
29.27
24.39
36.59
Table 3: Performance analysis across different question types. AyurParam-2.9B-Instruct shows strong performance on MCQ questions and competitive results across all question formats.
This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via continual pre-training and merging. We address the scarcity of specialized non-English data by constructing a high-quality German medical corpus (FineMed-de) from FineWeb2. This corpus is used to continually pre-train and merge three well-known LLMs (ranging from 7B to 24B parameters), creating the DeFineMed model family. A comprehensive evaluation confirms that specialization dramatically enhances 7B model performance on German medical benchmarks. Furthermore, the pairwise win-rate analysis of the Qwen2.5-based models demonstrates an approximately 3.5-fold increase in the win-rate against the much larger Mistral-Small-24B-Instruct through domain adaptation. This evidence positions specialized 7B models as a competitive, resource-efficient solution for complex medical instruction-following tasks. While model merging successfully restores instruction-following abilities, a subsequent failure mode analysis reveals inherent trade-offs, including the introduction of language mixing and increased verbosity, highlighting the need for more targeted fine-tuning in future work. This research provides a robust, compliant methodology for developing specialized LLMs, serving as the foundation for practical use in German-speaking healthcare contexts.
Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of medicine. We argue that robust cross-lingual medical transfer requires Hindi reasoning. To this end, we introduce HiMed, a Hindi reasoning medical corpus and benchmark suite covering both Western and Indian medicine. We further propose HiMed-8B, a Hindi-form medical reasoning LLM, through the design of decaying scaffolding reward. Extensive experiments demonstrate improvement in Hindi medical reasoning performance and reduction in the English--Hindi accuracy gap. Ablation studies validate the contribution of each training stage and reward component. All data and code are available on GitHub: https://github.com/FreedomIntelligence/HiMed.
Dingfeng Jiang, Han Yan, Chenze Ma +12
1The Chinese University of Hong Kong, Shenzhen · 2Indian Institute of Technology (Banaras Hindu University) Varanasi · 3Tongji University +4
Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in regions like rural India, where patients often express complex medical queries in native Indic languages and rely on multimodal inputs such as medical images. Existing English-centric MLLMs struggle to support such use cases, limiting equitable access to AI-driven healthcare assistance. To address this challenge, we introduce ArogyaBodha, a large-scale multilingual multimodal medical question-answer dataset constructed from eight heterogeneous sources, covering 31 body systems, six imaging modalities, and 21 clinical domains across English and seven major Indian languages. We further propose ArogyaSutra, an actor-critic-based multi-agent framework that integrates tool grounding with dual-memory mechanisms for step-wise, reasoning-aware decision making, and uses stored actor-critic simulation trajectories for distillation. Experiments show that our dataset and framework improve multilingual medical reasoning accuracy across all Indic languages, with ablations validating the contribution of each component. The source code and dataset are available at: https://iitp-cse.github.io/ArogyaSutra/
Tanmoy Kanti Halder, Akash Ghosh, Subhadip Baidya +2
1Indian Institute of Technology Patna · 3Prasannadeb Women’s College · 2Indian Institute of Technology Kanpur