Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.
Figures & tables
Fig. 1: The end-to-end PhysicsMate workflow. (1) Data Ingestion : page-level OCR of the NCTB Grade 9-10 physics textbook, Markdown normalization, and three-page contextual knowledge extraction yielding 811 semantic units; (2) Multi-Relational Knowledge Graph : a 10-type domain ontology with 1,760 nodes and 2,600 directed edges; (3) Benchmark Construction : 1,834 curriculum-grounded Bengali QA pairs stratified into Train/Val/Test; and (4) Efficient Adaptation & Edge Deployment : LoRA fine-tuning of Qwen3 models with 0.6B, 1.7B, and 4B parameters, along with 4-bit offline quantization.
Node Type
Count
Top Edge Type
Count
Concept
397
part_of
585
Question
252
applies_to
531
Quantity
224
illustrates
361
Figure
218
example_of
268
Example
188
depends_on
215
Entity
176
derived_from
118
TABLE I: Knowledge graph statistics. Node types correspond to four cognitive levels: declarative recall (definitions, laws), procedural application (formulas, examples), visual interpretation (figures), and curricular scaffolding (HSC topics).
Field
Description
node_type
Ontology category (concept, law, formula, etc.).
title
Canonical name of the curricular item.
content
Self-contained Bengali description of the knowledge unit.
formula
Mathematical expression, when the item includes one.
unit
SI or textbook-specified physical unit, where applicable.
source_pages
Textbook page identifiers used as evidence.
TABLE II: Fields stored for each extracted knowledge record. Every record is assigned one ontology type and retains full provenance back to the source textbook pages.
Model
Params
Acc. (%)
Token F1
BERTScore
Raw
FT
Raw
FT
Raw
FT
Cloud baselines (zero-shot)
GPT-4o-mini
> 8B
69.5
0.239
0.728
Gemini-2.5-FL
> 10B
82.9
0.267
0.733
Edge models (local, offline)
Qwen3-0.6B
0.6B
0.0
5.5
0.216
0.278
0.696
0.750
TABLE III: Results on the held-out test set ( n=275 ). Accuracy is LLM-judge (Gemini-2.5-Flash-Lite). Cloud baselines are zero-shot; edge models are evaluated with and without LoRA adaptation.
Fig. 2: Capacity-dependent adaptation gains. Fine-tuning produces monotonically increasing accuracy improvements (+5.5, +15.0, and +23.3 percentage points) and consistent gains in Token F1 and BERTScore as model size grows from 0.6B to 4B parameters.
Node Type
n
Raw F1
FT F1
Rel. Gain
Quantity
29
0.262
0.508
93.9%
Law
8
0.286
0.448
56.6%
Definition
28
0.228
0.356
56.1%
Example
35
0.244
0.358
46.7%
Concept
65
0.250
0.342
36.8%
Figure
67
0.263
0.320
21.7%
TABLE IV: Qwen3-4B Token F1 by knowledge-node type ( n=275 test questions). Structured knowledge types benefit most from adaptation.
Fig. 3: Token F1 by knowledge-node type for base vs. fine-tuned Qwen3-4B. The largest gains appear on structured physical quantities (93.9%) and scientific laws (56.6%); entity-level knowledge shows the smallest improvement (8.7%).
Category
Gold Reference
Base Qwen3-4B
Fine-Tuned Qwen3-4B
Optics (ch09_q084) Image properties for an object between a convex lens and its focus?
\bnfont প্রতিবিম্ব অবাস্তব, সোজা ও বস্তুর তুলনায় বড়ো হয়। [Virtual, erect, magnified.]
Fabricated rule: \bnfont প্রতিবিম্ব বামে দাঁড়ায় (বস্তুর বাম দিকে)… [Invents a spatial rule absent from the curriculum.]
\bnfont প্রতিবিম্ব অবাস্তব, সোজা ও বস্তুর তুলনায় বড় হয়। [Matches gold reference.]
Circuits (ch11_q031) Express power in terms of current and resistance.
P=I2R [Power dissipated as heat in the resistor.]
Incomplete: \bnfont V=I×R , যেখানে V= ভোল্ট, I= প্রবাহ… [Stops at Ohm’s law; never derives the power expression.]
\bnfont P=I2R বা P=V2/R দ্বারা ক্ষমতাকে প্রকাশ করা যায়।
TABLE V: Representative base vs. fine-tuned Qwen3-4B outputs. Fine-tuning corrects factual hallucination and incomplete derivation.
We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures vision-language reasoning: train-eval contamination, translation drift, and MCQ saturation. (1) Public training pools (UGPhysics-Train, SciInstruct, MMK12) pass single-stage 5-gram-Jaccard audits with zero hits across all six public physics evals; a three-stage audit (Jaccard -> mxbai-embed-large cosine -> Haiku-4.5 LLM-judge) surfaces 134 near-duplicates and 4,846 paraphrase candidates in SciInstruct alone. (2) A 17-pp Sonnet 4.5 delta on 59 paired Estonian-English olympiad problems (30.5% vs. 13.6%; sign test p=0.011, McNemar p=0.021, paired bootstrap 95% CI [+5.1, +28.9] pp). (3) A 46-pp format-and-novelty gradient on identical Sonnet weights between MCQ (79.7% on PhyX) and open-ended olympiad evaluation (33.4% on PhysOlym-A). We release four artifacts addressing these gaps: PhysCorp-A (6,432-record three-stage-audited multimodal corpus), PhysR1Corp (2,268-record closed-form RL pool), PhysOlym-A (500-problem, 99.8% novel-source held-out olympiad eval with native difficulty labels and an EN/ET bilingual subset), and Physics-R1, a reference GSPO+DAPO recipe cold-started from Qwen3-VL-8B-Thinking. Across 3 seeds, Physics-R1 lifts the audited corpus over the 8B base by +18.3 pp on PhysOlym-A liberal (8.0 -> 26.3 +/- 1.7; 7.1 pp behind Sonnet 4.5), +15.7 pp on PhysReason (23.9 -> 39.6 +/- 6.4; ahead of Qwen3-VL-32B and Gemini 2.5 Pro), +6.9 pp on OlympiadBench-Physics (46.2 +/- 1.5), and +4.1 pp on PhyX MCQ (77.8 +/- 0.3).
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.