Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.
Figures & tables
Fig. 1: The end-to-end PhysicsMate workflow. (1) Data Ingestion : page-level OCR of the NCTB Grade 9-10 physics textbook, Markdown normalization, and three-page contextual knowledge extraction yielding 811 semantic units; (2) Multi-Relational Knowledge Graph : a 10-type domain ontology with 1,760 nodes and 2,600 directed edges; (3) Benchmark Construction : 1,834 curriculum-grounded Bengali QA pairs stratified into Train/Val/Test; and (4) Efficient Adaptation & Edge Deployment : LoRA fine-tuning of Qwen3 models with 0.6B, 1.7B, and 4B parameters, along with 4-bit offline quantization.
Node Type
Count
Top Edge Type
Count
Concept
397
part_of
585
Question
252
applies_to
531
Quantity
224
illustrates
361
Figure
218
example_of
268
Example
188
depends_on
215
Entity
176
derived_from
118
TABLE I: Knowledge graph statistics. Node types correspond to four cognitive levels: declarative recall (definitions, laws), procedural application (formulas, examples), visual interpretation (figures), and curricular scaffolding (HSC topics).
Field
Description
node_type
Ontology category (concept, law, formula, etc.).
title
Canonical name of the curricular item.
content
Self-contained Bengali description of the knowledge unit.
formula
Mathematical expression, when the item includes one.
unit
SI or textbook-specified physical unit, where applicable.
source_pages
Textbook page identifiers used as evidence.
TABLE II: Fields stored for each extracted knowledge record. Every record is assigned one ontology type and retains full provenance back to the source textbook pages.
Model
Params
Acc. (%)
Token F1
BERTScore
Raw
FT
Raw
FT
Raw
FT
Cloud baselines (zero-shot)
GPT-4o-mini
> 8B
69.5
0.239
0.728
Gemini-2.5-FL
> 10B
82.9
0.267
0.733
Edge models (local, offline)
Qwen3-0.6B
0.6B
0.0
5.5
0.216
0.278
0.696
0.750
TABLE III: Results on the held-out test set ( n=275 ). Accuracy is LLM-judge (Gemini-2.5-Flash-Lite). Cloud baselines are zero-shot; edge models are evaluated with and without LoRA adaptation.
Fig. 2: Capacity-dependent adaptation gains. Fine-tuning produces monotonically increasing accuracy improvements (+5.5, +15.0, and +23.3 percentage points) and consistent gains in Token F1 and BERTScore as model size grows from 0.6B to 4B parameters.
Node Type
n
Raw F1
FT F1
Rel. Gain
Quantity
29
0.262
0.508
93.9%
Law
8
0.286
0.448
56.6%
Definition
28
0.228
0.356
56.1%
Example
35
0.244
0.358
46.7%
Concept
65
0.250
0.342
36.8%
Figure
67
0.263
0.320
21.7%
TABLE IV: Qwen3-4B Token F1 by knowledge-node type ( n=275 test questions). Structured knowledge types benefit most from adaptation.
Fig. 3: Token F1 by knowledge-node type for base vs. fine-tuned Qwen3-4B. The largest gains appear on structured physical quantities (93.9%) and scientific laws (56.6%); entity-level knowledge shows the smallest improvement (8.7%).
Category
Gold Reference
Base Qwen3-4B
Fine-Tuned Qwen3-4B
Optics (ch09_q084) Image properties for an object between a convex lens and its focus?
\bnfont প্রতিবিম্ব অবাস্তব, সোজা ও বস্তুর তুলনায় বড়ো হয়। [Virtual, erect, magnified.]
Fabricated rule: \bnfont প্রতিবিম্ব বামে দাঁড়ায় (বস্তুর বাম দিকে)… [Invents a spatial rule absent from the curriculum.]
\bnfont প্রতিবিম্ব অবাস্তব, সোজা ও বস্তুর তুলনায় বড় হয়। [Matches gold reference.]
Circuits (ch11_q031) Express power in terms of current and resistance.
P=I2R [Power dissipated as heat in the resistor.]
Incomplete: \bnfont V=I×R , যেখানে V= ভোল্ট, I= প্রবাহ… [Stops at Ohm’s law; never derives the power expression.]
\bnfont P=I2R বা P=V2/R দ্বারা ক্ষমতাকে প্রকাশ করা যায়।
TABLE V: Representative base vs. fine-tuned Qwen3-4B outputs. Fine-tuning corrects factual hallucination and incomplete derivation.