Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning
Authors: Zhenchao Tang, Fang Wang, Haohuai He, Jiale Zhou, Tianxu Lv, Jun Zhu, Shouzhi Chen, Minghao Yang, +8 more
Organizations: Tencent AI for Life Sciences Lab, Shenzhen, China. · Sun Yat-sen University, Shenzhen, China. · The Hong Kong Polytechnic University, Hong Kong SAR, China. · Westlake University, Hangzhou, China. · The Hong Kong University of Science and Technology, Hong Kong SAR, China.
Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather than the sparse aleatoric stylistic variation typical of general text. Based on this discovery, we propose Balanced Fine-Tuning (BFT), a dual-scale post-training method that combines group-normalized token reweighting with sequence-level reallocation toward knowledge-dense samples exhibiting dense epistemic uncertainty. Across medical evaluation, biological reasoning, sparse-reward RL, and biological representation tasks, BFT provides more consistent gains than SFT and DFT under a shared training setup. When replacing the default closed-source backbones in GeneAgent (GPT-4o) and VCWorld (Gemini-2.5-Flash), the BFT-aligned 70B model delivers stronger performance across biological process reasoning and chemical perturbation prediction. Critically, all BFT variants further improve after subsequent GRPO with sparse rewards, while SFT and DFT degrade, suggesting that epistemic-aware post-training provides a more robust policy initialization. Beyond text generation, BFT-aligned LLMs produce more accurate and professional biomedical profile texts; after encoding these profiles with a text embedding model, the resulting representations support gene-level, cell-level, and perturbation-response tasks, suggesting that BFT-enhanced generation can facilitate biological representation and, in turn, broader biomedical downstream tasks.
Figures & tables
Figure 1: Overview of BFT. a : Semantic composition of low-confidence tokens inside Group A (sparse-low) and Group B (dense-low). b : 4096-token trace illustrating one sparse-low window (Group A) and one dense-low window (Group B). c : BFT applies token- and sample-level adjustments.
Figure 2: Medical competence and biological process reasoning. a : Medical competence is measured across axis-wise (left) and theme-wise (right). b : Biological reasoning benchmark.
Method
C32
HepG2C3A
HOP62
Hs 766T
PANC-1
Average
VCWorld (Gemini-2.5-Flash)
0.69 ± .01
0.67 ± .02
0.72 ± .01
0.69 ± .02
0.62 ± .02
0.68
VCWorld (Base 70B)
0.42 ± .02
0.37 ± .02
0.41 ± .02
0.39 ± .02
0.40 ± .02
0.40
DFT 70B
0.54 ± .03
0.46 ± .02
0.49 ± .02
0.45 ± .03
0.51 ± .02
0.49
SFT 70B
0.58 ± .02
0.53 ± .02
0.58 ± .02
0.55 ± .01
0.57 ± .02
0.56
BFT-sample 70B
0.61 ± .02
0.58 ± .02
0.62 ± .02
0.58 ± .02
0.59 ± .02
0.60
BFT-token 70B
0.68 ± .02
0.65 ± .02
0.67 ± .01
0.66 ± .02
0.64 ± .02
0.66
Table 1: Chemical perturbation reasoning accuracy for joint Up/Down/No classification. Upper block: original VCWorld; middle: aligned 70B replacements; lower: after GRPO. Cells are mean ± SD over eight stochastic inference passes of one checkpoint (decoding, not training, variability). Displayed averages are rounded to two decimals; tests use unrounded cell-line means. Best in each column is shown in bold .
Figure 3: General capability retention, window sensitivity and preference transmission. a : MMLU and CMMLU evaluations. b : BFT performance under different window sizes g , measured on the 14B model; absolute values are therefore not comparable to the 70B results reported in Table 1 . c : Transmission of an injected owl preference.
Bio conservation
Batch correction
Aggregate score
Method
Isolated labels
KMeans NMI
cLISI
Silhouette batch
iLISI
Graph conn.
Batch correction
Bio conservation
Total
scMODAL
0.54
0.66
1.00
0.87
0.46
0.69
0.62
0.67
0.65
BFT
0.52
0.61
0.98
0.96
0.34
0.52
0.50
0.62
0.57
Harmony
0.48
0.44
0.98
0.75
0.08
0.48
0.38
0.58
0.50
BBKNN
0.52
0.51
0.99
0.48
0.00
0.45
0.19
0.60
0.44
Table 2: Comparison of multimodal integration on the scIB benchmark Luecken et al. (2022) . Bio conservation (Isolated labels, KMeans NMI, cLISI) and Batch correction (Silhouette batch, iLISI, Graph connectivity) are reported per-metric; aggregate scores are reported separately, with Total = 0.4 Batch + 0.6 Bio.
Figure 4: Unified representation across genetic and cellular levels. a : UMAP visualization of gene embeddings. b : Gene-level downstream performance on biological property prediction and gene interaction prediction. c : Cell-level representation performance on phenotype and cell-type.
Figure 5: Single-cell perturbation-response prediction across four datasets. BFT-generated text embeddings are used as inputs to the STATE decoder; profile generation is zero-shot with respect to the evaluated perturbation datasets, while the STATE decoder is retained.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Domain / prompt
Group-B prev.
Entity A
Entity B
Causal-B
PubMed Central full text
biomedical, constructed x
20.4%
12.5%
60.9%
30.6%
Our synthetic supervision
biomedical, natural x
18.6%
11.8%
57.4%
31.3%
UltraChat-200k
conversational, natural x
2.9%
3.6%
4.1%
11.1%
NuminaMath-CoT
mathematical, natural x
7.5%
3.0%
3.2%
42.6%
Appendix
Table 3: Matched diagnostic across real and synthetic biomedical text and two general-domain controls. Group-B prevalence measures how often dense-low windows occur. Entity A → B is the within-corpus change in the share of biomedical entities among low-confidence tokens; this contrast, rather than an absolute cross-corpus entity rate, is our primary quantity. Causal-B is the share of Group-B low-confidence tokens matched by the fixed causal-connective list.
Tokens retained at full weight
VCWorld accuracy
All tokens (SFT)
0.562
Random selection, size-matched to dense windows
0.571
Sparse-low windows only
0.508
Dense-low windows only
0.674
Full BFT (continuous)
0.696
Appendix
Table 4: Retention counterfactual on VCWorld (70B). The four discrete controls differ only in which tokens receive unit rather than DFT weight; random selection is size-matched to Group B. Full BFT is the continuous, threshold-free reference.
Method
VCWorld accuracy
DFT
0.472
SFT
0.531
BFT
0.649
Appendix
Table 5: VCWorld accuracy after supervised training directly on raw NCBI text, without instruction synthesis.
Method
Seed 1
Seed 2
Seed 3
SD
DFT 70B
0.490
0.466
0.508
0.021
SFT 70B
0.562
0.541
0.578
0.019
BFT 70B
0.696
0.679
0.708
0.015
Appendix
Table 6: VCWorld average accuracy across three independent fine-tuning seeds. The final column is the standard deviation across the three seed-level means, not decoding variability.
Sample-level rule
VCWorld accuracy
No sample coefficient (BFT-token)
0.660
Original minimum coefficient
0.696
Batch-normalized coefficient sb
0.700
Fixed-count coefficient sb(K) , K=8
0.690
Appendix
Table 7: Sample-level controls on VCWorld (70B). All rows use the same supervised recipe. Best in each column is shown in bold .
Base model
DFT
SFT
BFT
Qwen2.5-14B
0.42
0.47
0.57
Llama-3.1-8B
0.37
0.42
0.52
DeepSeek-R1-Distill-70B
0.49
0.56
0.70
Appendix
Table 8: VCWorld average accuracy across model families under an identical recipe. The DeepSeek-R1-Distill row is the 70B result from Table 1 ; the other two families are smaller, so absolute values are lower and only the within-row ordering is meaningful.
Objective
VCWorld accuracy
DFT + BFT sample coefficient
0.528
iw-SFT [ 21 ]
0.578
UFT [ 22 ]
0.591
Anchored SFT [ 23 ]
0.607
ProFit [ 24 ]
0.618
BFT (ours)
0.696
Appendix
Table 9: VCWorld average accuracy for alternative reweighting objectives, 70B, shared recipe.
Method
Entities
Causal
Length
Drug/Mut.
SFT 70B
32.4±8.2
8.6±2.4
3187±557
75%
DFT 70B
28.1±9.5
7.2±2.8
2756±549
65%
BFT 70B
55.4±7.1
15.8±3.1
5458±635
95%
Appendix
Table 10: Automated response analysis for SFT, DFT, and BFT-aligned 70B models on 20 biomedical causal reasoning prompts. Count-valued columns report mean ± standard deviation across prompts, not repeated inference runs. Entities : unique biomedical entities per response (SciSpacy en_core_sci_lg ). Causal : causal connective count from a 45-term indicator list. Length : response character count. Drug/Mut. : fraction of responses mentioning ≥ 1 specific drug or mutation. Metric extraction is deterministic conditional on the generated responses.
Reasoning budget N
DFT
SFT
BFT
256 tokens
0.441
0.512
0.598
512 tokens
0.462
0.537
0.641
1024 tokens
0.481
0.554
0.672
2048 tokens (main setting)
0.490
0.562
0.696
4096 tokens
0.491
0.564
0.698
Appendix
Table 11: VCWorld average accuracy (70B) under a shared reasoning budget N and up to 32 additional forced-completion tokens. The 2048-token row matches Table 1 ; 4096 tests sensitivity to a larger cap.
Method
Triples
KB-alignable
Support rate
Adversarial reading
SFT 70B
608
209 (34.4%)
71.3%
24.5%
DFT 70B
579
193 (33.3%)
68.9%
23.0%
BFT 70B
951
328 (34.5%)
82.6%
28.5%
BFT 70B, length-matched
641
215 (33.5%)
81.4%
27.3%
Appendix
Table 12: Targeted knowledge-base support check for extracted relation triples from 20 NSCLC trajectories. The length-matched row prefix-truncates BFT responses to SFT’s mean 3,187-character length before extraction. Rates apply only to the KB-alignable subset.
Encoder
Method
TF Range
Dosage
No-methy
Lys4
GGI
PPI
-
scGPT
0.557
0.853
0.804
0.847
0.635
0.661
-
GenePT
0.585
0.891
0.910
0.939
0.806
0.681
ada-002
DFT
0.519
0.808
0.761
0.793
0.598
0.604
SFT
0.534
0.838
0.769
0.812
0.616
0.622
BFT (Ours)
0.629
0.913
0.927
0.910
0.823
0.769
BioBERT
DFT
0.495
0.806
0.755
0.777
0.566
0.583
Appendix
Table 13: Quantitative evaluation of biological representations at the gene level. We compare the performance of embeddings derived from different fine-tuning strategies (SFT, DFT, and BFT) across two text-embedding encoders. We use OpenAI text-embedding-ada-002 , the encoder used by GenePT for NCBI gene summaries; BFT instead generates profile text before encoding. This comparison holds the encoder fixed, not the source text. BioBERT is included as a domain-specific encoder. All values are reported as AUC. Bold indicates the best performance in each column.
Phenotype Clustering
Cell Type Clustering
Encoder
Method
ARI
AMI
ASW
ARI
AMI
ASW
-
PCA
0.14
0.16
0.011
0.15
0.28
0.031
-
scGPT
0.11
0.13
0.011
0.26
0.32
0.039
-
GenePT
0.13
0.12
0.019
0.30
0.46
0.041
ada-002
DFT
0.09
0.11
0.003
0.10
0.21
0.016
SFT
0.11
0.13
0.006
0.12
0.25
0.021
Appendix
Table 14: Quantitative evaluation of cell-level clustering performance across phenotype and cell type labels. We compare embeddings derived from SFT, DFT, and BFT using text-embedding-ada-002 (the encoder GenePT uses for NCBI summaries; our method encodes BFT-generated profiles) and BioBERT (a domain-specific encoder), against traditional baselines (PCA, scGPT, and GenePT). Bold indicates the best performance in each column.
Figure 6: Workflow for extracting biological embeddings from LLM-BFT. a : LLM-BFT generates profile texts for entities of interest (e.g., a specific gene). The gene profile text is encoded by OpenAI text-embedding-ada-002 to obtain gene embeddings. GenePT uses this encoder on NCBI gene summaries, whereas our method encodes BFT-70B-generated profiles. b : For a single-cell dataset, gene embeddings are weighted by gene expression values to generate cell embeddings.
Figure 7: UMAP visualization of the multi-gene input task. a : For GGI, the input embedding of the classifier is directly concatenated from the embeddings of two genes. b : For PPI, the input embedding of the classifier is directly concatenated from the embeddings of two proteins.
Figure 8: UMAP visualization of cell-level embeddings. a : PCA embeddings of the raw data, colored by cell type labels (cell type heterogeneity), patient labels (batch labels), and phenotype labels (disease heterogeneity), respectively. b : Cell embeddings derived from LLM-BFT, colored by cell type labels (cell type heterogeneity), patient labels (batch labels), and phenotype labels (disease heterogeneity), respectively.
Figure 9: UMAP visualization of single-cell multimodal data integration results. Rows 1 to 4 represent different integration methods, respectively. Columns 1 to 3 correspond to different coloring labels (modality, cell type, and donor), respectively.
Despite the success of large language models (LLMs) on general-purpose tasks, their performance in highly specialized domains such as biomedicine remains unsatisfactory. A key limitation is the inability of LLMs to effectively leverage biomedical tools, which clinical experts and biomedical researchers rely on extensively in daily workflows. While recent general-domain tool-calling datasets have substantially improved the capabilities of LLM agents, existing efforts in the biomedical domain largely rely on in-context learning and restrict models to a small set of tools. To address this gap, we introduce BioTool, a comprehensive biomedical tool-calling dataset designed for fine-tuning LLMs. BioTool comprises 34 frequently used tools collected from the NCBI, Ensembl, and UniProt databases, along with 7,040 high-quality, human-verified query-API call pairs spanning variation, genomics, proteomics, evolution, and general biology. Fine-tuning a 4-billion-parameter LLM on BioTool yields substantial improvements in biomedical tool-calling performance, outperforming cutting-edge commercial LLMs such as GPT-5.1. Furthermore, human expert evaluations demonstrate that integrating a BioTool-fine-tuned tool caller significantly improves downstream answer quality compared to the same LLM without tool usage, highlighting the effectiveness of BioTool in enhancing the biomedical capabilities of LLMs. The full dataset and evaluation code are available at https://github.com/gxx27/BioTool
Large Language Models such as GPT-4o and GPT-5 achieve strong zero-shot performance on biomedical claim verification, but cost and opacity limit scalable use. We fine-tune three small LLMs: Phi-3-mini (3.8B), Qwen2.5-3B, and Mistral-7B, via QLoRA on SciFact and HealthVer, providing the first study of QLoRA models against GPT-4o and fine-tuned BioLinkBERT encoders. Mistral-7B QLoRA surpasses both GPT-4o and GPT-5 (up to 12% F1 gain) at a fractional cost using just 1,008 training examples. We conduct extensive in-domain and cross-domain evaluation: models trained on SciFact tested on HealthVer and vice versa, at matched sizes to isolate dataset structure from data quantity. We identify a previously unreported structural artifact in SciFact that inflates in-domain scores, and show through bidirectional out-of-domain evaluation that training on structurally sound data enables robust cross-domain transfer. We plan to release all code and adapter checkpoints.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Ada Fang, Nikitha Thoduguli, Lukas Fesser +3
Harvard University · Massachusetts Institute of Technology