Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.
Figures & tables
Fig. 1: Overview of FORGE. The generated continuation is scanned prefix by prefix until toxicity exceeds τ at the onset t⋆ , and leave-one-out attribution within the window pins the responsible position p . Candidate classification (the bridge ) partitions the top- Kbase candidates at p into a suppression set T− and an acceptable set T+ with floor token j⋆ , turning the defect into a repair target that a pluggable backend (QP on the lm_head , or null-space projection on MLP layers) can consume. Verification commits an edit only if sentence toxicity does not rise; the cursor then advances and the sequence is rescanned.
Model
Method
Bias ↓ ( Δ% )
Tox ↓ ( Δ% )
PPL ↑ ( Δ% )
GPT-2
Original
0.4618 (—)
0.1390 (—)
45.578 (—)
LoRA
0.1073 ( +76.8% )
0.0448 (+67.8%)
(+26.5%)
Full FT †
0.0355 (+92.3%)
0.0280 (+79.9%)
(+407.2%) †
FORGE-NS
0.0274 ( +94.1% )
0.0119 ( +91.4% )
(+3.1%)
FORGE-QP
0.1278 (+72.3%)
0.0404 ( +70.9% )
(+7.4%)
Qwen2.5-1.5B
Original
0.4581 (—)
0.1167 (—)
14.457 (—)
TABLE I: Main results on five LLMs. Bias/Tox: reduction vs. baseline ( ↑ better); PPL: increase vs. baseline ( ↓ better). † : language collapse (PPL increase >100% ); such methods are excluded from bold/underline ranking since their low scores reflect degenerate output, not true debiasing.
Model
Resp.-layer depth peak
In FORGE-NS range?
Better backend
Llama-3.2-1B
0.67
Yes (upper)
FORGE-NS
Qwen2.5-3B
0.91–0.97
No
FORGE-QP
TABLE II: Toxicity-responsible layer depth vs. backend performance. Depth is the causal-probe relative-depth peak (first-layer spurious signal excluded).
Fig. 2: Low-resource stability on GPT-2. Axis-level bias (full eval set, 1,886 prompts) vs. training-set size N ; gray dotted line is the unrepaired baseline. Both FORGE backends converge and lock onto a stable low bias; the gradient baselines (LoRA, Full FT) oscillate and never converge, and Full FT’s low bias comes with severe perplexity collapse (Table I ).