cs.CLOct 4, 2026

FORGE: Verification-Gated Behavioral Repair for Generative Language Models

Authors: Hsin-Ling Hsu, Min-Yu Chen, Nai-Chia Chen, Yan-Ru Chen, Yi-Ling Chang, Fang Yu

Organizations: Department of Management Information Systems National Chengchi University Taipei, Taiwan

Abstract

Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.

Figures & tables

Explore similar work

CardsList
  1. CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

    Apr 16, 2026Yian Wang, Yuen Chen, Agam Goyal +1DetoxificationHarmful Language

  2. Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

    Sep 27, 2026Gert Lek, Abele Malan, Chaoyi Zhu +3Large Language Model SafetyMasked Diffusion Language Models

  3. COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models

    May 28, 2026Arya Fayyazi, Mehdi Kamal, Massoud PedramLarge Language Model BiasFrozen Language Model