Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations
Organizations: RICS · Daegu Gyeongbuk Institute of Science and Technology · IPAI · Samsung Advanced Institute of Technology, Samsung Electronics Co., Ltd · AIIS · Department of Intelligence and Information, Seoul National University
Abstract
As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-injection approaches. However, these lines of work have largely evolved within individual paradigms, leaving their cross-paradigm relationships and trade-offs underexplored, especially for recently emerging embedding-based adaptations. This survey presents a unified framework that categorizes task adaptation methods by where and how task information is encoded: model weights, input prompts, or injected task embeddings. We provide a comprehensive taxonomy that integrates these paradigms, analyze their key strengths and limitations to explain how different adaptation paradigms have evolved, clarify relationships across paradigms, and highlight open problems for future research.
Figures & tables
| Adaptation Type | Task Information Encoding and Injection | Key Strengths | Key Limitations |
|---|---|---|---|
| Weight-Based (PEFT) | Loc: Task information encoded in model weights Build: Partial or additional weights trained Inject: Task-specific weights applied at inference | Strong and stable task performance Well established PEFT libraries Lower inference overhead than ICL | Requires training and parameter access Task-specific modules must be loaded Difficult mixed-task batching |
| Prompt-Based (ICL) | Loc: Task information encoded in prompts Build: Instructions or demonstrations in prompts Inject: Task-specific prompts used as input | Training-free adaptation Applicable to proprietary/API models Natural-language task specification | Repeated inference incurs optimization cost Long prompts increase deployment cost Sensitive to prompt formulation |
| Embedding-Based (ICL-derived) | Loc: Task information encoded in embeddings Build: Embeddings derived from demonstrations Inject: Embeddings injected into model activations | Compact task representations Lower inference overhead than ICL | Requires access to internal activations Requires task embedding construction |
| Embedding-Based (Learned) | Loc: Task information encoded in embeddings Build: Embeddings learned via optimization Inject: Embeddings injected into model activations | Compact task representations Lower inference overhead than ICL | Requires access to internal activations Optimization can be unstable Training may require many iterations |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Adaptation Type | Method | # Trainable Parameters | Big-Bench Hard | Total Runtime |
|---|---|---|---|---|
| Weight-Based Adaptation | LoRA ( Hu et al., 2022 ) | 3407.87K | 60.39 (0.48) | 171.2 min |
| (IA) 3 ( Liu et al., 2022a ) | 524.29K | 60.29 (0.70) | N/A | |
| Prompt-Based Adaptation | 10-shot ICL ( Brown et al., 2020 ) | - | 47.17 (0.98) | 248.4 min |
| Embedding-Based Adaptation (ICL-derived) | FV ( Todd et al., 2024 ) | - | 17.82 (0.37) | 411.2 min |
| MTV ( Huang et al., 2024 ) | 1.02K | 42.54 (0.71) | 328.8 min | |
| I2CL ( Li et al., 2025d ) | 0.13K | 50.60 (1.12) | 178.7 min |
| Category | Benchmark | Output Type | Size | Description |
|---|---|---|---|---|
| Natural Language Understanding | TREC ( Li and Roth, 2002 ) | Classification | 6.0K | TREC Question Classification is a question classification benchmark consisting of open-domain questions annotated with a hierarchical taxonomy of six coarse and fifty fine-grained semantic classes, designed to evaluate question intent understanding and answer type prediction. |
| CoNLL-2003 ( Sang and De Meulder, 2003 ) | Classification | 22K | CoNLL-2003 is a widely used named entity recognition benchmark focusing on the identification of four entity types including persons, locations, organizations, and miscellaneous names, annotated in English and German newswire text to evaluate information extraction. | |
| Subj ( Pang and Lee, 2005 ) | Classification | 10K | Subjectivity detection dataset consisting of 5,000 subjective movie reviews and 5,000 objective plot summaries designed to test the ability of language models to distinguish between fact-based objective descriptions and opinion-based subjective statements. | |
| SST-5 ( Socher et al., 2013 ) | Classification | 12K | SST-5 is a fine-grained sentiment analysis dataset providing five sentiment labels ranging from very negative to very positive, designed to evaluate models’ ability to capture subtle emotional shifts and semantic compositionality in movie review phrases. | |
| SNLI ( Bowman et al., 2015 ) | Classification | 570K | SNLI is a large-scale dataset of human-written English sentence pairs manually labeled into entailment, contradiction, and neutral categories, designed to evaluate models’ ability to capture fundamental logical and semantic relationships. | |
| DBPedia ( Zhang et al., 2015 ) | Classification | 630K | DBPedia is a large-scale topic classification benchmark consisting of Wikipedia articles labeled with 14 non-overlapping DBpedia ontology classes, designed to evaluate models’ ability to perform document-level topic and semantic classification. |
| Category | Benchmark | Output Type | Size | Description |
|---|---|---|---|---|
| Mathematics | AQuA-RAT ( Ling et al., 2017 ) | Multiple Choice Question / Open-Ended Generation | 100K | AQuA-RAT is an algebraic word problem 5-way multiple-choice benchmark where each question is paired with a step-by-step natural-language rationale (often including human-readable math expressions), designed to evaluate interpretable multi-step mathematical reasoning beyond predicting only the final answer. |
| MATH ( Hendrycks et al., 2021 ) | Open-Ended Generation | 13K | MATH is a competition-level mathematics problem-solving benchmark covering diverse topics (e.g., algebra, geometry, number theory, counting & probability, and precalculus), where each problem includes a full step-by-step solution and a final answer, designed to evaluate advanced mathematical reasoning and multi-step problem solving. | |
| SVAMP ( Patel et al., 2021 ) | Open-Ended Generation | 1.0K | SVAMP is an adversarial challenge benchmark of one-unknown arithmetic math word problems created by applying small but systematic variations to existing problems, designed to evaluate robust multi-step arithmetic reasoning beyond keyword matching and shallow statistical cues. | |
| GSM8K ( Cobbe et al., 2021 ) | Open-Ended Generation | 8.8K | GSM8K is a grade-school math word problem benchmark consisting of natural-language questions with final numeric answers, designed to evaluate multi-step arithmetic reasoning required to solve the problems. | |
| TabMWP ( Lu et al., 2022a ) | Multiple Choice Question / Open-Ended Generation | 38K | TabMWP is a tabular math word problem benchmark where each question is paired with a tabular context provided in multiple formats (e.g., table image and structured text), designed to evaluate joint reasoning over tables and natural-language descriptions for deriving correct numerical answers (with gold step-by-step solutions available). | |
| MATH500 ( Lightman et al., 2023 ) | Open-Ended Generation | 0.5K | MATH500 is a curated evaluation subset of the MATH dataset, designed to provide an efficient yet diverse test of competition-style mathematical problem solving. |
| Category | Benchmark | Output Type | Size | Description |
|---|---|---|---|---|
| Question Answering | SQuAD v1.1 ( Rajpurkar et al., 2016 ) | Open-Ended Generation | 98K | SQuAD v1.1 is an extractive reading comprehension benchmark built from questions written on Wikipedia passages, where each question is paired with an answer that is a contiguous text span from the given passage, designed to evaluate models’ ability to perform span-based question answering. |
| TriviaQA ( Joshi et al., 2017 ) | Open-Ended Generation | 96K | TriviaQA is a reading-comprehension benchmark built from trivia questions paired with evidence documents from Wikipedia and the web, designed to test answering complex, compositional questions despite substantial mismatch between questions and supporting evidence. | |
| HotpotQA ( Yang et al., 2018 ) | Open-Ended Generation | 113K | HotpotQA is a multi-hop QA benchmark where each question is paired with supporting facts across multiple Wikipedia articles, designed to test whether models can integrate evidence from more than one document to answer complex questions. | |
| Natural Questions ( Kwiatkowski et al., 2019 ) | Open-Ended Generation | 323K | Natural Questions is an open-domain QA benchmark built from real Google search queries paired with Wikipedia pages, providing annotations for both long answers (passages) and short answers (specific entities) to test end-to-end question answering grounded in retrieved evidence. | |
| Summarization | CNN/DailyMail ( Hermann et al., 2015 ) | Open-Ended Generation | 312K | CNN/DailyMail is a news summarization benchmark consisting of full news articles paired with human-written highlights (bullet-style summary sentences), designed to evaluate models’ ability to produce concise summaries that capture the key information in long-form news reports. |
| XSum ( Narayan et al., 2018 ) | Open-Ended Generation | 227K | XSum is an extreme abstractive summarization benchmark pairing BBC news articles with single-sentence summaries that capture what the article is about, designed to measure models’ ability to produce highly condensed, gist-focused summaries rather than detail-preserving paraphrases. |
| Category | Benchmark | Output Type | Size | Description |
|---|---|---|---|---|
| Safety and Trustworthiness | HateSpeech18 ( De Gibert et al., 2018 ) | Classification | 11K | HateSpeech18 is a sentence-level hate speech dataset sampled from posts on the Stormfront white-supremacist forum and manually labeled as hate vs non-hate, designed to evaluate models’ ability to detect hate speech in highly domain-specific, ideologically skewed online discussions. |
| RealToxicityPrompts ( Gehman et al., 2020 ) | Open-Ended Generation | 100K | RealToxicityPrompts is a dataset of 100k naturally occurring, sentence-level prompts drawn from English web text and paired with toxicity annotations, designed to evaluate whether language models produce toxic continuations when conditioned on real-world prompts spanning a range of toxicity levels. | |
| CrowS-Pairs ( Nangia et al., 2020 ) | Multiple Choice Question | 1.5K | CrowS-Pairs is a social bias evaluation benchmark consisting of sentence pairs that differ in whether they express a stereotype, covering nine bias categories (e.g., race, gender, religion, age), designed to measure models’ tendency to prefer stereotypical statements over less-stereotyping alternatives. | |
| ToxiGen ( Hartvigsen et al., 2022 ) | Classification | 274K | ToxiGen is a large-scale machine-generated toxicity dataset consisting of toxic and benign statements about minority identity groups, created to surface implicit and adversarially crafted toxic language beyond explicit slurs or profanity, and designed to evaluate models’ ability to detect subtle harmful content rather than relying on group-mention shortcuts. | |
| ParaDetox ( Logacheva et al., 2022 ) | Open-Ended Generation | 20K | ParaDetox is a parallel text detoxification dataset pairing toxic sentences with human-written non-toxic paraphrases, designed to test whether models can remove toxicity while preserving the original meaning. | |
| TruthfulQA ( Lin et al., 2022 ) | Multiple Choice Question / Open-Ended Generation | 0.8K | TruthfulQA is a question answering benchmark consisting of questions spanning 38 categories that are crafted to trigger common misconceptions, designed to evaluate whether models can produce truthful, misconception-resistant answers rather than imitating popular human falsehoods. |