Multimodal Large Language Models as Image Classifiers
Organizations: Visual Recognition Group, Czech Technical University in Prague
Abstract
Multimodal Large Language Model (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and Vision-Language Models (VLMs) report conflicting conclusions, and we show these conflicts stem from protocols that either inflate or underestimate performance. Across the most common evaluation protocols, we identify and fix key issues: model outputs that fall outside the provided class list and are discarded, inflated results from weak multiple-choice distractors, and open-world setting that underperforms only due to poor output mapping. We additionally quantify the impact of commonly overlooked design choices - batch size, image ordering, and text encoder selection - showing they substantially affect accuracy. Evaluating on ReGT, our multilabel reannotation of 625 ImageNet-1k classes, reveals that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models. Much of the reported MLLM underperformance on classification is thus an artifact of noisy ground truth and flawed evaluation protocol rather than genuine model deficiency. Models less reliant on supervised training signals prove most sensitive to annotation quality. Finally, we show that MLLMs can assist human annotators: in a controlled case study, annotators confirmed or integrated MLLM predictions in approximately 50% of difficult cases, demonstrating their potential for large-scale dataset curation. This work is part of the Aiming for Perfect ImageNet-1k project, see https://klarajanouskova.github.io/ImageNet/.
Figures & tables
| Cat. | Definition | Cat. | Definition |
|---|---|---|---|
| # | 31250 | 18071 | 16177 | 1894 | 11834 | 10756 | 1078 | 1345 |
|---|---|---|---|---|---|---|---|---|
| % | 100 | 57.8 | 51.8 | 6.1 | 37.9 | 34.4 | 3.5 | 4.3 |
| ImGT | ReGT | Im Re | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Task | ||||||||||
| PaliGemma2-mix-28B/448 | OW | 37.11 (13) | 47.94 +10.8 , (13) | 39.74 (13) | 41.49 (13) | 24.76 (13) | 54.55 (13) | 56.24 (13) | 37.66 (5) | 35.78 | |
| LLaVA-OneVision-72B-Chat | OW | 62.00 (12) | 70.58 +8.6 , (12) | 67.11 (12) | 70.67 (12) | 36.69 (4) | 72.52 (12) | 75.87 (12) | 39.05 (3) | 59.39 | |
| InternVL3.5-38B | CW+ | 64.31 (11) | 72.57 +8.3 , (11) | 67.69 (11) | 71.05 (11) | 39.02 (1) | 76.91 (11) | 80.43 (11) | 41.74 (1) | 61.53 | |
| Qwen3-VL-235B-A22B-Inst | OW | 68.74 (10) | 76.68 +7.9 , (10) | 73.61 (10) | 77.79 (10) | 37.91 (3) | 78.71 (10) | 82.46 (10) | 41.28 (2) | 65.87 | |
| GPT-4o-2024-08-06 | CW+ | 76.40 (9) | 82.36 +6.0 , (9) | 81.11 (9) | 86.12 (9) | 38.28 (2) | 82.27 (9) | 86.74 (9) | 37.66 (5) | 72.79 | |
| ImGT | ReGT | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | ||||||||||
| PaliGemma 2 | OW | 37.11 | 47.94 | 39.74 | 41.49 | 24.76 | 54.55 | 56.24 | 37.66 | |
| LLaVA-OV | OW | 62.00 +18.6 | 70.58 +17.8 | 67.11 +17.5 | 70.67 +18.7 | 36.69 +7.1 | 72.52 +20.3 | 75.87 +22.0 | 39.05 +4.1 | |
| CW+ | 52.66 +9.3 | 62.82 +10.1 | 59.54 +10.0 | 62.32 +10.4 | 35.80 +6.2 | 63.61 +11.4 | 65.84 +11.9 | 41.37 +6.4 | ||
| CW | 43.40 | 52.75 | 49.59 | 51.94 | 29.57 | 52.20 | 53.92 | 34.97 | ||
| InternVL3.5 | OW | 59.23 -4.9 | 68.18 -4.2 | 62.44 -5.1 | 65.78 -5.1 | 33.90 -5.1 | 73.33 -3.5 | 76.66 -3.7 | 40.07 -1.4 | |
| Model | ||||||||
|---|---|---|---|---|---|---|---|---|
| InternVL3.5 | 0.9 | 0.8 | 0.8 | 0.8 | 0.8 | 0.7 | 1.2 | 2.5 |
| GPT-4o | 5.3 | 4.0 | 3.4 | 8.7 | 6.0 | 5.6 | 9.9 | 16.4 |
| Qwen3-VL | 10.5 | 9.3 | 8.7 | 13.8 | 10.9 | 10.6 | 14.8 | 24.2 |
| LLaVA-OV | 26.8 | 24.1 | 23.6 | 28.1 | 29.0 | 28.8 | 31.2 | 42.5 |
| ImGT | ReGT | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Distractors | ||||||||||
| GPT-4o | ImGT + random | 99.62 ±0.07 | 90.95 ±0.05 | 89.49 ±0.07 | 99.67 ±0.07 | 0.09 ±0.18 | 91.85 ±0.10 | 99.75 ±0.11 | 0.00 ±0.00 | |
| ImGT + confEVA(ImGT) | 90.66 ±0.22 | 85.12 ±0.21 | 84.27 ±0.26 | 93.87 ±0.29 | 0.00 ±0.00 | 84.31 ±0.39 | 91.56 ±0.42 | 0.00 ±0.00 | ||
| ReGT + confEVA(ReGT) | 51.35 ±0.21 | 84.26 ±0.37 | 92.45 ±0.36 | 93.90 ±0.31 | 79.66 ±1.50 | 70.09 ±0.89 | 70.78 ±0.83 | 62.14 ±3.36 | ||
| ImGT + ReGT + random | 91.77 ±0.26 | 94.89 ±0.13 | 94.79 ±0.16 | 99.67 ±0.07 | 51.97 ±1.35 | 94.31 ±0.27 | 99.80 ±0.09 | 30.56 ±2.97 | ||
| ImGT + ReGT + confEVA(ImGT,ReGT) | 88.31 ±0.29 | 92.52 ±0.15 | 91.59 ±0.20 | 96.36 ±0.15 | 49.73 ±1.27 | 92.86 ±0.31 | 98.16 ±0.20 | 31.24 ±3.30 | ||
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Vision Backbone | Lang. Encoder | Training Strategy |
|---|---|---|---|
| PaliGemma2-mix-28B/448 | SigLIP so400M | Gemma2 27B | Joint multimodal pretraining + finetuning |
| LLaVA-OneVision-72B-Chat | SigLIP so400M | Qwen2 72B | Projector-based alignment + supervised finetuning |
| InternVL3.5-38B | InternViT 6B | Qwen3 32B | Progressive multimodal scaling + instruction tuning |
| Qwen3-VL-235B-A22B-Inst | SigLIP 2 so400M | Qwen3 MoE | Large-scale multimodal pretraining + instruction tuning |
| GPT-4o-2024-08-06 | (undisclosed) | (undisclosed) | Proprietary multimodal pretraining + post-training alignment (RLHF) |
| ImGT | ReGT | WeaselGT | |||||||
| Model | Task | ||||||||
| ImGT | CW | 100.00 | 100.00 | 91.20 -8.8 | 70.44 -29.6 | 100.00 | 00.00 | ||
| LLaVA-OV | CW | 42.56 | 25.79 | 52.00 +9.4 | 28.93 +3.1 | 35.71 | 12.77 | ||
| MC | 82.17 | 45.91 | 90.01 +7.8 | 61.01 +15.1 | 59.82 | 63.83 | |||
| Intern-VL3.5 | CW | 63.68 | 44.65 | 72.96 +9.3 | 61.01 +16.4 | 60.71 | 61.70 | ||
| MC | 85.89 | 55.97 | 90.53 +4.6 | 75.47 +19.5 | 75.89 | 74.47 | |||
| weasel | mink | polecat | domestic ferret | |||||||||
| ImGT | WeaselGT | ImGT | WeaselGT | ImGT | WeaselGT | ImGT | WeaselGT | |||||
| ImGT | 100.00 | 84.62 -15.4 | 100.00 | 85.00 -15.0 | 100.00 | 57.14 -42.9 | 100.00 | 61.54 -38.5 | ||||
| LLaVA-OV | 40.00 | 76.92 +36.9 | 60.00 | 65.00 +5.0 | 0.00 | 0.00 +0.0 | 0.00 | 0.00 +0.0 | ||||
| Intern-VL3.5 | 30.00 | 65.38 +35.4 | 32.00 | 37.50 +5.5 | 0.00 | 0.00 +0.0 | 92.00 | 100.00 +8.0 | ||||
| Qwen3-VL | 36.00 | 76.92 +40.9 | 72.00 | 80.00 +8.0 | 0.00 | 3.57 +3.6 | 42.00 | 43.08 +1.1 | ||||
| GPT-4o | 46.00 | 100.00 +54.0 | 94.00 | 95.00 +1.0 | 10.00 | 17.86 +7.9 | 92.00 | 100.00 +8.0 | ||||
| ImGT | ReGT | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Batch | |||||||||||
| LLaVA-OV | 1 | 42.56 | 52.00 | 46.31 | 48.10 | 30.56 | 53.75 | 55.20 | 36.84 | 168 | |
| 5 | 34.88 | 42.56 | 37.78 | 40.19 | 16.67 | 41.67 | 43.44 | 21.05 | 233 | ||
| 10 | 27.20 | 35.36 | 30.97 | 32.28 | 19.44 | 32.92 | 34.39 | 15.79 | 258 | ||
| Qwen3-VL | 1 | 64.96 | 73.12 | 69.32 | 73.10 | 36.11 | 75.00 | 77.83 | 42.11 | 28 | |
| 5 | 63.52 | 72.00 | 70.45 | 74.37 | 36.11 | 70.42 | 72.85 | 42.11 | 44 | ||
| ImGT | ReGT | ||||
|---|---|---|---|---|---|
| Model | In-Batch Ordering | ||||
| Qwen3-VL | Random | 63.47 | 70.83 | 734 | |
| Same-Class | 76.96 | 78.19 | 470 | ||
| GPT-4o | Random | 75.78 | 80.82 | 324 | |
| Same-Class | 86.03 | 84.61 | 155 |
| ImGT | ReGT | |||||
|---|---|---|---|---|---|---|
| Img Pos. | Batch | |||||
| LLaVA-OV | 1st | 1 | 44.44 | 60.32 | 18 | |
| 1st | 5 | 42.86 -1.6 | 58.73 -1.6 | 20 +2 | ||
| 1st | 10 | 33.33 -11.1 | 52.38 -7.9 | 23 +5 | ||
| 5th | 1 | 55.56 | 69.84 | 11 | ||
| 5th | 5 | 26.98 -28.6 | 34.92 -34.9 | 28 +17 |
| ImGT | ReGT | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Response | ||||||||||
| LLaVA-OV | ID | 7.68 | 14.24 | 9.38 | 9.81 | 5.56 | 9.58 | 9.95 | 5.26 | 0 | |
| Class name | 42.56 +34.9 | 52.00 +37.8 | 46.31 | 48.10 | 30.56 | 53.75 | 55.20 | 36.84 | 168 | ||
| InternVL3.5 | ID | 57.60 | 67.20 | 62.22 | 64.24 | 44.44 | 70.00 | 72.85 | 36.84 | 6 | |
| Class name | 63.68 +6.1 | 72.96 +5.8 | 67.05 | 69.30 | 47.22 | 77.92 | 80.54 | 47.37 | 6 | ||
| Qwen3-VL | ID | 48.00 | 57.28 | 54.55 | 56.96 | 33.33 | 55.42 | 58.37 | 21.05 | 0 | |
| OOP | Correctly Mapped | ||||
|---|---|---|---|---|---|
| # | % | # | % | ||
| InternVL3.5 | 271 | 0.87 | 83 | 30.63 | |
| 145 | 0.80 | 34 | 23.45 | ||
| 130 | 0.80 | 34 | 26.15 | ||
| 15 | 0.79 | 0 | 0.00 | ||
| 92 | 0.78 | 15 | 16.30 | ||
| Model | Partial | ImageNet | Abstain | Wrong |
|---|---|---|---|---|
| InternVL3.5 | 21.77 | 0.0 | 1.11 | 77.12 |
| GPT-4o | 35.34 | 3.22 | 6.38 | 55.07 |
| Qwen3-VL | 26.81 | 1.67 | 0.0 | 71.52 |
| LLaVA-OV | 36.29 | 2.59 | 0.04 | 61.08 |
| ImGT | ReGT | ||||||||||
| Model | Batch | ||||||||||
| GPT-4o | 10 | 74.69 ±0.19 | 81.32 ±0.18 | 79.87 ±0.19 | 84.49 ±0.21 | 39.25 ±0.82 | 80.89 ±0.35 | 84.72 ±0.36 | 36.33 ±0.58 | 30.13 ±1.11 | |
| ImGT | ReGT | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Emb. Space | ||||||||||
| PaliGemma 2 | Sentnce-BERT | 30.60 | 41.41 | 32.58 | 34.01 | 20.38 | 48.24 | 49.58 | 34.88 | |
| Sentence-BERT † | 31.89 | 42.81 | 33.86 | 35.32 | 21.44 | 49.98 | 51.33 | 36.55 | ||
| Qwen3-Embedding-8B | 33.24 | 43.53 | 35.49 | 37.05 | 22.23 | 49.37 | 50.99 | 33.21 | ||
| Qwen3-Embedding-8B † | 36.52 | 47.51 | 39.05 | 40.71 | 24.92 | 54.45 | 56.16 | 37.38 | ||
| SigLIP 2 | 34.56 | 45.04 | 36.45 | 38.23 | 21.28 | 51.90 | 53.51 | 35.81 | ||
| ImGT | ReGT | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Distractors | ||||||||||
| LLaVA-OV | ImGT + random | 99.28 ±0.15 | 90.75 ±0.11 | 89.36 ±0.13 | 99.52 ±0.14 | 0.18 ±0.25 | 91.52 ±0.18 | 99.39 ±0.20 | 0.00 ±0.00 | |
| ImGT + confEVA(ImGT) | 79.93 ±0.32 | 76.02 ±0.30 | 73.13 ±0.45 | 81.46 ±0.50 | 0.00 ±0.00 | 76.96 ±0.47 | 83.58 ±0.51 | 0.00 ±0.00 | ||
| ImGT + confBERT(ImGT) | 88.59 ±0.29 | 82.65 ±0.26 | 79.66 ±0.42 | 88.73 ±0.46 | 0.00 ±0.00 | 84.65 ±0.32 | 91.93 ±0.35 | 0.00 ±0.00 | ||
| ReGT + confEVA(ReGT) | 44.28 ±0.26 | 78.36 ±0.39 | 81.42 ±0.43 | 81.49 ±0.48 | 80.83 ±1.40 | 70.90 ±0.69 | 71.77 ±0.64 | 60.78 ±3.86 | ||
| ReGT + confBERT(ReGT) | 49.19 ±0.25 | 84.92 ±0.42 | 89.08 ±0.42 | 89.21 ±0.47 | 87.99 ±1.24 | 76.75 ±0.82 | 76.57 ±0.91 | 78.78 ±2.66 | ||
| GPT-4o Prompt |
|---|
| You are an image classifier. You will receive up to 50 images in order (image "1" = first, "2" = second, etc.). You are also provided with a list of class names: {class_list}. Classification Rules: - For each image, return the single class name that best represents the main subject of the image. - Choose only one class per image - the most relevant or dominant one. - Only return classes from the provided list. Output Rules: - Return exactly one output per image. - Each output must be only a single class name (no separators or lists). - Do not include explanations, confidence scores, or extra text. |
| LLaVA-OV & InternVL3.5 & Qwen3-VL Prompt |
| You are an image classifier. You will receive up to 50 images in order (image "1" = first, "2" = second, etc.). You are also provided with a list of class names: {class_list}. Your output will be automatically structured as JSON with keys "1", "2", "3", etc. corresponding to the order of images in the request. Each value should be the predicted class name for that image. Classification Rules: - For each image, return the class name only from the provided list. - Only return classes from the provided list. Output Rules: - Return exactly one JSON key per image ("1", "2", "3", etc.). - Each value must be only class names. - Do not include explanations, confidence scores, or extra text. |
| GPT-4o Prompt |
|---|
| You are an open-set fine-grained image classifier. You will receive up to 50 images in order (image "1" = first, "2" = second, etc.). Classification Rules: - For each image, identify the dominant object. - Return the most fine-grained, specific label that accurately describes that object (e.g., "golden retriever puppy", "1950s red convertible", "blue morpho butterfly", "ceramic coffee mug with floral pattern"). - Use natural-language labels that reflect detailed visual distinctions such as species, make/model, style, color, or material. - Avoid generic terms like "dog", "car", or "bird" when a more specific subtype or description is visually inferable. - If the dominant object cannot be clearly identified, return a concise descriptive label of its appearance (e.g., "abstract metal sculpture", "blurry human silhouette"). - Focus only on the dominant object, even if multiple are present. Output Rules: - Return exactly one output per image. - The output must contain only the final label (no punctuation beyond normal text, no explanations, confidence scores, or extra text). |
| LLaVA-OV & InternVL3.5 & Qwen3-VL Prompt |
| You are an open-set fine-grained image classifier. You will receive up to 50 images in order (image "1" = first, "2" = second, etc.). Your output will be automatically structured as JSON with keys "1", "2", "3", etc. corresponding to the order of images in the request. Each value should be the predicted label for that image. Classification Rules: - For each image, identify the dominant object. - Return the most fine-grained, specific label that accurately describes that object (e.g., "golden retriever puppy", "1950s red convertible", "blue morpho butterfly", "ceramic coffee mug with floral pattern"). - Use natural-language labels that reflect detailed visual distinctions such as species, make/model, style, color, or material. - Avoid generic terms like "dog", "car", or "bird" when a more specific subtype or description is visually inferable. - If the dominant object cannot be clearly identified, return a concise descriptive label of its appearance (e.g., "abstract metal sculpture", "blurry human silhouette"). - Focus only on the dominant object, even if multiple are present. Output Rules: - Return exactly one output per image. - The output must contain only the final label (no punctuation beyond normal text, no explanations, confidence scores, or extra text). |
| LLaVA-OV & InternVL3.5 & Qwen3-VL & GPT-4o Prompt |
| You are an image classifier. You will receive one image. You are also provided with four multiple-choice options (A, B, C, D). What is the main object in this image? {dynamic_choices} Classification Rules: - For the image, return the letter (A, B, C, or D) that corresponds to the correct option. - Only return one letter. Output Rules: - Return exactly one letter (A, B, C, or D). - Do not include explanations, the class name, or extra text. - Your answer must be only the letter. |