cs.CVOct 8, 2026

ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification

Authors: Mohammad Zare, Pirooz Shamsinejadbabaki

Organizations: Artificial Intelligence Lab at AriooBarzan Engineering Team Shiraz, Iran · Department of Computer Engineering and Information Technology Shiraz University of Technology, Shiraz, Iran

Abstract

Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.

Explore similar work

CardsList
  1. Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

    Mar 25, 2026Dipam Goswami, Simone Magistri, Gido M. van de Ven +4Prototype-Based LearningVLM Adaptation

  2. Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

    May 21, 2026David Méndez, Roberto Confalonieri, Natalia Díaz RodríguezCross-Modal LearningVLM Adaptation

  3. OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification

    Aug 30, 2026Ilán Carretero, Gustavo Jesús Angulo, Rocío del Amor +1Prototype-Based ClassificationFine-Grained Image Classification