cs.CVSep 28, 2026

AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection

Authors: Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, Luc Van Gool

Organizations: INSAIT, Sofia University “St. Kliment Ohridski”

Abstract

Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we propose AHMAD, a simple yet effective framework for generalist multitask learning that integrates different key vision tasks: semantic segmentation, instance segmentation, depth estimation, keypoint detection, and object detection. Our approach incorporates these five tasks into a unified structure: a shared encoder-decoder with several lightweight task-specific projectors. Under the multitask learning paradigm, we observed a complementary performance gain, achieving a state-of-the-art PQ of 53.1 and an mIoU of 66.5 for COCO-val panoptic and semantic segmentation, respectively. Additionally, for top-down keypoint detection, which typically incurs high computational overhead due to multiple forward passes, we introduce a knowledge distillation-based method that enables a single forward pass over the entire image, greatly improving efficiency. Ultimately, our model delivers a lightweight yet effective generalist multitask learning framework, demonstrating strong performance across five vision tasks.

Figures & tables

Explore similar work

CardsList
  1. GKDT: General Keypoint Detection Transformer

    Jul 1, 2026Changsheng Lu, Yuxin Chen, Haokun Gui +5Keypoint DetectionOpen-Vocabulary Object Detection

  2. DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

    Jul 17, 2026Jehun Kang, Jungha Wang, Youngjun Hwang +1Dense PredictionDepth Estimation

  3. MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

    Jun 29, 2026Shuo Zhou, Zhaoxin Li, Xiujuan ChaiDense PredictionMulti--Task Learning