Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

Yachun Mi1,2,Yu Li1,Yanting Li1,Chen Hui3,Tong Zhang1,Zhixuan Li2,Chenyue Song1,Wei Yang Bryan Lim2,Shaohui Liu1†
1Harbin Institute of Technology, 2Nanyang Technological University, 3Nanjing University of Information Science and Technology
Corresponding author.
International Conference on Machine Learning (ICML), 2026
Overview of the Q-CLIP framework.

Q-CLIP turns video quality assessment into cross-modal quality matching, adapting frozen vision-language encoders with a compact Shared Cross-Modal Adapter and learnable quality prompts.

Abstract

Q-CLIP is a vision-language framework for video quality assessment that preserves the broad semantic knowledge of pretrained encoders while adapting them efficiently to visual quality. It introduces shared cross-modal adapters for both image and text branches, learnable quality-level prompts, and frame-difference-guided sampling. The resulting model trains only a small fraction of its parameters and achieves strong performance across in-domain and cross-dataset VQA benchmarks.

Shared Cross-Modal Adaptation

Architecture of the Shared Cross-Modal Adapter.

E-SCMA shares lightweight parameters across encoder layers, while P-SCMA adapts the projection heads. Together they align visual evidence with ordered quality concepts without fully fine-tuning the foundation model.

Quality-Aware Frame Sampling

Comparison of video frame-sampling strategies.

Frame-difference-guided sampling selects frames from intervals with diverse temporal change, giving the model more informative quality observations than uniform or random sampling.

What the Model Sees

Attention visualizations comparing Q-CLIP and CLIP.

Attention maps show that Q-CLIP focuses more consistently on regions related to perceptual degradation, while retaining the semantic awareness inherited from CLIP.

BibTeX

@inproceedings{mi2026qclip,
  title={{Q-CLIP}: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation},
  author={Mi, Yachun and Li, Yu and Li, Yanting and Hui, Chen and Zhang, Tong and Li, Zhixuan and Song, Chenyue and Lim, Wei Yang Bryan and Liu, Shaohui},
  booktitle={International Conference on Machine Learning},
  year={2026}
}