Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
Published in Neural Computing and Applications, 2026
Keywords:
Contrastive learning, Data imbalance, NR-VQA, Self-supervised learning, Video classification, ViViT
Contributions:
- We propose the Combined-SSL framework, which is a theoretical contribution that leverages the reciprocal relationship between the VQA task and Contrastive Learning to classify videos objectively.
- We develop the SSL-V3 model to implement the Combined-SSL mechanism.
- We present a hierarchical VQA head designed to regress video quality score, incorporating Sequence Score Regressor and Video Score Regressor modules.
- We design a new Loss function to address class imbalance issues and control training at batch and subject levels.
Production: SSL-V3, Combined-SSL mechanism, hierarchical VQA head, new loss.
Used Framework: 
BibTex:
@article{sun2026contrastive,
title={Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition},
author={Sun, Jian and Mahoor, Mohammad},
journal={Neural Computing and Applications},
volume={38},
number={5},
pages={107},
year={2026},
publisher={Springer}
}
Recommended citation: Sun, J., Mahoor, M. Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition. Neural Comput & Applic 38, 107 (2026).
Download Paper | Download Slides
