Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition

Published in Neural Computing and Applications, 2026

Keywords:

Contrastive learning, Data imbalance, NR-VQA, Self-supervised learning, Video classification, ViViT

Contributions:

  • We propose the Combined-SSL framework, which is a theoretical contribution that leverages the reciprocal relationship between the VQA task and Contrastive Learning to classify videos objectively.
  • We develop the SSL-V3 model to implement the Combined-SSL mechanism.
  • We present a hierarchical VQA head designed to regress video quality score, incorporating Sequence Score Regressor and Video Score Regressor modules.
  • We design a new Loss function to address class imbalance issues and control training at batch and subject levels.

Production: SSL-V3, Combined-SSL mechanism, hierarchical VQA head, new loss.

Used Framework: PyTorch

BibTex:

@article{sun2026contrastive,
  title={Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition},
  author={Sun, Jian and Mahoor, Mohammad},
  journal={Neural Computing and Applications},
  volume={38},
  number={5},
  pages={107},
  year={2026},
  publisher={Springer}
}

Recommended citation: Sun, J., Mahoor, M. Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition. Neural Comput & Applic 38, 107 (2026).
Download Paper | Download Slides