ICCV Workshops (NIVT) 2023
Anindya Mondal* β’
Sauradip Nag* β’
Joaquin M. Prada β’
Xiatian Zhu β’
Anjan Dutta*
* Equal contribution / Corresponding authors
University of Surrey, United Kingdom
- Interactive Project Webpage: Available at https://mondalanindya.github.io/MSQNet/.
- ICCV 2023 Workshops: MSQNet accepted to the NIVT workshop at ICCV 2023.

Figure 1: Large action variation across diverse actors (animals and humans). MSQNet eliminates the need for actor pose estimation, offering an actor-agnostic framework.
Existing action recognition methods are typically actor-specific due to topological and morphological differences across actors (e.g., humans vs. quadrupeds, birds, reptiles). This typically demands specialized pose estimation models and limits generalization.
To solve this, we propose actor-agnostic multimodal multi-label action recognition and formulate MSQNet (Multimodal Semantic Query Network):
- Actor-Agnostic: Formulates multi-label video classification in a DETR-style detection framework without requiring actor pose or bounding boxes.
- Multimodal Query Formulation: Fuses CLIP text representations of class names with CLIP video features to formulate rich action queries.
- State-of-the-Art Performance: Outperforms prior actor-specific models on Animal Kingdom (+47.85% mAP boost), Charades, Thumos14, Hockey, and HMDB51, with strong zero-shot transfer capabilities.

Figure 2: Architecture of MSQNet comprising a Spatio-Temporal Video Encoder, Multimodal Query Encoder, and Multimodal Transformer Decoder.
| Dataset | Metric | Prior Art | Prior Score | MSQNet (Ours) | Improvement |
|---|---|---|---|---|---|
| Animal Kingdom | mAP | CARe (ICCV '21) | 25.25% | 73.10% | +47.85% |
| Charades | mAP | ActionCLIP | 44.30% | 47.57% | +3.27% |
| Thumos 14 | Accuracy | BMN | 62.12% | 83.16% | +21.04% |
| Hockey | Multilabel Acc | AFAC | 96.30% | 96.95% | +0.65% |
| HMDB51 | Accuracy | VideoMAE V2-g | 88.10% | 93.25% | +5.15% |
| Split | Method | Thumos 14 (Acc) | Charades (mAP) | HMDB51 (Acc) |
|---|---|---|---|---|
| Reported Prior SOTA | VideoCOCA / BIKE | - | 25.80% | 61.40% |
| 50% Seen Split | MSQNet (Full Model) | 63.98% | 30.91% | 59.24% |
| 75% Seen Split | MSQNet (Full Model) | 75.33% | 35.59% | 69.43% |
git clone https://github.com/mondalanindya/MSQNet.git
cd MSQNet
pip install -r requirements.txtconda env create -f environment.yml
conda activate msqnetpip install -e .Run the included verification script to confirm dependencies, model modules, and forward passes:
python verify_environment.pyDatasets can be placed under ./datasets or referenced directly using --data_dir <path> or the environment variable MSQNET_DATA_DIR.
Expected folder hierarchy:
datasets/
βββ AnimalKingdom/
β βββ action_recognition/
β βββ annotation/
β β βββ train_light.csv
β β βββ val_light.csv
β βββ dataset/
β βββ image/
β βββ [video_id_001]/
β β βββ 00001.jpg
β β βββ 00002.jpg
β β βββ ...
βββ Charades/
β βββ Charades_v1_train.csv
β βββ Charades_v1_480/
β βββ ...
βββ Hockey/
β βββ period1-gray/
β βββ ...
βββ THUMOS14/
β βββ annotation/
β βββ ...
βββ Volleyball/
python multi-label-action-main/utility/lighter_annotations.py \
--input /path/to/train.csv \
--output ./datasets/AnimalKingdom/action_recognition/annotation/train_light.csv
python multi-label-action-main/utility/lighter_annotations.py \
--input /path/to/val.csv \
--output ./datasets/AnimalKingdom/action_recognition/annotation/val_light.csvpython multi-label-action-main/utility/extract_frames.py \
--video_dir /path/to/videos \
--pattern "*.avi" \
--output_dir ./datasets/HockeyTo test the pipeline without downloading large datasets:
python multi-label-action-main/utility/create_dummy_dataset.py --output ./datasetsYou can train and evaluate directly from the root using run.py or through the scripts in scripts/.
# Single GPU training
python run.py \
--dataset animalkingdom \
--model msqnet \
--data_dir ./datasets \
--batch_size 16 \
--epochs 100 \
--total_length 16 \
--train True
# Multi-GPU Distributed (DDP) training
python multi-label-action-main/dist_main.py \
--dataset animalkingdom \
--model msqnet \
--data_dir ./datasets \
--batch_size 8 \
--total_length 16 \
--distributed TrueOr run the bash / batch scripts:
# Linux / macOS
bash scripts/train_msqnet.sh animalkingdom msqnet ./datasets 16 100 16
# Windows
scripts\train_msqnet.bat animalkingdom msqnet ./datasetspython run.py \
--dataset animalkingdom \
--model msqnet \
--data_dir ./datasets \
--checkpoint ./checkpoints/msqnet_msqnet_animalkingdom.pth \
--total_length 16 \
--train FalseOr run the evaluation scripts:
# Linux / macOS
bash scripts/eval_msqnet.sh ./checkpoints/msqnet_msqnet_animalkingdom.pth animalkingdom
# Windows
scripts\eval_msqnet.bat ./checkpoints/msqnet_msqnet_animalkingdom.pth animalkingdom

Video Demonstration: MSQNet multi-label temporal action prediction across unconstrained video snippets.

Attention Rollouts: GradCAM heatmaps showing how multimodal queries focus attention onto active bodies and interacting regions.

t-SNE Embeddings: Action class clusters before and after the multimodal transformer decoder on Animal Kingdom and Charades.
A self-contained, responsive academic project page is provided in index.html (and docs/index.html).
- Live URL: https://mondalanindya.github.io/MSQNet/
- Open
index.htmllocally in any browser for an interactive experience.
If you find our work useful, please consider citing:
@InProceedings{Mondal_2023_ICCV,
author = {Mondal, Anindya and Nag, Sauradip and Prada, Joaquin M and Zhu, Xiatian and Dutta, Anjan},
title = {Actor-Agnostic Multi-Label Action Recognition with Multi-Modal Query},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops},
month = {October},
year = {2023},
pages = {784-794}
}This repository is licensed under the MIT License.