Skip to content

Repository files navigation

MSQNet: Actor-Agnostic Multi-Label Action Recognition with Multi-Modal Query

ICCV Workshops (NIVT) 2023
Anindya Mondal* β€’ Sauradip Nag* β€’ Joaquin M. Prada β€’ Xiatian Zhu β€’ Anjan Dutta*
* Equal contribution / Corresponding authors
University of Surrey, United Kingdom

Project Page CVF Paper arXiv Poster Video Talk License Python PyTorch

Leaderboard on Papers With Code

PWC PWC PWC PWC PWC PWC PWC PWC


πŸ“Œ News


πŸ“– Overview

Actor Variations
Figure 1: Large action variation across diverse actors (animals and humans). MSQNet eliminates the need for actor pose estimation, offering an actor-agnostic framework.

Existing action recognition methods are typically actor-specific due to topological and morphological differences across actors (e.g., humans vs. quadrupeds, birds, reptiles). This typically demands specialized pose estimation models and limits generalization.

To solve this, we propose actor-agnostic multimodal multi-label action recognition and formulate MSQNet (Multimodal Semantic Query Network):

  1. Actor-Agnostic: Formulates multi-label video classification in a DETR-style detection framework without requiring actor pose or bounding boxes.
  2. Multimodal Query Formulation: Fuses CLIP text representations of class names with CLIP video features to formulate rich action queries.
  3. State-of-the-Art Performance: Outperforms prior actor-specific models on Animal Kingdom (+47.85% mAP boost), Charades, Thumos14, Hockey, and HMDB51, with strong zero-shot transfer capabilities.

MSQNet Pipeline
Figure 2: Architecture of MSQNet comprising a Spatio-Temporal Video Encoder, Multimodal Query Encoder, and Multimodal Transformer Decoder.


πŸ”¬ Benchmark Results

Supervised Action Recognition

Dataset Metric Prior Art Prior Score MSQNet (Ours) Improvement
Animal Kingdom mAP CARe (ICCV '21) 25.25% 73.10% +47.85%
Charades mAP ActionCLIP 44.30% 47.57% +3.27%
Thumos 14 Accuracy BMN 62.12% 83.16% +21.04%
Hockey Multilabel Acc AFAC 96.30% 96.95% +0.65%
HMDB51 Accuracy VideoMAE V2-g 88.10% 93.25% +5.15%

Zero-Shot Action Recognition

Split Method Thumos 14 (Acc) Charades (mAP) HMDB51 (Acc)
Reported Prior SOTA VideoCOCA / BIKE - 25.80% 61.40%
50% Seen Split MSQNet (Full Model) 63.98% 30.91% 59.24%
75% Seen Split MSQNet (Full Model) 75.33% 35.59% 69.43%

πŸ› οΈ Environment Setup

Option A: Using pip

git clone https://github.com/mondalanindya/MSQNet.git
cd MSQNet
pip install -r requirements.txt

Option B: Using conda

conda env create -f environment.yml
conda activate msqnet

Option C: Editable Package Install

pip install -e .

Verify Environment

Run the included verification script to confirm dependencies, model modules, and forward passes:

python verify_environment.py

πŸ“‚ Dataset Setup

Datasets can be placed under ./datasets or referenced directly using --data_dir <path> or the environment variable MSQNET_DATA_DIR.

Expected folder hierarchy:

datasets/
β”œβ”€β”€ AnimalKingdom/
β”‚   └── action_recognition/
β”‚       β”œβ”€β”€ annotation/
β”‚       β”‚   β”œβ”€β”€ train_light.csv
β”‚       β”‚   └── val_light.csv
β”‚       └── dataset/
β”‚           └── image/
β”‚               β”œβ”€β”€ [video_id_001]/
β”‚               β”‚   β”œβ”€β”€ 00001.jpg
β”‚               β”‚   β”œβ”€β”€ 00002.jpg
β”‚               β”‚   └── ...
β”œβ”€β”€ Charades/
β”‚   β”œβ”€β”€ Charades_v1_train.csv
β”‚   β”œβ”€β”€ Charades_v1_480/
β”‚   └── ...
β”œβ”€β”€ Hockey/
β”‚   β”œβ”€β”€ period1-gray/
β”‚   └── ...
β”œβ”€β”€ THUMOS14/
β”‚   β”œβ”€β”€ annotation/
β”‚   └── ...
└── Volleyball/

Generating Animal Kingdom Lighter Annotations

python multi-label-action-main/utility/lighter_annotations.py \
    --input /path/to/train.csv \
    --output ./datasets/AnimalKingdom/action_recognition/annotation/train_light.csv

python multi-label-action-main/utility/lighter_annotations.py \
    --input /path/to/val.csv \
    --output ./datasets/AnimalKingdom/action_recognition/annotation/val_light.csv

Extracting Video Frames

python multi-label-action-main/utility/extract_frames.py \
    --video_dir /path/to/videos \
    --pattern "*.avi" \
    --output_dir ./datasets/Hockey

Testing with a Synthetic Toy Dataset

To test the pipeline without downloading large datasets:

python multi-label-action-main/utility/create_dummy_dataset.py --output ./datasets

πŸš€ Training & Evaluation

You can train and evaluate directly from the root using run.py or through the scripts in scripts/.

1. Training MSQNet

# Single GPU training
python run.py \
    --dataset animalkingdom \
    --model msqnet \
    --data_dir ./datasets \
    --batch_size 16 \
    --epochs 100 \
    --total_length 16 \
    --train True

# Multi-GPU Distributed (DDP) training
python multi-label-action-main/dist_main.py \
    --dataset animalkingdom \
    --model msqnet \
    --data_dir ./datasets \
    --batch_size 8 \
    --total_length 16 \
    --distributed True

Or run the bash / batch scripts:

# Linux / macOS
bash scripts/train_msqnet.sh animalkingdom msqnet ./datasets 16 100 16

# Windows
scripts\train_msqnet.bat animalkingdom msqnet ./datasets

2. Evaluating a Checkpoint

python run.py \
    --dataset animalkingdom \
    --model msqnet \
    --data_dir ./datasets \
    --checkpoint ./checkpoints/msqnet_msqnet_animalkingdom.pth \
    --total_length 16 \
    --train False

Or run the evaluation scripts:

# Linux / macOS
bash scripts/eval_msqnet.sh ./checkpoints/msqnet_msqnet_animalkingdom.pth animalkingdom

# Windows
scripts\eval_msqnet.bat ./checkpoints/msqnet_msqnet_animalkingdom.pth animalkingdom

🎨 Qualitative Results & Visualizations

MSQNet Real-Time Predictions
Video Demonstration: MSQNet multi-label temporal action prediction across unconstrained video snippets.

GradCAM Attention Comparison
Attention Rollouts: GradCAM heatmaps showing how multimodal queries focus attention onto active bodies and interacting regions.

t-SNE Embeddings
t-SNE Embeddings: Action class clusters before and after the multimodal transformer decoder on Animal Kingdom and Charades.


🌐 Webpage

A self-contained, responsive academic project page is provided in index.html (and docs/index.html).


πŸ“š Citation

If you find our work useful, please consider citing:

@InProceedings{Mondal_2023_ICCV,
    author    = {Mondal, Anindya and Nag, Sauradip and Prada, Joaquin M and Zhu, Xiatian and Dutta, Anjan},
    title     = {Actor-Agnostic Multi-Label Action Recognition with Multi-Modal Query},
    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops},
    month     = {October},
    year      = {2023},
    pages     = {784-794}
}

πŸ“„ License

This repository is licensed under the MIT License.