Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AnimateProject: AnimateDiff + QR (HLG) ControlNet + Tile ControlNet / IP-Adapter video generation

This project generates short videos (multi-frame image sequences) using AnimateDiff (SD 1.5) and (Multi)ControlNet to achieve:

  • QR control / scannability via your trained HLG QR ControlNet (local checkpoint)
  • Style / identity reference using Tile ControlNet (reference image as control input) and/or IP-Adapter (reference conditioning)
  • Per-frame decay (frame decay) to gradually weaken Tile ControlNet or IP-Adapter influence over time and reduce over-conditioning in later frames

The main entry point is my_script.py. The script make_gif_from_folder.py can also build a GIF from a folder of frames.


Directory structure (high level)

  • my_script.py: Main generation script (AnimateDiff + QR ControlNet + optional Tile/IP-Adapter + HLG conditioning).
  • my_script2.py: Older variant (no HLG conditioning flow; different defaults).
  • animatediff_controlnet.py: Custom pipeline implementation (currently my_script.py imports the diffusers built-in AnimateDiffControlNetPipeline; you can switch if needed).
  • make_gif_from_folder.py: Create a GIF from images in a folder (sorted by filename / numeric index).
  • ref/: Reference images (for example boy.png, penguin.png).
  • code/: Example QR / intermediate assets (for example ternary_output.png).
  • *.gif, *.png: Mostly experiment outputs and debug images.

Prerequisites

Hardware

  • An NVIDIA GPU (CUDA available) is strongly recommended for reasonable inference speed.
  • CPU can run but is typically very slow and may run out of memory.

External dependency (another repo / folder)

HLG conditioning in my_script.py imports the HLG extractor from:

  • /scratch1/intern/si2025_2/Hsueh/QRcode/src/hlg.py

If you run on a different machine/environment, ensure that repo exists and is readable, or update hlg_repo_root inside my_script.py.

Models and weights

By default, my_script.py loads:

  • SD 1.5 base model: --sd_model_id (default SG161222/Realistic_Vision_V6.0_B1_noVAE)
  • AnimateDiff motion adapter: --motion_adapter_id (default guoyww/animatediff-motion-adapter-v1-5-2)
  • QR HLG ControlNet (local checkpoint): --qr_controlnet_path
    • Default: /scratch1/intern/si2025_2/Hsueh/QRcode/outputs/hlg_controlnet/checkpoint-54000
  • Optional Tile ControlNet: --tile_controlnet (default lllyasviel/control_v11f1e_sd15_tile)
  • Optional IP-Adapter: --ip_adapter_repo + --ip_adapter_weight (default h94/IP-Adapter / ip-adapter-plus_sd15.bin)

Note: downloading from Hugging Face may require network access and proper cache/auth setup.


Setup (suggested)

This repo includes a requirements.txt for the Python packages (excluding PyTorch). The recommended flow is to use a conda env named animateqr, then install dependencies with pip.

conda create -n animateqr python=3.10 -y
conda activate animateqr

# Install PyTorch that matches your CUDA/driver setup (choose ONE approach).
# Option A (conda):
#   conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia -y
# Option B (pip):
#   pip install torch

pip install -U pip
pip install -r requirements.txt

# Optional: memory efficient attention (only if your environment supports it)
# pip install xformers

Tested environment (this machine)

The existing conda env animateqr on this machine contains (high level):

  • Python 3.10.19
  • torch 2.11.0.dev20251227+cu128, torchvision 0.25.0.dev20251227+cu128, torchaudio 2.10.0.dev20251227+cu128
  • diffusers 0.32.0
  • transformers 4.57.3
  • accelerate 1.12.0
  • safetensors 0.7.0
  • numpy 2.2.6
  • pillow 12.0.0

Common issue: transformers cannot import due to huggingface-hub 1.x

If you see an error like:

  • ImportError: huggingface-hub>=0.34.0,<1.0 is required ... but found huggingface-hub==1.x

Fix it by downgrading huggingface-hub:

conda activate animateqr
pip install -U "huggingface-hub<1.0"

If you see ModuleNotFoundError: diffusers.pipelines.animatediff ..., your diffusers version is likely too old. Upgrade:

pip install -U diffusers

Quick start

1) Run with defaults (QR (HLG) + Tile reference)

python my_script.py \
  --qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
  --ref_image /scratch1/intern/si2025_2/arthur/AnimateProject/ref/boy.png \
  --output my_run

The output filename includes a timestamp and key parameters to avoid overwriting. Example shape:

  • <YYYYMMDD_HHMMSS>__<prefix>__mode-<...>__qr<...>__tile<...>__ipa<...>__F<...>__S<...>__G<...>.gif

2) QR ControlNet only (no reference injection)

python my_script.py \
  --ref_injection none \
  --qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
  --output my_run_qr_only

3) Tile ControlNet plus IP-Adapter (double reference constraints)

python my_script.py \
  --ref_injection tile_controlnet \
  --enable_ipadapter \
  --ref_image /scratch1/intern/si2025_2/arthur/AnimateProject/ref/boy.png \
  --qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
  --tile_scale 0.8 \
  --ipadapter_scale 0.6 \
  --output my_run_tile_plus_ipa

4) Save HLG conditioning debug images (what is actually fed into QR ControlNet)

python my_script.py \
  --save_hlg_debug \
  --save_hlg_debug_split \
  --output my_run_debug_cond

Core concepts and flags (based on my_script.py)

Image size constraints (important)

  • --width, --height: default 512x512
  • When --qr_conditioning hlg (default):
    • Must be square (width == height)
    • Must be divisible by --hlg_a (default hlg_a=16)

If these are not satisfied, my_script.py raises an error early.

QR conditioning: --qr_conditioning

  • hlg (default): extract a 3-scale HLG map via extract_hlg_map(...) from Hsueh/QRcode, then feed it to the QR ControlNet
  • raw: feed the input --qr_image directly as conditioning (legacy / comparison)

Reference injection mode: --ref_injection

Modes:

  • none: no reference image
  • tile_controlnet (default): load Tile ControlNet and feed --ref_image as control image
  • ipadapter: use IP-Adapter as the main reference mechanism (still requires --enable_ipadapter to actually enable it)

If you want Tile and IP-Adapter at the same time, use:

  • --ref_injection tile_controlnet plus --enable_ipadapter

Control strengths

  • --qr_scale: QR (HLG) ControlNet strength (higher means stronger QR conditioning)
  • --tile_scale: Tile ControlNet strength (higher means stronger adherence to reference texture/structure)
  • --ipadapter_scale: IP-Adapter strength (higher means stronger reference influence)

Per-frame decay (frame decay)

You can specify a start scale and an end scale, with interpolation across frames:

  • Tile: --tile_scale_end <float>
  • IP-Adapter: --ipadapter_scale_end <float>
  • Schedule: --frame_decay_schedule {linear,cosine,power}
  • Exponent (power only): --frame_decay_power <float>

Example: decay Tile from 0.8 down to 0.2:

python my_script.py \
  --tile_scale 0.8 \
  --tile_scale_end 0.2 \
  --frame_decay_schedule power \
  --frame_decay_power 1.5 \
  --output my_run_tile_decay

Quality and speed

  • --num_frames: number of frames (common motion adapters support up to about 32; the script checks via infer_motion_max_frames(...))
  • --num_inference_steps: diffusion steps (more is slower, often more stable)
  • --guidance_scale: classifier-free guidance scale (CFG)
  • --seed: random seed
  • --decode_chunk_size: VAE decode chunk size (VRAM saving)

VRAM and performance knobs

  • --offload {model,sequential,none}:
    • model (default): balanced in many cases
    • sequential: lowest VRAM usage, usually slowest
    • none: fastest but uses the most VRAM
  • --attention_slicing: lower VRAM but slightly slower
  • --xformers: enable memory efficient attention if available

Outputs and filename conventions

my_script.py encodes key parameters in the output filename to make experiment tracking easier. The name includes:

  • mode: mode-<ref_injection>
  • scales: qr<...> tile<...> ipa<...>
  • optional per-frame decay: tileE<...> ipaE<...> decay-<schedule> (and p<...> when schedule is power)
  • F frames, S steps, G guidance

If you enable --save_hlg_debug, the script also saves the actual conditioning image used for the QR ControlNet:

  • __cond-hlg.png or __cond-raw.png
  • with --save_hlg_debug_split, it also saves __r.png, __g.png, __b.png

Utility: make a GIF from a folder

make_gif_from_folder.py loads png/jpg/webp files from --input_dir, sorts them by filename (natural numeric order), and writes a GIF:

python make_gif_from_folder.py \
  --input_dir /path/to/frames_dir \
  --output_gif /path/to/out.gif \
  --fps 8

Troubleshooting

1) Failed to import HLG extractor ...

This means the script cannot import /scratch1/intern/si2025_2/Hsueh/QRcode/src/hlg.py.

  • Verify the repo exists and is readable
  • Or update hlg_repo_root inside my_script.py to match your environment

2) HLG errors: HLG expects square images / must be divisible by --hlg_a

Set --width and --height to the same value, and ensure the value is divisible by --hlg_a (for example 512 and 16).

3) num_frames exceeds motion module limit

Lower --num_frames to 16 or 32, or use a motion adapter that supports a longer sequence length.

4) CUDA out of memory (OOM)

Try, in order:

  • Lower --width/--height (for example 512 -> 448/384)
  • Lower --num_frames
  • Enable --attention_slicing
  • Set --offload sequential
  • Lower --num_inference_steps

Optional: use the custom pipeline in animatediff_controlnet.py

Currently, my_script.py imports:

  • from diffusers.pipelines.animatediff import AnimateDiffControlNetPipeline

If you want to use the custom pipeline implementation in this repo, change the import to:

from animatediff_controlnet import AnimateDiffControlNetPipeline

Then verify that the custom pipeline __call__ signature matches how my_script.py passes kwargs.

MatrixQR

About

Video generation that keeps a QR code scannable while it is stylised, AnimateDiff with QR and Tile ControlNet

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages