This project generates short videos (multi-frame image sequences) using AnimateDiff (SD 1.5) and (Multi)ControlNet to achieve:
- QR control / scannability via your trained HLG QR ControlNet (local checkpoint)
- Style / identity reference using Tile ControlNet (reference image as control input) and/or IP-Adapter (reference conditioning)
- Per-frame decay (frame decay) to gradually weaken Tile ControlNet or IP-Adapter influence over time and reduce over-conditioning in later frames
The main entry point is my_script.py. The script make_gif_from_folder.py can also build a GIF from a folder of frames.
my_script.py: Main generation script (AnimateDiff + QR ControlNet + optional Tile/IP-Adapter + HLG conditioning).my_script2.py: Older variant (no HLG conditioning flow; different defaults).animatediff_controlnet.py: Custom pipeline implementation (currentlymy_script.pyimports the diffusers built-inAnimateDiffControlNetPipeline; you can switch if needed).make_gif_from_folder.py: Create a GIF from images in a folder (sorted by filename / numeric index).ref/: Reference images (for exampleboy.png,penguin.png).code/: Example QR / intermediate assets (for exampleternary_output.png).*.gif,*.png: Mostly experiment outputs and debug images.
- An NVIDIA GPU (CUDA available) is strongly recommended for reasonable inference speed.
- CPU can run but is typically very slow and may run out of memory.
HLG conditioning in my_script.py imports the HLG extractor from:
/scratch1/intern/si2025_2/Hsueh/QRcode/src/hlg.py
If you run on a different machine/environment, ensure that repo exists and is readable, or update hlg_repo_root inside my_script.py.
By default, my_script.py loads:
- SD 1.5 base model:
--sd_model_id(defaultSG161222/Realistic_Vision_V6.0_B1_noVAE) - AnimateDiff motion adapter:
--motion_adapter_id(defaultguoyww/animatediff-motion-adapter-v1-5-2) - QR HLG ControlNet (local checkpoint):
--qr_controlnet_path- Default:
/scratch1/intern/si2025_2/Hsueh/QRcode/outputs/hlg_controlnet/checkpoint-54000
- Default:
- Optional Tile ControlNet:
--tile_controlnet(defaultlllyasviel/control_v11f1e_sd15_tile) - Optional IP-Adapter:
--ip_adapter_repo+--ip_adapter_weight(defaulth94/IP-Adapter/ip-adapter-plus_sd15.bin)
Note: downloading from Hugging Face may require network access and proper cache/auth setup.
This repo includes a requirements.txt for the Python packages (excluding PyTorch). The recommended flow is to use a conda env named animateqr, then install dependencies with pip.
conda create -n animateqr python=3.10 -y
conda activate animateqr
# Install PyTorch that matches your CUDA/driver setup (choose ONE approach).
# Option A (conda):
# conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia -y
# Option B (pip):
# pip install torch
pip install -U pip
pip install -r requirements.txt
# Optional: memory efficient attention (only if your environment supports it)
# pip install xformersThe existing conda env animateqr on this machine contains (high level):
- Python 3.10.19
- torch 2.11.0.dev20251227+cu128, torchvision 0.25.0.dev20251227+cu128, torchaudio 2.10.0.dev20251227+cu128
- diffusers 0.32.0
- transformers 4.57.3
- accelerate 1.12.0
- safetensors 0.7.0
- numpy 2.2.6
- pillow 12.0.0
If you see an error like:
ImportError: huggingface-hub>=0.34.0,<1.0 is required ... but found huggingface-hub==1.x
Fix it by downgrading huggingface-hub:
conda activate animateqr
pip install -U "huggingface-hub<1.0"If you see ModuleNotFoundError: diffusers.pipelines.animatediff ..., your diffusers version is likely too old. Upgrade:
pip install -U diffuserspython my_script.py \
--qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
--ref_image /scratch1/intern/si2025_2/arthur/AnimateProject/ref/boy.png \
--output my_runThe output filename includes a timestamp and key parameters to avoid overwriting. Example shape:
<YYYYMMDD_HHMMSS>__<prefix>__mode-<...>__qr<...>__tile<...>__ipa<...>__F<...>__S<...>__G<...>.gif
python my_script.py \
--ref_injection none \
--qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
--output my_run_qr_onlypython my_script.py \
--ref_injection tile_controlnet \
--enable_ipadapter \
--ref_image /scratch1/intern/si2025_2/arthur/AnimateProject/ref/boy.png \
--qr_image /scratch1/intern/si2025_2/arthur/AnimateProject/code/ternary_output.png \
--tile_scale 0.8 \
--ipadapter_scale 0.6 \
--output my_run_tile_plus_ipapython my_script.py \
--save_hlg_debug \
--save_hlg_debug_split \
--output my_run_debug_cond--width,--height: default 512x512- When
--qr_conditioning hlg(default):- Must be square (width == height)
- Must be divisible by
--hlg_a(defaulthlg_a=16)
If these are not satisfied, my_script.py raises an error early.
hlg(default): extract a 3-scale HLG map viaextract_hlg_map(...)fromHsueh/QRcode, then feed it to the QR ControlNetraw: feed the input--qr_imagedirectly as conditioning (legacy / comparison)
Modes:
none: no reference imagetile_controlnet(default): load Tile ControlNet and feed--ref_imageas control imageipadapter: use IP-Adapter as the main reference mechanism (still requires--enable_ipadapterto actually enable it)
If you want Tile and IP-Adapter at the same time, use:
--ref_injection tile_controlnetplus--enable_ipadapter
--qr_scale: QR (HLG) ControlNet strength (higher means stronger QR conditioning)--tile_scale: Tile ControlNet strength (higher means stronger adherence to reference texture/structure)--ipadapter_scale: IP-Adapter strength (higher means stronger reference influence)
You can specify a start scale and an end scale, with interpolation across frames:
- Tile:
--tile_scale_end <float> - IP-Adapter:
--ipadapter_scale_end <float> - Schedule:
--frame_decay_schedule {linear,cosine,power} - Exponent (power only):
--frame_decay_power <float>
Example: decay Tile from 0.8 down to 0.2:
python my_script.py \
--tile_scale 0.8 \
--tile_scale_end 0.2 \
--frame_decay_schedule power \
--frame_decay_power 1.5 \
--output my_run_tile_decay--num_frames: number of frames (common motion adapters support up to about 32; the script checks viainfer_motion_max_frames(...))--num_inference_steps: diffusion steps (more is slower, often more stable)--guidance_scale: classifier-free guidance scale (CFG)--seed: random seed--decode_chunk_size: VAE decode chunk size (VRAM saving)
--offload {model,sequential,none}:model(default): balanced in many casessequential: lowest VRAM usage, usually slowestnone: fastest but uses the most VRAM
--attention_slicing: lower VRAM but slightly slower--xformers: enable memory efficient attention if available
my_script.py encodes key parameters in the output filename to make experiment tracking easier. The name includes:
- mode:
mode-<ref_injection> - scales:
qr<...> tile<...> ipa<...> - optional per-frame decay:
tileE<...> ipaE<...> decay-<schedule>(andp<...>when schedule is power) Fframes,Ssteps,Gguidance
If you enable --save_hlg_debug, the script also saves the actual conditioning image used for the QR ControlNet:
__cond-hlg.pngor__cond-raw.png- with
--save_hlg_debug_split, it also saves__r.png,__g.png,__b.png
make_gif_from_folder.py loads png/jpg/webp files from --input_dir, sorts them by filename (natural numeric order), and writes a GIF:
python make_gif_from_folder.py \
--input_dir /path/to/frames_dir \
--output_gif /path/to/out.gif \
--fps 8This means the script cannot import /scratch1/intern/si2025_2/Hsueh/QRcode/src/hlg.py.
- Verify the repo exists and is readable
- Or update
hlg_repo_rootinsidemy_script.pyto match your environment
Set --width and --height to the same value, and ensure the value is divisible by --hlg_a (for example 512 and 16).
Lower --num_frames to 16 or 32, or use a motion adapter that supports a longer sequence length.
Try, in order:
- Lower
--width/--height(for example 512 -> 448/384) - Lower
--num_frames - Enable
--attention_slicing - Set
--offload sequential - Lower
--num_inference_steps
Currently, my_script.py imports:
from diffusers.pipelines.animatediff import AnimateDiffControlNetPipeline
If you want to use the custom pipeline implementation in this repo, change the import to:
from animatediff_controlnet import AnimateDiffControlNetPipelineThen verify that the custom pipeline __call__ signature matches how my_script.py passes kwargs.