AIREITER

Wan2.2 Animate Explained: Modes, Specs, API Pricing & Setup

Last Updated: 2026-08-14 01:20:49

Wan2.2 Animate is Alibaba's open-source 14B character-animation model that does two jobs in one framework: animate a still character from a driving video, or replace an actor in existing footage while keeping the scene intact. Released September 19, 2025 under Apache-2.0, it runs locally through ComfyUI or the official CLI. But clean output depends on matching FPS, pose alignment, and masking, not merely pressing generate.

What Is Wan2.2 Animate?

Wan2.2-Animate-14B is Wan-AI's MoE character-animation variant of the Wan2.2 I2V-A14B image-to-video model. Per the model card, roughly 14B parameters are active per denoising step (27B total across two experts). The model takes two inputs: a character image and a driving/reference video. Depending on the mode you select, it outputs the character performing the motions from the video, or the character replacing the original actor in the scene.

The official weights ship in BF16; an FP8-quantized version (the Wan2_2-Animate-14B_fp8 file used in the ComfyUI workflow) stores weights at 8-bit precision instead of 16-bit, halving weight memory and making the model viable on consumer GPUs.

Move Mode vs Mix Mode — Two Jobs, One Model

Wan2.2 Animate ships with two operating modes. ComfyUI calls them Move and Mix; the CLI and official docs call them animation and replacement. Both share the same pipeline (skeleton extraction, implicit facial-feature extraction, character image as visual reference), but they differ in what stays and what gets replaced.

Move mode (animation) transfers body motion and facial expressions from the driving video to your character image. The character image is the subject; the video provides the performance. A still portrait becomes a dancing avatar; a concept character inherits the clip's choreography. Body motion is driven by spatially aligned skeleton signals, and facial expressions are retargeted from the source actor.

Mix mode (replacement) does the inverse: it replaces the actor in an existing video with your character image while preserving the original scene. A dedicated Relighting LoRA adapts the inserted character to the source footage's lighting and color tone so the composite does not look pasted in.

AspectMove (Animation)Mix (Replacement)
What changesStill image gets animatedVideo actor gets replaced
What's preservedCharacter identity from imageScene, lighting, color from video
Driving inputMotion + expressions from videoPerformance from video
CLI flag--retarget_flag --use_flux--replace_flag --use_relighting_lora
Pose retargetingRecommended (handles body-proportion differences)Disabled by default (character-environment interaction risk)
Best forAnimating concept characters, avatarsActor replacement, branded insertion

System Requirements & Hardware

The official repository is Wan-Video/Wan2.2 on GitHub, requiring PyTorch ≥2.4.0. FlashAttention is needed for speed but can be installed last if it fails the first time. Because no official fixed VRAM minimum is published, treat the FP8 path plus model offload as the consumer-GPU route and verify the exact ComfyUI workflow's memory notes before downloading.

The model card's efficiency results are in figures rather than text. Single-GPU runs use model offload and dtype conversion. Community reports offer partial benchmarks: one r/comfyui user produced dance animations on an RTX 3070 using masking and inpainting, and a linked variant advertises 8 GB VRAM editing. These are partial setups, not the full Animate-14B pipeline.

Running Wan2.2 Animate Locally

There are two local paths: the official CLI and ComfyUI. Both require downloading weights first:

git clone https://github.com/Wan-Video/Wan2.2
huggingface-cli download Wan-AI/Wan2.2-Animate-14B --local-dir ./Wan2.2-Animate-14B

CLI Path

You preprocess the driving video to extract skeletons and facial landmarks, then run inference. The mode you select determines the preprocessing flags (full argument lists are in the model card's run section):

  • Animation (Move): --retarget_flag --use_flux
  • Replacement (Mix): --replace_flag --use_relighting_lora --iterations 3 --k 7 --w_len 1 --h_len 1

Single-GPU runs use --refert_num 1, per the model card's example commands. The card also warns: "We do not recommend using LoRA models trained on Wan2.2." Weight changes from third-party LoRAs can cause unexpected behavior in the animation pipeline.

ComfyUI Path

The ComfyUI native workflow needs these model files in their respective directories:

FileDirectoryRole
Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensorsdiffusion_models/Kijai FP8 main weights
umt5_xxl_fp8_e4m3fn_scaled.safetensorstext_encoders/UMT5 XXL text encoder (FP8)
clip_vision_h.safetensorsclip_visions/CLIP Vision encoder
wan_2.1_vae.safetensorsvae/Wan2.1 VAE
lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16.safetensorsloras/LightX2V 4-step acceleration LoRA

You also need two custom nodes: ComfyUI-KJNodes and comfyui_controlnet_aux (which provides the DWPose Estimator for skeleton preprocessing). Two practical constraints apply: output width or height must be divisible by 16, and each Video Extend node adds 77 frames (~4.8 seconds) for clips longer than the base generation. To activate Mix mode you keep all connections; for Move mode you disconnect background_video and character_mask from the output subgraph node.

Hosted API Pricing Across Providers

If local setup is too heavy, several providers host Wan2.2 Animate with pay-per-use pricing. The rates look similar at first glance, but billing models differ in ways that materially affect your bill:

Provider480p720pBilling modelDuration limits
fal.ai$0.04/s$0.08/sFrames normalized to 16 fps; high-FPS source costs more—
WaveSpeed$0.20 / 5s$0.40 / 5sPer output run, tiered in 5-second blocks5–120s
APIXO$0.04/s$0.08/sPer detected reference-video second5s min, 120s cap
Replicate——$3 per 1,000 inference-seconds (compute time, not output)—

WaveSpeed's 5-second blocks normalize to the same effective rate ($0.04/s at 480p, $0.08/s at 720p), so fal.ai, WaveSpeed, and APIXO are price-competitive on the headline number. The catch is what gets measured: fal.ai normalizes frame counts to 16 fps, so a 30 fps driving video costs more per wall-clock second than the rate suggests. Replicate bills GPU inference time rather than output duration: a 10-second output that takes 40 seconds of compute costs $0.12. Verify which modes each endpoint exposes before committing; the fal.ai URL above points specifically to the Move/animation endpoint.

What Actually Works (and What Doesn't)

The model's showcase demos look polished, but community experience reveals a gap between demo quality and what you get from a first local run:

"what's the secret, is it post processing, frame interpolation" — u/Grand-Summer9946, r/comfyui

Based on that thread and the official preprocessing guide, here is what works and what breaks.

What works well: Single-person dance and motion-transfer clips, where one character moves against a relatively static background. Facial-expression reenactment is strong when the source video has clear, front-facing performance.

What breaks:

  • Multi-person scenes. Mask extraction is designed for single-person videos only; tracking the wrong pose is a common failure.
  • Body-proportion mismatch. Differences between the character image and the driving video produce artifacts and deformations in Mix mode.
  • FPS mismatch. When the driving video and the workflow configuration disagree on FPS, expect stuttering, slow-motion, or zoom artifacts.
  • Mask granularity. A coarser mask preserves more background but can cause shape leakage; a finer mask constrains generation and risks background inconsistency.

Raw output is usable for previews and concept tests. Production-quality results may require masking, inpainting, and frame interpolation layered on top, as the cited community thread discusses.

Wan2.2 Animate vs Other Wan Models

Wan is a fast-moving model family, and the version naming can be confusing. Per the model card, here is where Wan2.2 Animate sits:

ModelWhat it doesRelation to this article
Wan2.2-Animate-14BCharacter animation + actor replacement (video-to-video)This article's subject
Wan-Animate-2Community-reported successor modelDiscussed on r/LocalLLaMA (2026)
Wan 2.7 Image ProImage generation and editing (not video)Different modality: static images
Wan 3General-purpose text-to-video / image-to-videoNewer general video, no dedicated character-animation mode
Wan2.2-S2V-14BSpeech-to-video (audio-driven animation)Companion model for audio-driven motion

FAQ

Can Wan2.2 Animate handle multiple characters at once?

No. The official preprocessing guide builds mask extraction for single-person videos, so both modes support one character per run. Multi-person input can fail or track the wrong pose. For multiple characters, process each separately and composite in post-production.

What resolutions does it support?

Supported workflow resolutions are 480p and 720p; the official preprocessing samples use 1280×720. Nothing above 720p is documented, so higher resolution requires upscaling as a separate step.

What license applies to the output?

The model code and weights ship under Apache-2.0. Wan-AI claims no rights over generated content but holds users responsible for lawful, non-harmful use. The model card prohibits generating unlawful content, targeting vulnerable populations, or malicious personal-data use.

Does Wan2.2 Animate generate audio?

No. Animate produces video only. If you need audio-driven animation (lip-sync from speech), that is the separate Wan2.2-S2V-14B model, which added CosyVoice TTS support on September 5, 2025 per the Wan2.2 model card's update log.


Bottom line: Use Move for stills, Mix for actor replacement. Run locally if you have GPU and ComfyUI skills; use a hosted API (fal.ai, WaveSpeed, APIXO at $0.04–0.08/s) for speed, but budget for fal.ai's 16-fps normalization and Replicate's compute-time billing.