You're staring at an RTX 3060 with 12 GB VRAM, trying to decide whether to queue a clip in MiniMax H3 or LTX 2.3. Both are open-weights video models with native audio. Both run locally through ComfyUI. But one generates a 5-second 1080p clip in under 40 seconds while the other takes 11 minutes for 10 seconds of 480p - and the slow one produces physics that the Reddit community is calling "not close." The real question isn't which model is better. It's which model fits the shot, the hardware, and the delivery deadline you have right now.
MiniMax H3 vs LTX 2.3 at a Glance
| Spec | MiniMax H3 | LTX 2.3 |
|---|---|---|
| Architecture | 33B DiT + Qwen3-VL-32B encoder | DiT-based, new VAE |
| Max resolution | 2K | 4K |
| Max duration | 15 seconds | 20 seconds (10s at 4K/1440p) |
| Audio | Native stereo sound | Native audio, cleaner vocoder, audio-to-video endpoint |
| LoRA support | Not yet | Style LoRAs, HDR IC-LoRA, character LoRAs |
| Portrait mode | Via aspect ratio control | Native 9:16, trained on portrait data |
| Min VRAM (local) | 8 GB (INT8 quantized) | ~12 GB practical (OOM reports on lower cards) |
| License | MiniMax Community License | Apache 2.0 weights; LTX Model License for commercial use over $10M revenue |
| Released | API July 31, 2026; weights Aug 3, 2026 | March 2026 |
Both models landed as open-weights within months of each other. LTX 2.3 has a five-month head start with a mature LoRA pipeline; H3's weights are eight days old at time of writing.
Render Speed and Hardware: The Real Bottleneck
Speed is where the gap becomes unworkable for some pipelines - and where raw model size misleads. H3 is technically the bigger model (21 GB diffusion weights, 16 GB text encoder), yet multiple users report it running more stably than LTX 2.3 on consumer hardware:
"Even though it is technically bigger than LTX 2.3, it runs faster and has less problems with memory. I've recently tried 2.3 again and was annoyed by memory consumption." - Reddit user dobomex761604, r/StableDiffusion
But stability ≠ speed. Concrete wall-clock timings from the same Reddit comparison thread:
| Hardware | Model | Duration / Resolution | Wall-clock |
|---|---|---|---|
| RTX 3060 12 GB | H3 (INT8) | 10s / 480p | ~11 min |
| RTX 3060 12 GB | H3 (INT8) | 5s / ~480p | ~3-4 min |
| RTX 5090 24 GB | H3 (full) | 15s / 720p / 20 steps | 48 min |
| RTX 4080 | LTX 2.3 | 15s / 1080p | 15-20 min |
| RTX 4090 | LTX 2.3 | 5s / 1080p | 25-40s |
| RTX 4090 | Wan 2.2 (14B FP8) | 5s / 1080p | 90-120s |
A practical tip from user srikantpatnaik: entering the ComfyUI subgraph and dropping steps from 20 to 10 saves roughly 40% of H3's generation time with manageable quality loss. The seconds-per-resolution ratio is H3's biggest weakness right now.
On the VRAM side, H3 surprises. The INT8-pruned model from ComfyUI's own distribution runs on cards as low as 8 GB VRAM with 16 GB system RAM, though you're limited to small resolutions (0.3-0.4 megapixels, roughly 480-600p) and 5-second clips per user Aadi_880's testing notes. LTX 2.3 users report frequent OOM errors on 8 GB cards, with one user calling their experience "a disaster" for memory.
Prompt Adherence and Physics: H3's Decisive Win
If speed is LTX's domain, prompt comprehension is where H3 takes a commanding lead. Several participants in the linked Reddit comparison thread preferred H3 for prompt adherence and physics:
"MiniMax H3 is significantly easier to use. Pretty decent scenes can be generated without making convoluted, heavily descriptive prompt. That alone seals the deal for me." - Reddit user nvidiot, r/StableDiffusion
"From my tests. Minimax understands physics and prompt following very well. It doesn't shy away from graphic scenes like action or blood. I can get my result with 4 sentences compared to LTX which needed 2 paragraphs." - Reddit user Fit_Satisfaction2953
The physics gap is specifically about object permanence. LTX 2.3 frequently lets objects pass through each other or vanish mid-scene. H3 uses a 33B DiT with the 50th-layer hidden states from Qwen3-VL-32B as its text encoder, and community testers attribute its stronger spatial tracking to this larger architecture. User SX2k7 noted: "Things just disappear or move through each other for no reason way too often with LTX."
Users on the same Reddit thread also report H3 is less censored than LTX and Wan, allowing action scenes, blood, and physically intense moments that those models filter by default.
Identity Preservation in Image-to-Video
H3's Ref2Vid (reference-to-video) capability is a standout feature for identity preservation. Feed it a character image, and community testers report the model maintains facial identity across a full clip, even when the character walks off-screen and returns.
LTX 2.3's image-to-video mode has improved in this release (less freezing, fewer Ken Burns-style fake motion artifacts), but identity drift remains a known issue:
"Yep I tried an I2V where the character was out of shot for 10 seconds of the shot and then looked back at the end, identical face. LTX be like, 'who dat?'" - Reddit user kemb0
For product shots, character consistency, and brand work where the same face must survive a 10-second clip, H3's Ref2Vid is the stronger option among open-weights models.
Audio-Reactive LoRA Workflows: LTX's Customization Edge
This is where LTX 2.3 still holds a structural advantage that H3 cannot match yet. The LTX ecosystem has had months to accumulate custom tooling:
- Style LoRAs: Fine-tune the model on a specific visual aesthetic (stop-motion, handcrafted, miniature) and swap them in per shot
- HDR IC-LoRA: Generate directly in HDR with extended dynamic range, or convert SDR footage to EXR for finishing pipelines
- Audio-to-video endpoint: Provide an audio clip and LTX generates matching visuals - useful for music-driven content and social clips where sound-to-motion sync matters
- ComfyUI seedhunter and director workflows: Community-built multi-pass pipelines that iterate on seeds to find the best output, then refine composition
- Portrait 9:16 native: Trained on portrait-orientation data, not cropped from landscape - critical for vertical social content
- 24/48 FPS options: Frame-rate flexibility for matching deliverable specs
H3's workflow is comparatively bare: text-to-video, image-to-video, and reference-to-video with audio. No LoRA training pipeline, no style adapters, no HDR output. The model is eight days old at time of writing, so these gaps may close - but today, creators with established LTX LoRA libraries and tuned ComfyUI graphs pay a real switching cost to move.
User l2ddit captured the tension: "I am so glad I can now delete all those LTX loras that didn't do anything to fix the issues. The only thing LTX did better was non-English voice sync."
Cost Breakdown: Local vs API
Running either model locally is free in licensing terms, but time is the real cost. Through hosted APIs, the pricing split is wide:
| Endpoint | LTX 2.3 (fal.ai) | MiniMax H3 (hosted) |
|---|---|---|
| Text-to-video 1080p | $0.06/s | Not yet published via API |
| Text-to-video 4K/2K | $0.24/s (4K) | Not yet published via API |
| Fast variant 1080p | $0.04/s | N/A |
| Audio-to-video | $0.10/s | Included with video |
| Image-to-video | $0.06-0.24/s | Included with video |
| Extend / Retake | $0.10/s | Not yet available |
For a concrete budget: a 5-second 1080p clip on LTX 2.3 through fal.ai costs $0.30. The Fast variant drops that to $0.20. MiniMax H3's hosted API pricing had not been publicly published at time of writing. Locally, both are free to run, but H3's 11-minute render on a 3060 adds real iteration cost.
LTX 2.3 weights are available on HuggingFace under Apache 2.0; the LTX Model License governs commercial embedding for companies over $10M annual revenue. H3 weights are on HuggingFace under the MiniMax Community License. Both permit local, offline use. LTX 2.3 API pricing is sourced from fal.ai's model page.
Production Routing: When to Use Each
Most creators don't need to pick one model. A hybrid pipeline plays to each model's strengths:
Route to H3 when:
- The shot requires complex physics - collisions, fast motion, physical interactions
- Character identity must hold across a long clip (product reveals, narrative scenes)
- You want native stereo audio without a separate sound pass
- Prompt simplicity matters (4 sentences vs 2 paragraphs)
- Your GPU is on the lower end (8-12 GB VRAM) and you can accept longer render times
Route to LTX 2.3 when:
- You need rapid iteration - storyboards, timing tests, draft previews
- A custom LoRA defines your visual style
- Portrait 9:16 social content is the deliverable
- 4K resolution is required (H3 caps at 2K)
- You need audio-to-video sync where an existing audio track drives the visuals
- Speed matters more than peak per-frame quality
A practical flow: use LTX Fast to validate blocking and timing, then queue the approved shot in H3 for the final render with physics, identity, and audio. Post-process with an upscaler only after the shot is approved. The render-time cost of re-running H3 makes it worth getting the shot right first.
FAQ
Is MiniMax H3 better than LTX 2.3?
H3 wins on quality, prompt adherence, and identity preservation. LTX 2.3 wins on speed, LoRA customization, and 4K output. Route by shot type, not by brand.
Can MiniMax H3 run on consumer GPUs?
Yes. The INT8-pruned model runs on 8 GB VRAM + 16 GB RAM at roughly 480-600p and 5 seconds. Full-precision benefits from 24 GB VRAM.
Does H3 support LoRA fine-tuning?
Not yet. As of the August 3, 2026 weights release, H3 has no LoRA pipeline. LTX 2.3 supports Style LoRAs, HDR IC-LoRAs, and character LoRAs.
Which model has better audio?
Both generate native audio. H3 produces stereo sound with richer ambient detail. LTX 2.3 improved its vocoder for cleaner output and offers an audio-to-video endpoint that generates visuals from an input audio clip.
What VRAM do I need for each model?
H3 INT8: 8 GB VRAM + 16 GB RAM minimum for 480p. Full-precision H3 benefits from 24 GB VRAM. LTX 2.3: practical minimum ~12 GB VRAM, with OOM reports on 8 GB cards.