
Four serious open-weight AI video models exist in August 2026: LTX-2.5 (22B), MiniMax H3 (33B), Alibaba's Wan family, and Tencent's HunyuanVideo. All publish downloadable weights — but "open" doesn't mean free to run. Between GPUs, licenses, and API-only components, the cheapest way to actually use the best of them is a hosted deployment like Soloa's LTX-2.5 at $0.01/sec.
| Model | Size | Max Output | Audio | License | Hosted From |
|---|---|---|---|---|---|
| LTX-2.5 (LTX / Lightricks) | 22B | 4K · 20s | ✅ Native | Free under $10M ARR | $0.01/sec on Soloa |
| MiniMax H3 (Hailuo 3.0) | 33.1B | 768p open · 2K API | ✅ Stereo | Community License, <$20M revenue | $0.1625/sec (official API, 2K) |
| Wan 2.2 / 2.5 (Alibaba) | Various | 1080p · 10s (4K preview) | ✅ (2.5) | Apache 2.0 (2.2) | Varies by provider |
| HunyuanVideo 1.5 (Tencent) | 13B | 720p · 15s | ✅ | Custom community license | Varies by provider |
LTX-2.5 is a 22B open-weights model with weights on Hugging Face, free for organizations under $10M ARR. It generates up to 4K, 20-second clips with native synchronized audio, and it's the fastest model in the class — about 6.8 seconds for a 10-second 720p clip on top-end NVIDIA hardware, faster than real time. Its new Diffusion Video Decoder and native multishot generation (consistent characters across cuts) closed most of the quality gap with closed models. Full breakdown in our LTX-2.5 deep dive.
MiniMax launched H3 on July 31, 2026 and open-sourced it on August 3 — a 33.1B transformer that reads text, images, video, and audio as one unified context and outputs 4–15 second clips with native stereo sound. The open release ships two checkpoints (FL2VA for text/image-to-video, Ref2VA for up to 9 reference images, 3 video clips, and 3 audio clips), with quantizations down to NVFP4. ComfyUI support landed the same day.
The catches: the open checkpoints default to 768p — the 2K upscaler (H3-Regenerate-2K) and the prompt preprocessor (H3-Context-IR) remain API-only. The text encoder alone is 48GB at full precision (14.6GB quantized), so this is not a consumer-GPU model. The official API runs $0.1625/sec at 2K with ~4-minute generation latency, and the Community License caps free commercial use at $20M annual revenue.
Alibaba's Wan family is the licensing favorite: Wan 2.2 is Apache 2.0 — no revenue caps, no restrictions, genuinely free for commercial use. The newer Wan 2.5 generates 1080p clips up to 10 seconds with synchronized dialogue, ambient sound, and music in one pass (4K in preview). If your legal team needs a clean license above all else, Wan 2.2 is the answer; if you want the newest capabilities, check each Wan release's terms individually — they vary by version.
Tencent's HunyuanVideo 1.5 packs strong cinematic realism into just 13B parameters using a Causal 3D VAE, producing 15-second 720p clips with audio integration. It's the most practical to run on modest hardware of the four, though its custom community license carries more restrictions than Apache 2.0 — read it before shipping a product on it.
Open weights eliminate the model vendor's margin — not the compute. To actually generate video you need:
Say clips render at real-time speed on your rented GPU and you pay $1/hour: that's ~$0.017 per second of video if the GPU never sits idle — before your time. This is exactly why hosted open-weights deployments make sense: Soloa keeps LTX-2.5 on always-warm servers and charges $0.01/sec at 720p — at or below realistic DIY cost, with zero setup, and it's included in every subscription (20–50 generations/day). Open-model economics, closed-model convenience.
Premium closed models — Seedance 2.5, Veo 3.1 — still lead outright quality for hero shots, at 15–40x the per-second price. See the full market pricing table in our cost-per-second comparison.
LTX-2.5 for overall quality-per-dollar and speed; MiniMax H3 for multimodal reference control; Wan 2.2 for unrestricted Apache 2.0 licensing. "Best" depends on whether you optimize for output, control, or license terms.
The core 33B checkpoints (FL2VA, Ref2VA) are downloadable under the MiniMax H3 Community License — free for non-commercial use and for companies under $20M revenue. But the 2K upscaler and prompt preprocessor stay API-only, so the full flagship experience isn't fully open.
LTX-2.5 and HunyuanVideo run on 16GB+ VRAM cards (RTX 4080/4090 class). MiniMax H3 is realistically a multi-GPU or cloud job even quantized. Hosted access at $0.01/sec is cheaper than upgrading a GPU for most people.
Yes, within each license: Wan 2.2 (Apache 2.0) has no limits, LTX-2.5 is free under $10M ARR, H3 under $20M revenue. Above those thresholds you negotiate a commercial license — or use a hosted platform that handles licensing for you.
Because compute isn't. Once you count GPU rental, idle time, and setup, DIY costs roughly what Soloa charges for LTX-2.5 ($0.01/sec at 720p) — without the browser interface, warm servers, image-to-video modes, and daily generation allowances.
Bottom line: 2026 is the year open-weight video caught up. LTX-2.5 leads on speed and cost, H3 on multimodal control, Wan on licensing. And the cheapest way to use the best of them isn't a GPU in your closet — it's $0.01/sec on Soloa.
50+ AI models for image, video, voice, and music. One subscription, no switching between tools.