
LTX-2.5 is a 22-billion-parameter open-weights AI video model from LTX (spun out of Lightricks). It generates up to 20-second clips at up to 4K with native synchronized audio, and it's faster than real time — a 10-second clip renders in about 6.8 seconds on top-end hardware. The official API costs $0.09–$0.30/sec, but Soloa runs it from $0.01/sec — the cheapest video generation on the market.
| LTX-2.5 at a Glance | Details |
|---|---|
| Developer | LTX (spun out of Lightricks) |
| Size | 22B parameters, open weights on Hugging Face |
| Output | Up to 4K, clips up to 20 seconds, 24/25/48/50fps |
| Audio | Native synchronized audio in the same pass |
| Speed | 10s 720p clip in ~6.8s on dual NVIDIA GB200s |
| Modes | Text-to-video, image-to-video, first/last-frame, audio-to-video |
| License | Free for orgs under $10M ARR; commercial license above |
| Cheapest access | Soloa: $0.005–$0.02/sec, no setup |
LTX-2.5 is the newest open-weights foundation model in the LTX video family. Where most frontier video models (Veo, Kling, Seedance) are closed APIs, LTX publishes its weights on Hugging Face — anyone can download, inspect, fine-tune, or self-host the model. It's free to use for organizations under $10 million in annual recurring revenue; larger companies negotiate a commercial license.
Two things set it apart from both closed rivals and other open models:
On the official API (fal.ai and other providers), LTX-2.5 comes in two variants:
| Variant | 720p | 1080p | 1440p | 4K |
|---|---|---|---|---|
| LTX-2.5 Fast | $0.09/sec | $0.13/sec | $0.19/sec | $0.30/sec |
| LTX-2.5 Pro | $0.12/sec | $0.17/sec | — | — |
Pro targets film-grade pipelines (EXR output, ACES color, 6/8/10-second fixed durations at 24/25/50fps). Fast covers everything else, up to 4K and 20 seconds, with an "auto" duration mode that predicts the right clip length from the described action. Both support 16:9 and 9:16 and include audio at every tier.
Download the weights and run them locally with a 16GB+ VRAM GPU. There's no per-clip fee and no watermark, but you're paying in hardware (or ~$0.40–$2/hour in cloud GPU rental), setup time, cold starts, and maintenance. Worth it for engineering teams; overkill for creators. Full breakdown in our open-weight model comparison.
Simple pay-per-second access, priced like other frontier APIs. A 20-second 1080p clip on Fast costs $2.60; ten iterations of it, $26.
Soloa operates its own always-warm LTX-2.5 deployment and passes the self-hosting economics on to subscribers:
| Resolution | Soloa | Official API (Fast) | Savings |
|---|---|---|---|
| 480p | $0.005/sec | — | — |
| 720p | $0.01/sec | $0.09/sec | 9x cheaper |
| 1080p | $0.02/sec | $0.13/sec | 6.5x cheaper |
It's included with every Soloa subscription — plans include 20–50 generations per day — and supports text-to-video, image-to-video, and first/last-frame-to-video with native audio, straight from the browser. No GPU, no API keys, no cold starts. That 20-second 1080p clip costs $0.40; ten iterations, $4.
On the quality leaderboards, ByteDance's Seedance 2.5 and Google's Veo 3.1 still lead on cinematic fidelity. But the cost gap is enormous: Seedance 2.5 text-to-video runs $0.293/sec at 720p — about 29x Soloa's LTX-2.5 rate — and Veo 3.1 Standard with audio is $0.40/sec, 40x more. For social content, drafts, b-roll, product clips, and iteration-heavy work, LTX-2.5's quality-per-dollar is unmatched in 2026. See the full market table in our cost-per-second comparison.
Generation speed sounds like a convenience feature until you count what it does to output. A model that takes 3–4 minutes per clip (typical for heavyweight closed models — MiniMax H3 averages ~264 seconds) lets you test maybe 10 ideas in an hour. LTX-2.5 on warm infrastructure returns clips in seconds to tens of seconds, so a one-hour session yields 40–60 iterations. Combined with $0.01/sec pricing, that turns AI video from a "one careful shot" tool into an actual creative sandbox — the same shift that made digital photography better than film for most working photographers.
LTX-2.5 uses a fine-tuned text encoder that reads prompts like a director reads a shot description — full sentences beat tag lists, and camera, lighting, and audio all belong in the prompt. We wrote a dedicated LTX-2.5 prompt guide with 30+ copy-paste prompts and a camera-movement vocabulary. And because Soloa charges $0.01/sec, you can afford the 5–10 iterations that separate a decent clip from a great one — a luxury that costs real money on $0.09–$0.53/sec APIs.
The weights are free to download and self-host for organizations under $10M ARR. Running them still costs GPU money. The cheapest hosted access is Soloa at $0.005–$0.02/sec, included with any subscription.
Faster than real time on data-center hardware — about 6.8 seconds for a 10-second 720p clip on dual NVIDIA GB200s. On hosted platforms, generations typically complete in well under a minute.
Yes. Ambient noise, music, and dialogue are generated natively in the same pass at every resolution tier — one of the few open models with built-in audio.
Yes — image-to-video with optional end-frame pinning, plus first/last-frame-to-video, which interpolates a clip between two stills. Both modes are available on Soloa's LTX-2.5 tab.
A GPU with at least 16GB VRAM for the base model. For most creators, hosted access at $0.01/sec costs less than the electricity and time of running it yourself.
Bottom line: LTX-2.5 made frontier-quality video generation open and fast. Soloa made it cheap. Generate your first clip from $0.01/sec.
50+ AI models for image, video, voice, and music. One subscription, no switching between tools.