
LTX-2.5 reads prompts like a film director reads a shot description — full sentences describing camera, subject, action, environment, and style, not comma-separated tags. Master that structure and your usable-clip rate jumps from 1-in-10 to 1-in-3. This guide gives you the exact formula, a camera-movement vocabulary, audio cues, multishot syntax, and templates you can paste straight into Soloa's LTX-2.5 generator.
LTX-2.5 uses a fine-tuned language-model text encoder that parses scene direction, spatial relationships, and causality. The reliable prompt structure layers five components in order:
| Layer | What to Specify | Example Fragment |
|---|---|---|
| 1. Camera & lens | Movement, angle, focal length | "A slow dolly-in on a 35mm anamorphic lens…" |
| 2. Subject | Who/what + defining attributes | "…a silver-haired ceramicist in a clay-streaked apron…" |
| 3. Action & timeline | What happens, in what order | "…shapes a spinning bowl, then pauses to inspect the rim…" |
| 4. Environment & lighting | Setting, light source and direction | "…in a sunlit studio, golden afternoon light raking through dusty air…" |
| 5. Style & fidelity | Look, grade, render quality | "…documentary realism, shallow depth of field, fine skin texture." |
One more layer is unique to LTX-2.5: audio. The model generates sound in the same pass as the picture, so ambient noise, music, and dialogue belong in the prompt: "the wheel hums softly, rain taps the skylight, no music."
Name a camera move explicitly — it's the single highest-leverage phrase in the prompt. Moves LTX-2.5 executes reliably:
| Camera Move | Effect | Use For |
|---|---|---|
| Slow dolly in | Smooth approach toward subject | Emotional intensity, reveals |
| Dolly zoom (vertigo effect) | Zoom in while pulling back | Psychological tension |
| Low-angle tracking shot | Ground-level parallel movement | Action, power shots |
| 360-degree orbit | Revolves around the focal point | Product and character reveals |
| FPV drone flythrough | Fast multi-axis aerial movement | Landscapes, establishing shots |
| Crane down / jib down | Vertical descent | Wide-to-intimate transitions |
| Macro rack focus | Focus shifts foreground ↔ background | Detail storytelling, time-lapses |
| Locked-off tripod shot | No movement at all | Dialogue, UGC-style content |
Pair the move with a lens and lighting direction: 85mm portrait lens, softbox key from camera left. Lens language (35mm anamorphic, 100mm macro), film-stock references, and volumetric terms (fog, dust, steam, rain reflections) all steer the render noticeably.
A slow 360-degree orbit around a matte-black wireless earbud case standing on wet slate, macro lens, water droplets beading on the lid. The case opens as the camera passes the front, LEDs pulsing soft blue. Studio darkness with a single overhead softbox, reflections gliding across the surface. Crisp product-photography fidelity. Audio: low ambient synth hum, a soft click as the lid opens.
A locked-off smartphone-style shot of a woman in her late 20s with curly hair and a mustard cardigan, sitting in a bright kitchen, speaking directly to camera with natural hand gestures. Morning light from a window on the right, slight handheld sway. Casual vlog realism, natural skin texture. Audio: her voice saying "I tried this for a week — here's what happened," faint kitchen ambience.
An FPV drone flythrough skimming low over turquoise shallows toward a limestone cliff at golden hour, then rising fast to reveal a fishing village strung with lights. Sun flares through haze on a 24mm wide lens, warm cinematic grade with teal shadows. Audio: wind rush, distant gulls, a swelling ambient score.
A macro rack focus starting on steam curling from a ramen bowl, shifting focus to chopsticks lifting noodles in slow motion, broth droplets catching warm tungsten light. Dark izakaya background with soft bokeh lanterns. Rich editorial food styling, high detail. Audio: gentle sizzle, quiet restaurant murmur, chopstick clink.
A 12-second three-shot sequence. Shot 1: wide establishing shot of a neon-lit repair shop on a rainy street at night, puddle reflections rippling. Cut to Shot 2: close-up of a mechanic's oil-stained hands soldering a drone rotor, sparks lighting her safety glasses amber. Cut to Shot 3: over-the-shoulder as the drone lifts off her palm and hovers between them. Preserve character appearance, rain intensity, neon color palette, and lighting direction across all cuts. Audio: rain patter throughout, solder hiss in shot 2, rising rotor whine in shot 3.
Shot 1… Cut to Shot 2… transitions — LTX-2.5 natively generates multiple cuts in one pass.Weak: cyberpunk woman, walking, neon lights, rainy street, 8k, cinematic
Comma-stacked tags produce jerky motion and generic composition because the encoder gets no spatial or causal relationships to work with. Rewrite as direction:
Strong: "A low-angle tracking shot follows a woman in a translucent rain jacket walking through a neon-soaked market alley at night, magenta and cyan signs reflecting in the puddles around her boots, steam drifting between stalls. Cinematic grade, shallow depth of field. Audio: rain, distant synth music, wet footsteps."
| Parameter | Recommended | Notes |
|---|---|---|
| Guidance (CFG) | 3.0–4.2 | Above ~5.0 causes oversaturation and color burn |
| Steps | 25–35 (full) · 8–12 (distilled) | Speed vs detail trade-off |
| Denoise (i2v) | 0.45–0.65 | 0.45 preserves the source image; 0.65 allows transformation |
| Duration | "auto" or explicit | Auto predicts clip length from the described action |
A widely used negative prompt for LTX-2.5: blurry, distorted anatomy, warped faces, flickering background, unnatural jitter, overexposed, static freeze frame, watermarks, text overlays. On Soloa these parameters are pre-tuned — you just write the prompt and pick a mode (text-to-video, image-to-video, or first/last-frame).
Even with perfect structure, great clips come from variations: nudge the camera move, swap the lighting, tighten the action beat. On premium APIs at $0.10–$0.53/sec, ten 10-second takes cost $10–$53 — so people stop iterating early and settle. On Soloa's LTX-2.5 at $0.01/sec (720p), those same ten takes cost $1. Cheap iteration is the difference between testing three ideas and testing thirty — see how the pricing compares market-wide in our cost-per-second breakdown, or get the full model rundown in LTX-2.5 explained.
Two to five full sentences. Long enough to cover camera, subject, action, environment, style, and audio — short enough that each shot has one clear beat. Single-word or tag-list prompts underperform badly.
Yes. Put the spoken line in quotes inside the prompt ("she says: 'let's begin'") and describe the voice tone. Audio is generated in the same pass, lip-synced to the speaker.
Use image-to-video or first/last-frame mode with a reference still, repeat the same character description verbatim across prompts, and add a consistency clause in multishot prompts. Both modes are available on Soloa.
Usually tag-stacked prompts, contradictory motion instructions, or guidance set too high. Rewrite as full-sentence direction, keep one camera move per shot, and stay in the CFG 3.0–4.2 range.
Iterate at low resolution first — on Soloa, 480p drafts cost $0.005/sec, so a hundred 10-second practice clips cost about $5. Lock the prompt, then re-render the winner at 1080p for $0.02/sec.
Bottom line: direct, don't describe — camera first, one action beat per shot, audio in the prompt. Then iterate ruthlessly, because at $0.01/sec on Soloa the retries that make clips great are basically free.
50+ AI models for image, video, voice, and music. One subscription, no switching between tools.