Published by Video Runtime · Source check 2026-09-21 · Evidence method · Suggest a correction
EDITORIALLY VERIFIED PRODUCT INFORMATION
About Stream LTX
A project page for a block-causal LTX-2.3 rebuild that generates joint audio and video in real time with mid-stream text instructions. One-second audio-video blocks use causal masks across six attention paths, teacher forcing with resample forcing, diffusion forcing and few-step distillation, plus a fixed-budget three-tier KV cache and separated global and local prompts. New text instructions can be handed to the model between one-second blocks, changing body action, dialogue, props or scene conditions mid-take.
Key features
The project page publishes five uncut console recordings with timecoded cue sheets linking each instruction to the moment it reached the model. The method targets genuinely joint audio-video streaming rather than silent video generation.
Getting started
Review the project page recordings and cue sheets; there is nothing to install until code or weights are released.
Real-time interaction
One-second audio-video blocks use causal masks across six attention paths, teacher forcing with resample forcing, diffusion forcing and few-step distillation, plus a fixed-budget three-tier KV cache and separated global and local prompts. New text instructions can be handed to the model between one-second blocks, changing body action, dialogue, props or scene conditions mid-take.
Pricing
Project page only; no licence, code or weights are released Repository code is published under No licence declared yet. Confirm separate model, dataset and dependency terms before reuse. No hosted price is inferred. Budget for the documented GPU, storage and setup requirements.
Limitations
No code or weights are released; the page documents an internal research prototype built during a 2026 Tencent research internship. The headline throughput counts steady-state DiT denoising only and excludes text encoding, VAE and audio decoding, with first-block latency reported separately.
EDITORIAL VIEW
Editor's Verdict
Stream LTX is included as a source-available implementation relevant to real-time generative video or interactive world systems.
Run the supplied example on representative hardware and review every linked licence before production use.
How It Works
- One-second audio-video blocks use causal masks across six attention paths, teacher forcing with resample forcing, diffusion forcing and few-step distillation, plus a fixed-budget three-tier KV cache and separated global and local prompts.
- New text instructions can be handed to the model between one-second blocks, changing body action, dialogue, props or scene conditions mid-take.
What We Like
- The project page publishes five uncut console recordings with timecoded cue sheets linking each instruction to the moment it reached the model.
- The method targets genuinely joint audio-video streaming rather than silent video generation.
Current Limitations
- No code or weights are released; the page documents an internal research prototype built during a 2026 Tencent research internship.
- The headline throughput counts steady-state DiT denoising only and excludes text encoding, VAE and audio decoding, with first-block latency reported separately.
Pricing & Access
- Repository code is published under No licence declared yet. Confirm separate model, dataset and dependency terms before reuse.
- No hosted price is inferred. Budget for the documented GPU, storage and setup requirements.
Verification Summary
- Official repository and licence reviewed on 2026-09-21. Video Runtime has not installed or benchmarked this project.
- Performance and compatibility statements below are attributed to the maintainers and are not site measurements.
Technical notes & official performance
Official Performance Data
- Author claim: 43 FPS on 4 × H800 at 512 × 768 with ten-minute takes. Video Runtime has not reproduced it and the page excludes setup and decode costs from the figure.
Technical Notes
- The author claims 43 FPS at 512 × 768 on 4 × H800 with ten-minute continuous takes; the figure is a steady-state denoising throughput claim, not an end-to-end result.
Sources
- Official GitHub repository ↗github · accessed 2026-09-21
- Code licence ↗github · accessed 2026-09-21
- Official project page ↗official-docs · accessed 2026-09-21
Access, stage & mechanism
These are separate source-checked facts, not a live uptime monitor. Unverified fields are left open; transport and persistent sessions do not establish how frames are generated.
- Release stage
- Not verified
- Access mode
- Not verified
- Source availability
- Not verified
- Transport
- Not verified
- Interaction
- Not verified
- Session mode
- Not verified
- Generation mechanism
- Not verified
- Continuity method
- Not verified
- Licence status
- Not verified
- Tasks
- Not classified
Read how official-source and hands-on records stay separate in our Methodology.