Category · the models

Realtime video models — the engines of live AI

The "realtime" label is the whole difference. Most AI video models generate offline, taking minutes for a clip you then upload. A realtime video model is post-trained for speed — today that means fal’s H3 Max build of MiniMax H3, which fal reports at roughly 35× the throughput of MiniMax’s official endpoint, producing about five seconds of 768p video in under three.

Listed · 6

The generation work that makes faster-than-playback video possible.

Benchmarks, with a caveat

On Artificial Analysis, H3 Max led image-to-video-with-audio at Elo ~1,204 as of early Sep 2026; on text-to-video it sits lower. The speed and throughput numbers are fal’s own and have not been independently reproduced — another outlet cites up to 50×. Resolution tops out at ~768p; MiniMax H3’s 2K regeneration pass is API-only.

The open-weight base

MiniMax H3 (Hailuo 3.0) is the open-weight 33B base underneath the wave — 15-second clips with native stereo audio, weights opened Aug 3 2026. Only the base is open (locally around 768p); the 2K and reference modules are API-only, and the licence excludes several regions including the US and EU.

Frequently asked questions

What is a realtime video model?

A video model that generates clips fast enough to sustain a live stream — faster than playback, rather than an offline render job.

Which model powers the 2026 live-AI wave?

fal’s H3 Max Live, a realtime-tuned build of MiniMax H3 — used by fal.live and InfiniteSlop.

Is MiniMax H3 open source?

The H3 base weights are open (Aug 3, 2026); 2K regeneration and reference modules are API-only, and the licence has regional exclusions.

How fast is H3 Max, really?

fal reports ~35× official H3 throughput and ~2.8s for a 5s/768p clip; those figures are vendor-reported and not yet independently reproduced.

Keep exploring