On Oct 5 Reka Labs released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch that unifies text, image, video and robot-action tokens in one network and one KV cache, with no tool calls or second model. Each transformer block has an understanding stream and a generation stream with shared attention, trained jointly with next-token prediction and flow matching. Reka reports base video generation at 0.79x real time (median) and a distilled 8-step variant (vs 99) returning a 5.3s clip in about a second. The checkpoint was trained on 320 H100s for three months. Research preview only; no public weights.

Key Takeaways

  • ✓Scale: 19B parameters trained from scratch on just 320 H100s for three months
  • ✓Speed: base video at 0.79x real time (median), ~6s to a watchable stream; distilled 99→8 denoising steps returns a 5.3s clip in ~1s
  • ✓Architecture: dual expert streams (understanding/generation) per block with shared attention and one KV cache; next-token + flow-matching objectives
  • ✓Limits: native video capped at 672x384; long-horizon drift, temporal grounding and edit stability are admitted weak points
  • ✓Access: research preview only, no open weights or public API; partnerships via [email protected]
Reka releases Rho-1 research preview: a 19B omni-reasoning model trained from scratch that understands and generates text, images and video and emits robot actions in one network
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Reka Labs' Rho-1 (blog) is a 19B model trained from scratch that collapses the usual multimodal agent pipeline — a planner LLM handing off to image, video and robotics specialists — into one network with one KV cache. Text and commands are discrete tokens; image/video latents, robot actions and proprioception are continuous tokens. Each transformer block splits into an understanding stream and a generation stream with shared attention; the understanding stream emits a handoff token and the generation stream renders from the full accumulated state via a diffusion head. Training jointly optimizes next-token prediction and flow matching.

Reka publishes no standard benchmark scores, only internal speed numbers: base video at 0.79x real time (median) with a watchable stream in ~6s, and a distilled variant that cuts denoising from 99 to 8 steps and returns a 5.3s clip in about a second. Robotics is shown on LIBERO simulation demos without success-rate tables. Admitted limits: 672x384 native video, long-horizon structural drift, weak temporal grounding and brittle edits. The checkpoint used 320 H100s for three months. Rho-1 is a research preview with no open weights or public API; collaboration is via [email protected]. For developers the takeaway is architectural: a single model with shared state as an alternative to multi-model orchestration for robotics, simulation and interactive video.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.