Table of Contents
🌏 中文版
In February 2024, OpenAI previewed Sora with a minute-long Tokyo street scene and realistic physical reflections. The industry suddenly took text-to-video seriously. Sora was not even a public product—only a few filmmakers could test it—but it set the agenda for the next two years. Two and a half years later, this fifteenth family deep dive in the AI Model Landscape Overview has different protagonists: Sora has announced its exit, and Veo, Kling, and Runway divide the market.
For benchmark interpretation, see the AI Model Evaluation Sources Guide. This is the fifteenth family deep dive in the AI Model Landscape Overview.
Four-family Evolution Timeline
Sora: From Shockwave to Retirement
Sora 1 opened to ChatGPT Plus and Pro users at sora.com in December 2024 and generated silent video only. The real turning point was Sora 2 on September 30, 2025: native synchronized audio, major gains in physical correctness, an invitation-only iOS app, and a cameo system for licensing real people's likenesses. It became an immediate hit. The reversal was abrupt. OpenAI notified developers in March 2026 that every Videos API would retire; the consumer app closed April 26, and the API stops accepting requests September 24, 2026. Official documentation lists no successor. Third-party estimates put lifetime compute cost far above revenue. OpenAI's “GPT-1 moment” description proved prophetic: it was a beginning, not a product.
Veo: Native Audio Defines the Battlefield
Veo followed the steadiest path: announced at I/O in May 2024; Veo 2 reached VideoFX and Vertex AI in December; Veo 3 on May 20, 2025 became the first native-audio model deployed at scale, generating dialogue, effects, and ambience together, alongside the Flow creation tool. Veo 3.1 on October 15 improved audio, added Ingredients to Video with three character-reference images, first/last-frame control, and Extend for chained clips. Gemini API release notes show 4K output and full-resolution vertical video arriving in January 2026, while Veo 2/3.0 endpoints were scheduled to close by late June 2026. Google also began a research collaboration with A24 in June 2026.
Kling: Kuaishou's Commercial Surprise
Kling 1.0 debuted in June 2024 and opened global testing in July, distinguished by physical realism. It then iterated faster than anyone: 1.5 (2024-09), 1.6 (2024-12), 2.0 Master (2025-04), 2.1 (2025-05, with major price cuts), 2.5 Turbo (2025-09, leading an Artificial Analysis snapshot), 2.6 (2025-12, native audio), and 3.0 (2026-02). Version 3.0 brought a unified multimodal architecture, native audio in five languages, multi-shot storyboards, character coreference, and video up to 15 seconds. Its annualized revenue run rate reached $240M in December 2025, after 19 months and 60 million creators. Kuaishou integrated it with short video and advertising, making it the only player to demonstrate “model as ecommerce infrastructure.”
Runway: An Enterprise Workflow Moat
Runway commercialized first: Gen-1 in February 2023 for video-to-video, Gen-2 in March for text-to-video, and Gen-3 Alpha in June 2024. Gen-4.5, released December 1, 2025, led an Artificial Analysis text-to-video snapshot at 1,247 Elo. A week later, TechCrunch reported native audio alongside the GWM-1 world model. In 2026, Runway moved toward a platform: Aleph 2.0 editing, Runway Agent, Studio timeline, and July's Runway Dev plus Media Router. Its API changelog retired Gen-3 endpoints at the end of July. Customers such as Lionsgate show the positioning: not a single model, but the layer embedded in professional production.
Second tier, briefly: Luma progressed from Dream Machine to Ray3.2, a 16-bit HDR pioneer capped at 1080p in documentation; Pika emphasizes rapid short-video iteration; MiniMax open-sourced Hailuo 3.0's 33B base weights with native 2K in August 2026; xAI Imagine video starts at $0.05/s, with flagship 1.5 reaching roughly $0.25/s at 1080p inside the X ecosystem.
Specifications and Pricing
| Model | Maximum duration | Maximum resolution | Audio | Price/second | API |
|---|---|---|---|---|---|
| Sora 2 / 2 Pro | 4–25s by tier | 1080p Pro | Native | $0.10; Pro $0.30–$0.70, batch half-price | Yes; retires 2026-09-24 |
| Veo 3.1 Quality / Fast | 8s base, chainable with Extend | 1080p / 4K since 2026-01 | Native | ~$0.40 / $0.15; no-audio about one-third less | Gemini API + Vertex AI |
| Kling 3.0 | 15s multi-shot | High-resolution video tiers; 2K/4K images | Native, five languages | Credits; 2.x about half flagship price or less | Official + fal/Replicate |
| Runway Gen-4.5 | Short clips; exact cap varies by source | 4K from Gen-4 Turbo | Native, added 2025-12 | Credits; entry subscription ~$12–15/month | Official API + Runway Dev |
| Hailuo 3.0 (H3) | 5–15s, extendable to ~30s | 2K | Native stereo | ~$0.13–$0.26 by resolution/channel | Official API; open weights |
Prices are approximate at publication; verify current official pricing before procurement.
Architectural Consensus: Everyone Is on the Same Path
All four converge on diffusion transformers over latents. A 3D VAE compresses video into spatiotemporal latents, splits them into patches, and a Transformer denoises them. Sora's spacetime patches were the earliest public description; Kling documents DiT + 3D VAE. Runway's Gen-4.5 A2D advances this with autoregressive planning by Qwen2.5-VL and parallel diffusion decoding.
The second consensus is audiovisual synchronization. After Veo 3, Kling 2.6, Sora 2, Gen-4.5, and H3 all adopted native audio within a year. Sound went from differentiator to admission ticket. The third is consistency control: reference images lock characters through Veo Ingredients, Kling element consistency, and Runway multi-shot control, replacing early videos in which every shot seemed to recast the actors.
Copyright, Watermarks, and Commercial Use
Three approaches coexist. OpenAI uses visible plus verifiable provenance: Sora downloads carry a dynamic watermark and C2PA credentials. Google embeds invisible SynthID into pixels and waveforms and says it has marked tens of billions of assets. Some platforms reserve watermark-free export for paid tiers. Paid plans generally permit commercial use. Two traps remain: metadata can be stripped, so failure to detect a signal does not prove human origin; and restrictions on real likenesses, cameos, and copyrighted characters vary by platform. Read service terms before commercial use rather than relying on demos.
Selection Advice
- Marketing shorts: Veo 3.1 is the default. Native audio removes a whole post-production stage; use Fast for drafts and Quality for final output. Kling 2.6/3.0 offers a better audio/price mix when cost or Chinese speech matters.
- Film previsualization: Runway. Gen-4.5 consistency and Aleph 2.0 editing support director revisions, while Studio and Adobe Firefly integration fit existing pipelines. Use Luma Ray3.2 when HDR/EXR output is a grading requirement.
- Programmatic batch generation: avoid credit subscriptions and use per-second APIs. Hailuo H3 starts around $0.18/s depending on channel and can be self-hosted for low-cost volume. Veo on Vertex AI suits enterprise contracts and data governance. Do not start a new Sora workflow while its API retirement counts down.
Overall
The story has three acts. Sora defined the category with a preview but did not win the market. Veo reset the minimum standard with native audio and forced everyone to follow within a year. Kling proved a Chinese short-video ecosystem could turn a video model into a money machine, while Runway showed that professional workflows retain customers better than a leaderboard lead. The next dividing line is already visible: world models such as Runway GWM-1 and Google's Genie line, plus open weights such as H3 and Wan, are moving the contest from who makes the prettiest clip to who can simulate a world—and who can be deployed independently.
References
- Sora 2 — OpenAI
- Video generation with Sora — OpenAI API Docs
- What to know about the Sora discontinuation — OpenAI Help
- Creating with Sora safely — OpenAI
- Veo — Google DeepMind
- Introducing Veo 3.1 — Google Developers Blog
- Gemini API Release Notes
- Veo updates coming to Flow — Google Blog
- Kling AI — Wikipedia
- Kling AI Annualized Revenue Run Rate Hits USD240 Million — PR Newswire
- Introducing Runway Gen-4.5 — Runway Research
- Runway releases its first world model — TechCrunch
- Runway API Changelog
- SynthID — Google DeepMind
- Luma model information
- AI Model Evaluation Sources Guide
- AI Model Landscape Overview
Loading...