MiniMax H3 Max Generates 5-Second Video in 3 Seconds, Enabling Real-Time Generation
Co-developed with fal, H3 Max generates a 5-second 768p video in under 3 seconds, making generation faster than playback.
MiniMax has rewritten the assumptions around video generation speed. With a model that outputs 5 seconds of video in under 3 seconds, it has achieved a state where generation outpaces playback. With waiting time eliminated, formats of streaming and short-form viewing that previously relied on pre-production are being replaced by on-demand, instantaneous generation.
In reporting by QbitAI’s Kressy, a live stream in which AI generates all material on the spot is introduced as a case symbolizing this shift. Despite production only beginning after receiving viewer requests, the stream is characterized by an uninterrupted supply of video.
Structure Where Generation Outpaces Playback Time
The core of H3 Max lies in its speed. It takes less than 3 seconds to generate a 5-second video at 768p. Throughput reaches approximately 35 times that of the base H3.
It takes H3 Max less than 3 seconds to generate a 5-second, 768p video. Throughput is about 35 times that of the original H3.
The significance of this figure goes beyond mere acceleration. While one video is playing, the next one is completed. By reversing the time relationship between generation and playback, continuous supply without pre-generation or waiting becomes possible. In conventional video generation, waiting for generation fragmented the experience, but that fragmentation has been structurally eliminated.
Behind this development is a shift in thinking: aligning not with the constraints of generation time on the creator side, but with the viewing experience. If generation outpaces playback speed, streams and feeds can theoretically continue indefinitely. For applications requiring real-time performance, crossing this threshold marks a practical turning point.
Acceleration Through Joint Development with fal
H3 Max was created through joint development by MiniMax and fal. MiniMax open-sourced its video generation model H3 at the end of July 2026, recording 24 million downloads in three weeks. fal is a company specializing in AI inference infrastructure with deep expertise in accelerating inference.
The collaboration between the two companies converged on a clear objective: to build on H3 and apply post-training and inference optimization specialized for real-time generation use cases. In short, it was an effort to integrate acceleration mechanisms into an existing high-performance model.
This optimization was achieved not by overhauling the model architecture, but through refinement on both the training and execution sides. According to published information, the acceleration was intended to be achieved without sacrificing quality, with adjustments made for continuous generation in production environments. With a widely adopted open-source foundation, the effects of the optimization were quickly validated and fed back to the developer community, creating a virtuous cycle.
High-Speed Generation Achieved While
Maintaining Quality
Speed and quality often conflict. However, H3 Max has demonstrated that both can be achieved. In two major evaluations — Artificial Analysis and Design Arena — it ranked first in image-to-video generation in both.
This result shows that the increase in speed did not come at the cost of quality. Even if generation speed exceeds playback speed, the viewing experience fails if the visual quality is not maintained. Topping the evaluation rankings is evidence that the threshold was crossed while preserving production-ready quality.
In video generation, multiple factors determine quality, including resolution, duration, naturalness of motion, and audio synchronization. H3 Max is distinguished by delivering 5 seconds of generation at a practical 768p resolution in under 3 seconds without compromising these elements. Not only the speed figures but also its standing on quality metrics will determine its viability for commercial use.
Three Streaming Use Cases Emerge in Succession
With the speed requirement met, applications emerged organically. Notably, all were built around the same time by different developers without prior coordination.
The first is Rehan Sheikh, an engineer at fal. He connected H3 Max to an overseas streaming platform and launched a channel called “Cross-Dimension TV.” With no schedule or pre-recorded material, an animated short may play one moment, followed the next by a 1960s-style marionette advertisement, then a space documentary. All video and audio are generated and streamed continuously by the model without interruption. A clip he posted on X garnered 5.4 million views.
The second is indie developer Pieter Levels, known as the creator of Nomad List and PhotoAI and adept at solo product launches. Levels built a live streaming site where viewers can freely send requests and AI creates content on the spot. There is no host or script; viewer requests become the content itself. Because the latency from request to generation is shorter than the playback time, the stream never stalls despite being made to order.
The third is Alex Koumpas, also an engineer at fal. His stream broadcasts a single continuous story. Viewers can intervene via text in the chat, for example with instructions like “make the protagonist turn around and walk into the forest” or “make red rain suddenly fall from the sky,” which are reflected in the visuals and cause the story to branch according to viewer input. Because generation outpaces the narrative’s progression, interactivity and continuity coexist.
In addition, one developer used Codex to build a short-form video application that can be described as an “AI TikTok.” The content stream is generated in real time by the video generation model and linked seamlessly. While the viewer watches the current video on screen, the next one is generated and switches naturally with a swipe. It also features the ability to extend videos up to 30 seconds when needed, enhancing immersion.
Speed as a Turning Point in the Video
Generation Race
The “AI real-time streaming” format demonstrated by H3 Max could become a watershed in the competition over video generation. This is because it is not merely faster, but directly opens a path to commercialization by eliminating intermediate waiting time.
This structural shift is being compared to the change Peter Steinberger brought to AI agents through OpenClaw. OpenClaw transformed AI agents from command-line tools for developers into open-source assistants accessible to anyone. While the idea itself was simple, after reaching 4.7 million downloads it became a movement that even leading industry executives could not ignore.
The nature of competition differs between video generation and code generation. Code has correct answers, making generative models prone to homogenization. Video, on the other hand, has no single correct answer. Aesthetic sensibilities vary from person to person, and even with the same prompt, ten models will produce ten different styles. Users choose based on “whose visuals they prefer.” This characteristic creates a structure where competition does not converge on a few winners and is not decided on price alone.
Furthermore, accumulated usage reinforces lock-in. For creators who have established a visual style with a particular model, what viewers support is that worldview, and what companies seek is that texture. Switching models would require rebuilding accumulated visual assets and viewer expectations from scratch. As usage grows, the model becomes deeply embedded in asset libraries, character settings, and production workflows, raising switching costs.
Ecosystem Expansion Driven by Open Source
Open source and an open ecosystem further strengthen this lock-in. Developers and creators in the community produce derivative works around the same base model, and that feedback is fed back into improving the model, forming a cycle. Once this cycle begins, it becomes difficult for latecomers to catch up on performance alone.
H3 is already running a prototype of this cycle. Within three weeks of being open-sourced, it reached 24 million downloads worldwide, with more than 300 derivative models. Variants such as FastH3, H3 480p, and H3 Max have been nurtured by different teams, each branching out in different directions. A structure is forming with H3 as the trunk and blossoms emerging everywhere.
MiniMax’s open-sourcing of its state-of-the-art video model has an aspect of democratizing technology. It is an act of delivering cutting-edge capabilities directly into the hands of all developers and raising the foundation for the entire industry. If developers and creators around the world build in different directions on this foundation and accumulate assets, workflows, and usage habits around the same ecosystem, the flywheel of commercialization will continue to spin autonomously.
Stock Surge Reflects Strong Market Endorsement
Capital markets reacted swiftly to this development. Last week, Morgan Stanley maintained its overweight rating on MiniMax with a target price of HK$900. On the day of the report, September 1, 2026, MiniMax’s Hong Kong shares closed up 16.18%, rising from HK$300.4 to HK$349.
The market movement suggests a reassessment of its commercialization capabilities. The open release of the technology, the expansion of the ecosystem, and the emergence of real-time as a new use case can be seen as having been recognized as a path to monetization. The rise in share price reflects not just a technological achievement but expectations for the sustainability of the business.
Along with a notice prohibiting unauthorized reproduction, MiniMax’s trajectory is set to be a focal point for the future video generation market. By crossing the threshold where generation exceeds playback, applications, ecosystem, and market valuation have begun to move in tandem.
Editorial Opinion
In the short term, we expect prototypes for streaming and short-form viewing premised on real-time generation to surge within 3 to 6 months. Once conditions are in place for generation to outpace playback, as with H3 Max, operations without reliance on pre-production or inventory become possible. We assess that developers will accelerate experimentation in areas such as streaming, interactive storytelling, and instant-generation feeds. It appears we have entered a phase where operational expertise in inference costs and stability will determine competitiveness.
In the long term, we see the competitive landscape for video generation being determined less by performance alone and more by ecosystem and stickiness. Preference-based selection and high switching costs encourage lock-in to specific models. As derivative models centered on open source accumulate, assets and workflows will become tied to particular foundations, raising barriers to entry for latecomers. We believe there is potential for integration into standard production workflows within 1 to 3 years.
The remaining question is whether real-time generation can establish itself as a sustainable revenue model. Can generation costs be covered by user payments or advertising? How will copyright and likeness rights be handled? How will variations in quality be controlled? These issues cannot be solved by technological speed alone. We assess that what is being tested is the design of who bears the costs and who captures the value.
References
- “3秒出片比播放还快,MiniMax打开了AI视频的实时商业化路径”, by 克雷西 — 量子位, 2026-09-01T12:02:39.000Z (ARR)
- Source URL: https://www.qbitai.com/2026/09/482512.html
Frequently Asked Questions
- How fast is H3 Max's generation speed?
- It generates a 5-second video at 768p in under 3 seconds. Throughput is said to be about 35 times that of the base H3. Because generation time is shorter than playback time, the next video is completed while the previous one is playing, enabling uninterrupted continuous supply.
- Who developed H3 Max and how?
- It was jointly developed by MiniMax and AI inference infrastructure company fal. Based on H3, which MiniMax released at the end of July 2026, they applied post-training and inference optimization specialized for real-time generation. It is an effort combining an open-source foundation with expertise in accelerating inference.
- What applications is real-time generation being used for?
- It is beginning to be used for live streams where all material is generated on the spot, on-demand production in response to viewer requests, streams where stories branch interactively, and short-form video feeds that switch with a swipe using AI-generated content. In all cases, waiting time is eliminated because generation outpaces playback.
Comments