Annual Offer
Annual option: 50% discount
AUG 202633B omni-modal · 4–15s · 24 FPS · 32 kHz stereo

MiniMax H3: 2K Video With Native Stereo Audio in One Pass

MiniMax H3 — shipped as Hailuo 3.0 on July 31, 2026 — is a 33-billion-parameter omni-modal model that reads text, images, video and audio as a single context, then returns video with the sound already generated inside it. This page covers what MiniMax H3 costs, what its open weights actually permit, and how it holds up against Veo 3.1, Sora 2 and Kling 3.0.

4–15 second clips at 24 FPS, up to 2KNative 32 kHz stereo — no second audio passOpen weights, four excluded territories
PROMPT REPRODUCTION LAB · 09 VIDEO RECIPES

MiniMax H3 Prompts You Can Reproduce, Not Just Admire

Nine real video examples pair the finished motion with the complete prompt behind it. Copy a recipe as-is, or send it straight to the video workspace and replace the subject, setting, camera or dialogue without starting from a blank box.

05 · By the numbers

Three MiniMax H3 Numbers Worth Memorizing

The figures that decide most MiniMax H3 build-or-buy conversations, current as of August 2026.

2K
Hosted output
Native Stereo
Audio in one pass
2K
Top MiniMax H3 output tier, reached by in-context regeneration of the native 768p base
4–15s
Clip range at 24 FPS with native 32 kHz stereo audio
42.5 GB
Smallest practical local MiniMax H3 weight set
01 · Workflows

Four Ways to Reach MiniMax H3 — and One Legal Catch

MiniMax H3 ships through a consumer app, a metered API, and a downloadable checkpoint. Each route carries a different price, a different ceiling, and a different license. Find the row that matches where you sit and what you are allowed to publish.

01

Hailuo AI App — the No-Setup Route

The consumer Hailuo app carries MiniMax H3 behind a credit system, with a free tier and paid monthly plans — check the app for current rates. It is the quickest way to see the model behave, but it exposes no endpoint, seed or resolution flag — which makes it weak ground for a reproducible comparison.

Prepare a Test Frame
02

Open Platform API — Billed Per Output Second

MiniMax bills H3 by output second, with the 2K tier priced above 768p. The model ID is MiniMax-H3. Prompts cap at 7,000 characters and the request body at 64 MB. The 768p tier has been intermittently gated, so confirm the live rate card before budgeting a batch.

Plan a Budget
03

Local Weights — 42.5 GB to Get Moving

MiniMax published H3 checkpoints in early August 2026. A minimum working set lands near 42.5 GB; both checkpoints at the smallest useful precision run about 63.4 GB, and full bf16 research weights reach roughly 123.6 GB. ComfyUI reports 12 GB VRAM as a floor with aggressive offload, though 64 GB of system RAM makes it bearable.

Check Hardware Notes
04

Blocked Where You Are? Read This First

The MiniMax H3 community license excludes the United States, the European Union, the United Kingdom and the Republic of Korea, and hosted access was withheld in those markets at launch. Inside one of them, MiniMax H3 is a model to understand rather than deploy — and Sora 2, Veo 3.1 or Seedance 1.5 Pro keep the work clean.

See Available Models
02 · Model overview

What Is MiniMax H3?

MiniMax H3 is a 33-billion-parameter dense omni-modal transformer that treats text, images, video and audio as one stream instead of bolting separate pipelines together. It generates 4- to 15-second clips at 24 FPS with 32 kHz stereo sound, reaching 768 pixels on the short edge locally and 2K through a hosted regeneration stage. Counting encoders and VAEs, the full MiniMax H3 inference stack sits closer to 69 billion parameters.

MiniMax H3 architecture overview covering the omni transformer, VAE and 2K regeneration stage

Sound Is Generated, Not Attached

The H3-Omni Transformer produces voice, effects and music inside the same pass as the picture, at a documented 32 kHz stereo. That removes an entire production stage compared with models that need a separate audio run — though it does not exempt the result from review for timing, intelligibility and channel balance.

H3-VAE Bought a 4× Longer Sequence

MiniMax rebuilt the tokenizer for H3, reporting a fourfold gain in effective sequence length. That is the structural reason a 15-second clip with synchronized sound is tractable at all, and why end-to-end training throughput rose nearly 30% against the previous generation.

Captions Compress 100K Tokens Into 4K

H3-Contextual Omni Representation uses language as a bridge between modalities: roughly 100K tokens of inference are distilled to about 4K on average before generation. This is what lets a mixed reference set — stills, clips and audio — read as one instruction rather than four competing ones.

2K Arrives by Regeneration, Not Upscaling

The base MiniMax H3 model renders at 768 pixels on the short edge. Rather than bolting on a super-resolution module, H3 regenerates that output in-context to reach 2K. Hosted API calls can request 2K directly; a local deployment produces 768p unless you run the regeneration stage yourself.

03 · Model comparison

MiniMax H3 vs Veo 3.1, Sora 2 and Kling 3.0

Four head-to-head reads using published rate cards and specifications as of August 2026. Prices and limits move, so verify with each provider before committing spend. Where a rival model is included in this site’s plans, that is stated plainly rather than implied.

FeatureVeo 3.1Sora 2MiniMax H3
Max resolution1080p2K
Clip lengthUp to 25s4–15s
Native audioYesNo32 kHz stereo
Reference inputsFramesVideo remixImages + video + audio
Access modelHostedHostedHosted + weights

MiniMax H3 vs Veo 3.1

Veo 3.1 is the closest match on capability, because it also generates audio with the picture. The split is resolution, price and availability: MiniMax H3 can reach 2K through its regeneration stage, while Veo 3.1 tops out near 1080p — and Veo’s Standard tier bills several times H3’s per-second rate, with even Fast priced above H3’s 2K output.

Veo 3.1 holds a genuine advantage in synchronized dialogue and in predictable camera language — years of published prompt patterns mean fewer surprises when a shot has to match a storyboard. MiniMax H3 answers with regenerated 2K output, a 15-second ceiling, and a reference contract that accepts up to nine images, three video clips and three audio files in a single call.

The deciding factor is usually jurisdiction rather than craft. Veo 3.1 is available worldwide through Google; MiniMax H3 is not licensed in the US, EU, UK or South Korea. Veo 3.1 is included in this site’s plans, so a team blocked from H3 can still run the equivalent brief without leaving the workflow.

Generate with MiniMax H3
CASE 01 · THE ONLY OTHER MODEL WITH SOUND BUILT IN
04 · Capabilities

MiniMax H3 Specifications

Documented model-family facts as of August 2026, drawn from the MiniMax H3 repository and Open Platform documentation. Verify against the exact variant and provider you intend to use — these describe the model, not this website’s interface.

/01

33B Dense Omni-Modal Transformer

MiniMax H3 runs a dense 33-billion-parameter architecture handling text, image, video and audio in one context. With encoders and VAEs included, the complete inference stack reaches roughly 69 billion parameters.

/02

2K Output, Metered Per Second

Hosted MiniMax H3 generation bills by output second, with the 2K tier priced above 768p. The Context-IR variant is priced separately by token.

/03

Native 32 kHz Stereo Audio

Voice, effects and ambience are generated with the picture rather than added afterwards. This is the capability that most clearly separates MiniMax H3 from Sora 2 and Kling 3.0, both of which need a separate audio stage.

/04

4 to 15 Seconds at 24 FPS

Durations run from four to fifteen seconds in whole numbers, at a fixed 24 frames per second. Native multi-shot modelling means a single MiniMax H3 clip can contain more than one camera setup.

/05

Six Documented Aspect Ratios

Official documentation names 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16 as examples rather than a closed set, and image-conditioned flows may use adaptive framing. Confirm which controls your endpoint exposes.

/06

Community License, Four Exclusions

MiniMax H3 weights are downloadable but not OSI open source. The license excludes the US, EU, UK and South Korea, requires visible attribution for commercial use, bars distillation, and gates high-revenue organizations behind written approval.

06 · How to use

How to Run a MiniMax H3 Test That Means Something

Five steps that keep a MiniMax H3 evaluation honest and repeatable. The Hero video workspace can now run text-to-video, image-to-video and multimodal reference briefs through KIE; the remaining steps keep the result measurable and reviewable.

MiniMax H3 evaluation worksheet listing visual, motion and audio criteria
STEP 01

Write the Acceptance Criteria First

List the subject traits, the required action, one camera behavior, the audible event and the failure conditions before you touch an endpoint. Without this, an attractive but off-brief MiniMax H3 result will pass review and cost you the reshoot later.

MiniMax H3 FL2VA and Ref2VA input route decision diagram
STEP 02

Pick FL2VA or Ref2VA Deliberately

MiniMax H3 exposes two input routes. FL2VA takes a text brief with zero, one or two images — one anchor sets the opening state, two define endpoints. Ref2VA is for mixed evidence: images, clips and audio that carry identity or behavior text cannot express.

MiniMax H3 reference budget showing per-category caps and twelve-file total
STEP 03

Respect the Twelve-File Ceiling

Ref2VA caps images at nine, video at three and audio at three — but a separate rule limits the combined set to twelve files, so those maxima cannot be summed to fifteen. Audio also requires an accompanying image or video. Give every reference one job and delete the rest.

MiniMax H3 prompt structure separating subject action, camera move and sound event
STEP 04

Name One Sound and One Camera Move

A prompt like “the cup meets the table with a soft ceramic knock” creates a testable relationship between sound and image; “cinematic audio” creates nothing you can score. Keep dialogue short enough to fit inside the 4-to-15-second window you selected.

MiniMax H3 local 768p and hosted 2K output paths labeled for review
STEP 05

Record Which Resolution Path You Used

A local MiniMax H3 deployment yields 768p until you run regeneration; a hosted call can request 2K directly. These are different artifacts and comparing them is meaningless. Log the provider, endpoint, variant and resolution against every clip you keep.

Decision Notes

MiniMax H3 Decision Notes: Nine Pipeline Scenarios

Nine composite adoption scenarios — where MiniMax H3's 2K pricing beats Veo 3.1, Sora 2 and Kling 3.0, and where the license, reference and hardware limits bite. Illustrative decision notes drawn from the published specs and comparisons above, not verified customer reviews.

We were paying Veo 3.1 Standard rates for 1080p social cuts. MiniMax H3 renders 2K at a fraction of that per-second rate, so the same monthly budget now covers three times the variants. Dialogue-heavy spots still go to Veo — H3 wins on cost per pixel, not on lip sync.
Social ad variants at 2KBudget moved from Veo 3.1 Standard
I read the license before I read the benchmarks, and I would tell anyone to do the same. H3 excludes the US, EU, UK and South Korea. Our Singapore team ships on it daily; our London team runs Veo 3.1 instead. Same brief, two pipelines.
License review before benchmarksUS / EU / UK / KR exclusions weighed
We lost an afternoon to this one. Nine images plus three videos plus three audio files is not a fifteen-file bundle — the combined MiniMax H3 ceiling is twelve, and every audio reference needs a visual alongside it. Once the team budgeted references properly, the rejections stopped.
Reference-cap collision9 + 3 + 3 inputs, 12-file ceiling
The first thing I check now is whether the story fits in fifteen seconds. If it does, MiniMax H3 is my default. If it needs the full twenty-five, I go to Sora 2 Pro and accept the bill — no amount of native audio saves a clip that gets cut off mid-beat.
Short-form narrative defaultFifteen-second ceiling as the filter
The 12 GB VRAM floor had me convinced this would run on a workstation. Then I read the host RAM figures and the 42.5 GB minimum working set. Two RTX 5090s still need about nine minutes per five-second clip, so hosted inference paid for itself before lunch.
Self-hosting reality check12 GB VRAM floor vs host RAM math
Worth knowing the 2K is not an upscale — MiniMax H3 regenerates its 768p base output in-context, which is why the detail actually holds up. It also means our local deployment stays at 768p until we run that second stage ourselves.
2K pipeline due diligenceIn-context regeneration, not upscale
Half my work is silent — background plates, product turntables, hero loops — and for those I still run Kling 3.0, which stays cheaper per second. MiniMax H3 earns its price the moment a brief actually names a sound, and then it saves me an entire audio pass.
Silent plates and product loopsKling 3.0 kept for audio-free work
MiniMax H3 ranked first in video editing with audio judged and second in text-to-video when we pulled arena tracking in August 2026. I read that as a tie at the top rather than a blowout — the confidence intervals overlap its nearest rival.
Arena-ranking sanity checkAudio-judged first place verified
Two things I flag for every client shipping MiniMax H3 output commercially: attribution has to be visible in the product interface rather than buried in a terms page, and past roughly $20M in annual revenue you need separate written authorization. Neither is difficult; neither is optional.
Commercial-terms flagVisible attribution and revenue clause

MiniMax H3 Pricing

08 · Pricing

Choose a Credit Term Before You Choose a Tier

Estimate the requests you expect to submit, read the cost shown by the relevant generator, then compare total charge, credit allowance, validity, and renewal cadence on the cards below.

/ Starter — Annual

Lowest annual charge and allowance in the recurring table

$9.9$19.8/month
credits9,600
  • MiniMax H3 Basic Yearly Plan
  • Monthly credit allocation: 800 credits; yearly total: 9600 credits
  • Monthly image maximum: 80 images
  • MiniMax H3 AI Generation
  • Unlimited MiniMax H3 downloads
  • 9600 MiniMax H3 credits with annual billing
  • Permanent MiniMax H3 history
  • License covering commercial use
Most Popular
/ Studio — Annual

Highest annual allowance for buyers with sustained usage

$44.9$89.8/month
credits72,000
  • MiniMax H3 Enterprise Yearly Plan
  • Monthly credit allocation: 6000 credits; yearly total: 72000 credits
  • Monthly image maximum: 600 images
  • MiniMax H3 AI Generation
  • Unlimited MiniMax H3 downloads
  • 72000 MiniMax H3 credits with annual billing
  • Permanent MiniMax H3 history
  • License covering commercial use
/ Pro — Annual

Annual middle tier between Starter and Studio allowances

$19.9$39.8/month
credits24,000
  • MiniMax H3 Pro Yearly Plan
  • Monthly credit allocation: 2000 credits; yearly total: 24000 credits
  • Monthly image maximum: 200 images
  • MiniMax H3 AI Generation
  • Unlimited MiniMax H3 downloads
  • 24000 MiniMax H3 credits with annual billing
  • Permanent MiniMax H3 history
  • License covering commercial use
09 · Updates

Track MiniMax H3 Changes

MiniMax H3 pricing, license terms and regional availability have all moved since launch. Subscribe for occasional notes when the documented specifications or access conditions change.

No spam. Unsubscribe in one click.
10 · FAQ

MiniMax H3: Frequently Asked Questions

Answers drawn from the MiniMax H3 repository, Open Platform documentation and independent August 2026 testing. Verify the live contract for your provider and territory before implementation.

Need more help?
Question not covered above? Our support team reads every message.

MiniMax H3 is a 33-billion-parameter dense omni-modal model released on July 31, 2026, and also sold as Hailuo 3.0. It reads text, images, video and audio as a single context and generates 4- to 15-second video at 24 FPS with native 32 kHz stereo audio. The API model ID is MiniMax-H3.

MiniMax H3 is billed per second of output on the MiniMax Open Platform, with 2K priced above the 768p tier; check the live rate card for current figures, since the 768p tier has been gated at times. The Context-IR variant is priced separately by token. On this site, MiniMax H3 generation is credit-metered — 18 credits per second at 768p and 29 credits per second at 2K — with credit packages listed on the pricing page.

Not in the OSI sense. MiniMax published downloadable H3 weights in early August 2026 under a community license. Commercial use requires visible “MiniMax H3” attribution in your interface, model distillation is prohibited, organizations above roughly $20M in annual revenue need separate written authorization, and four territories are excluded outright.

The MiniMax H3 community license excludes the United States, the European Union, the United Kingdom and the Republic of Korea, and hosted access was withheld in those markets at launch — MiniMax cited evolving AI regulation and copyright concerns. Teams in those territories should treat H3 as a model to understand, and use licensed alternatives such as Sora 2, Veo 3.1 or Seedance 1.5 Pro for production work.

Both generate audio with the picture, which puts them in a category of two. MiniMax H3 reaches 2K through its regeneration stage at the lower per-second list rate and accepts richer reference input. Veo 3.1 tops out near 1080p and bills more per second on both of its tiers, but leads on synchronized dialogue, camera predictability, and worldwide availability. For most teams the license question settles it before quality does.

Sora 2 Pro reaches about 25 seconds against H3’s 15-second ceiling and is stronger on physical plausibility, but it generates no audio in the same pass. On list price, MiniMax H3’s 2K rate sits just above Sora 2’s base tier and well below Sora 2 Pro. Once a separate audio stage is priced in, H3 often wins on cost per finished second.

A minimum working set of MiniMax H3 weights is about 42.5 GB, rising to 63.4 GB for both checkpoints and 123.6 GB at full bf16. ComfyUI reports a 12 GB VRAM floor with aggressive offloading, though a 24–32 GB card wants roughly 75 GB of host RAM for INT8 offload. Four H100 80GB cards complete a generation in about 13 seconds; two RTX 5090s take closer to nine minutes for five seconds.

Independent arena tracking in early August 2026 placed MiniMax H3 first in video editing with audio judged, second in text-to-video, and second or third in image-to-video depending on whether audio counted. MiniMax published no official benchmarks of its own, and the confidence intervals overlap its nearest competitor, so read the ranking as close rather than settled.

The Ref2VA route caps images at nine, video clips at three and audio files at three, with a separate combined limit of twelve files — so the category maxima cannot be summed. Reference video and audio each run 2 to 15 seconds, audio requires an accompanying image or video, prompts cap at 7,000 characters, and the request body at 64 MB.

Yes. After sign-in, the Hero workspace can create MiniMax H3 text-to-video, first/last-frame image-to-video, and multimodal Reference-to-Video tasks through KIE. Reference mode accepts images, MP4 or MOV video, and MP3 or WAV audio; generation supports 4–15 second duration and 768P or 2K output. Generation uses this site’s credit balance, while availability and permitted use still depend on your territory and the applicable MiniMax terms.

11 · Get started

Write the Shot, Then Run MiniMax H3

Turn a complete shot description, first and last frames, or a mixed image/video/audio reference set into a MiniMax H3 task in the Hero workspace. Compare Sora 2, Veo 3.1 and Seedance 1.5 Pro when territory, budget or delivery requirements call for another route.