MiniMax Music 3.0 launched August 13, 2026 with open weights on HuggingFace, a $0.15-per-song API, and songs up to five minutes. The API endpoint docs still reference Music 2.6, and local inference can demand up to 24 GB of VRAM.
What's New in Music 3.0 (and Why 5-Minute Songs Matter)
Music 3.0 replaces Music 2.6 as MiniMax's recommended music model. The headline change: songs up to five minutes in a single generation, with composition, arrangement, performance, and production all produced at once. MiniMax positions this as solving a specific problem in earlier versions: losing genre, instrumentation, and vocal consistency as songs progressed past the first section (official blog).
| Component | Parameters | Role |
|---|---|---|
| Global LLM (from Qwen3.5-8B) | 8B | Predicts musical structure and semantics across an entire song |
| Local LLM | 0.6B | Fills in fine acoustic detail within each frame |
| Flow-matching + Flow-VAE | 2.4B + 123M | Converts hidden states into continuous audio (not discrete tokens) |
Underneath, an 8-layer residual vector quantization (RVQ) stack separates coarse musical structure (Layer 1, 16,384-entry codebook) from fine acoustic detail (Layers 2–8, 1,024 entries each). The global LLM tracks song-level context while the local LLM handles per-frame acoustic prediction, and the flow-matching stack renders final audio from fused hidden states rather than decoding from discrete tokens alone. The blog claims this reduces the high-frequency artifacts and rigid phrasing that plagued earlier AI music models, but MiniMax has not published controlled benchmarks against Music 2.6, Suno, or Udio.
API Pricing, Model IDs, and Rate Limits
MiniMax's official platform lists two Music 3.0 model IDs with straightforward pay-as-you-go billing:
| Model ID | Price per Song | Rate Limit | Notes |
|---|---|---|---|
Music-3.0 | $0.15 | 120 RPM | Contact sales for higher limits |
Music-3.0-free | $0 | 3 RPM | Free trial tier |
Music-3.0 lyrics generation/editing | $0.01 | — | Auto-generate or optimize lyrics |
Both models generate a complete song (up to five minutes with vocals) per API call. The endpoint is POST /v1/music_generation, the same one Music 2.6 uses. As of launch day, the API reference docs still show music-2.6 in their indexed examples, meaning developers may need to test the Music-3.0 ID manually until documentation catches up.
Running Music 3.0 Locally: ComfyUI and Hardware Requirements
The open-weight release is what got the community most excited. Weights are available on HuggingFace at MiniMaxAI/MiniMax-Music3, with a ComfyUI-optimized distribution at Comfy-Org/MiniMax-Music-3.
Three hardware tiers are documented in the model card and corroborated by early community reports:
| Mode | VRAM | Approx. Speed | Quality |
|---|---|---|---|
| Full precision | <24 GB (e.g., RTX 4090) | Fastest | Best |
| CPU offloading | ~22 GB GPU + 32 GB+ RAM | Moderate | Near-full |
| Layer streaming | 8 GB GPU + 32 GB+ RAM | Slow | Constrained |
To get started locally: install ComfyUI 0.33.0 or later (confirmed by the ComfyUI community thread), download the model files from one of the HuggingFace repos above, load the ComfyUI workflow for MiniMax-Music3, select your precision mode based on available VRAM, and generate. CUDA is currently the only supported compute backend; no ROCm or MPS path exists yet (per community reports). A Diffusers integration PR is in progress, which would open up Python-native inference outside ComfyUI.
One Reddit user on r/comfyui reported generating 140 seconds of country-style music on an AMD RX 7900 GRE (16 GB VRAM, 32 GB RAM) in approximately 125 seconds. Another thread on r/StableDiffusion summed up the community sentiment bluntly:
"Today I have unsubscribed from Suno thanks to Minimax Music." — r/StableDiffusion
The license terms deserve a read before commercial use. The blog says "open weights" and "production-ready" but does not publish the license text, commercial-use clauses, or redistribution terms. The HuggingFace repository's LICENSE file is the authoritative source; verify it before shipping revenue-generating work.
How to Prompt Music 3.0: Structured Captions and Section Tags
Music 3.0's control interface centers on Structured Captions: detailed, temporally organized descriptions that tell the model what should happen in each section of the song. According to the official blog, a well-formed caption specifies:
- Genre and sub-genre (e.g., "progressive house / EDM")
- Tempo and key (e.g., "126 BPM, B-flat major")
- Instrumentation (e.g., "breathy male tenor, side-chained synths, club bass, hall reverb")
- Vocal delivery (e.g., "gravelly tenor with falsetto breaks")
- Emotional arc (e.g., "airy verse building to high-energy chorus")
- Production character (e.g., "vintage room mix," "glossy keys")
Lyrics use standard section tags that most music AI models recognize:
[intro]
[verse]
[pre-chorus]
[chorus]
[bridge]
[instrumental]
[solo]
[outro]
A Prompt Enhancement System can expand simpler requests ("upbeat pop song with female vocals") into the full Structured Caption format using music-industry terminology. This lets non-musicians access detailed control without knowing specialist vocabulary. Instrumental tracks can be generated without lyrics; vocal tracks require either user-supplied lyrics or the $0.01/song lyrics optimizer.
MiniMax Music 3.0 vs Suno and Udio
The comparison that matters: should you pay $0.15 per song to MiniMax, or subscribe to an established player?
| Feature | MiniMax Music 3.0 | Suno | Udio |
|---|---|---|---|
| Max song length | 5 min | ~4 min | ~2–4 min |
| Pricing model | $0.15/song (pay-as-you-go) | $10/mo Pro (500 songs) | $10/mo Standard (~600 songs) |
| Free option | 3 RPM API + open weights | 10 songs/day | Limited credits |
| Open weights | Yes (HuggingFace) | No | No |
| Local deployment | Yes (ComfyUI) | No | No |
| Commercial use | Check HuggingFace license | Included in paid plans | Included in paid plans |
The break-even point sits around 67 songs per month. Below that, MiniMax's pay-as-you-go at $0.15/song is cheaper than any subscription. Above it, Suno Pro or Udio Standard at $10/month for 500+ songs wins on raw cost. But MiniMax's open-weight local option is the wild card: unlimited free generation if you already own the GPU.
Community threads on r/SunoAI and r/comfyui paint Music 3.0 as competitive with Suno v4 on vocal naturalness and prompt adherence, with praise for electronic and pop genres. Rock, metal, and older acoustic styles drew more mixed reactions, with some users noting muddiness in dense mixes. Multilingual support (Mandarin and English in official demos, plus community-tested Bollywood rap) appears strong.
FAQ
Is MiniMax Music 3.0 free?
The Music-3.0-free API model generates songs at no cost (3 RPM limit). You can also run the open-weight model locally for free via ComfyUI.
Can Music 3.0 generate instrumental tracks?
Yes. Instrumental generation does not require lyrics; vocal generation requires user-supplied lyrics or the $0.01/song lyrics optimizer.
What is the API model ID for Music 3.0?
Music-3.0 (120 RPM, $0.15/song) and Music-3.0-free (3 RPM), both via POST /v1/music_generation. Note that the API docs still reference music-2.6 in some pages.
Does Music 3.0 support languages other than English?
Official demos include Mandarin and English. Community tests show successful generation in Hindi (Bollywood rap test). No official language list is published.
Can I use MiniMax Music 3.0 tracks commercially?
Check the HuggingFace LICENSE file and the API terms of service. The blog does not specify commercial-use rights.
How much VRAM do I need to run Music 3.0 locally?
Under 24 GB for full precision, ~22 GB with CPU offloading, or 8 GB in layer-streaming mode. See the hardware table above.
Which Route Should You Pick?
| Your Situation | Recommended Route | Cost |
|---|---|---|
| Testing or prototyping a few songs | Music-3.0-free API | $0 |
| Occasional production use (<67 songs/month) | Music-3.0 API | $0.15/song |
| High-volume generation (>67 songs/month) | Suno Pro or Udio Standard | $10/month |
| Own a 24 GB GPU, want unlimited free generation | Local ComfyUI with HuggingFace weights | $0 (compute only) |
| Building an app with programmatic music generation | Music-3.0 API via MiniMax platform | $0.15/song |
Two gaps remain: the API documentation lag (still on Music 2.6) and the unspecified license terms on the open weights. Both should resolve as the release stabilizes, but neither blocks the API from working today.