Chinese artificial intelligence company MiniMax has released MiniMax-Music3, an open-weights music generation model capable of creating complete songs of up to five minutes from lyrics and a detailed description of the desired music. The model is designed to generate vocals, instrumentation, arrangements and a full song structure in a single generation, marking another step toward more capable AI systems for end-to-end music production.
MiniMax-Music3 produces 32 kHz, 16-bit stereo WAV audio and is designed to maintain musical structure and audio quality over longer generations. Unlike basic AI music tools that focus on short clips or individual musical elements, the new model is aimed at producing complete tracks with evolving arrangements and expressive vocals. The release also makes the model weights and inference resources available, giving developers and creators the option to experiment with the technology outside a closed consumer application.
MiniMax-Music3 Generates Complete Five-Minute Songs
MiniMax-Music3 is built to generate a complete song from two main inputs: lyrics containing structural section tags and a detailed description of the music.
The description can specify elements such as genre, mood, instrumentation, vocal characteristics, arrangement and overall production style.
The model then combines those inputs to produce a finished audio track of up to five minutes in a single generation.
This approach is different from conventional AI music systems that may generate short clips that need to be extended or assembled manually.
| Feature | MiniMax-Music3 |
|---|---|
| Model type | Open-weights text-to-music |
| Maximum song length | Up to 5 minutes |
| Audio format | 32 kHz, 16-bit stereo WAV |
| Inputs | Lyrics + structured music description |
| Output | Complete song |
| Vocals | Supported |
| Weights | Available |
| Commercial use | Allowed under community license |
Structured Inputs Give Creators More Control
One of the important design choices behind Music3 is the separation of lyrics from the description of the music itself.
Lyrics can include section tags that indicate the structure of a song, such as verses, choruses and other sections.
The second input provides information about how the music should sound.
This separation allows users to control both the words and the musical direction without putting everything into a single unstructured prompt.
How Music3 Works
Lyrics with section tags
↓
Detailed music description
↓
MiniMax-Music3
↓
Vocal generation
+
Instrumentation
+
Arrangement
+
Song structure
↓
Complete music track
↓
Up to 5 minutes
The structure is particularly useful for creators who want more predictable song organization.
Open Weights Make the Model More Accessible
MiniMax has released the model with open weights, inference code and documented deployment options.
This is important because users are not limited to accessing the model through a single hosted interface.
Developers can experiment with the model, integrate it into their own workflows and investigate how the underlying system works.
The release also provides several serving options depending on available hardware.
MiniMax-Music3 can run through SGLang-Omni using two GPUs, through Diffusers with less than 24 GB of memory, or with group offloading using configurations requiring around 8 GB.
This gives the model a wider range of potential deployment scenarios.
Hybrid Architecture Combines Multiple AI Components
Music3 uses a Hybrid-LM architecture that combines an 8-billion-parameter Global LLM with a smaller 0.6-billion-parameter Local LLM.
The language components work with a continuous music-generation stack based on flow matching and a Flow-VAE.
The architecture is designed to handle musical structure and audio synthesis separately while allowing the systems to work together during generation.
The synthesis process operates on continuous hidden states rather than relying entirely on a conventional discrete audio-token decoding pipeline.
This design can help the system maintain longer musical sequences while producing detailed audio.
Flow Matching Handles Audio Generation
Flow matching is used as part of the continuous synthesis process.
In simple terms, the system learns how to transform a noisy representation into a coherent musical signal.
The approach has become increasingly important in generative AI because it can be applied to complex continuous data such as images, audio and other media.
For music generation, maintaining consistency across several minutes is particularly challenging because the model must preserve musical relationships while continuously introducing new sounds and arrangements.
Longer Songs Remain a Major Technical Challenge
Generating a short musical clip is considerably easier than maintaining a coherent five-minute composition.
Longer tracks require the model to remember what happened earlier in the song while continuing to produce appropriate melodies, vocals and instrumentation.
It also needs to avoid abrupt changes that make the track sound like a collection of unrelated clips.
Music3 is designed specifically around this long-form generation problem.
The objective is to produce tracks with evolving arrangements while maintaining a consistent overall musical identity.
Vocal Generation Is a Core Capability
Music3 can generate songs with vocals rather than limiting output to instrumental music.
The model can use provided lyrics and transform them into sung performances that fit the generated composition.
This makes the system relevant to a broader range of applications, including songwriting, advertising, social-media content, video production and entertainment.
The ability to combine lyrics, vocals and instrumental arrangements in one generation also reduces the number of separate tools required to produce a complete track.
Multilingual Music Creation Could Expand Its Use
The model can also be used for vocal generation in languages beyond English.
Reports following the release demonstrated Japanese vocal generation, highlighting the potential for the technology to support international creators.
Multilingual music generation could become increasingly important as AI music platforms expand beyond English-speaking markets.
For creators, the ability to generate songs in local languages could make AI music tools more useful for regional advertising, entertainment and online content.
MiniMax Is Expanding Beyond Text and Video
Music3 is part of MiniMax’s broader multimodal AI strategy.
The company has developed models covering text, audio, image, video and music.
Its earlier releases have focused heavily on general-purpose AI and video generation, while Music3 expands the company’s presence in the rapidly developing AI music market.
The strategy reflects a broader industry trend in which AI companies are attempting to provide multiple creative modalities through a common technology ecosystem.
AI Music Competition Is Intensifying
MiniMax is entering a market that already includes several established AI music platforms.
Companies such as Suno and Udio have helped popularize consumer-facing AI music generation, while other startups and research organizations are developing open and specialized models.
Competition is increasingly moving beyond simple song generation.
The focus is now shifting toward controllability, audio quality, longer compositions, editing capabilities, licensing and integration into professional workflows.
Open Models Could Change the Competitive Landscape
Closed AI music platforms generally control access to their models through web applications or APIs.
Open-weight models create a different model.
Developers can potentially run them locally, build custom interfaces and integrate them into specialized production pipelines.
That could encourage experimentation and lead to new applications that are difficult to build using closed systems.
However, open deployment also places greater responsibility on developers to manage hardware, inference performance, licensing and content policies.
Commercial Use Comes With Licensing Conditions
MiniMax-Music3 is released under a community license that permits commercial use, but it includes conditions.
Products using the model must prominently display “MiniMax-Music3” in the product interface.
Organizations whose aggregate annual revenue from products using the model exceeds $20 million must obtain separate prior written authorization from MiniMax.
These requirements are important for companies considering Music3 for commercial products.
The availability of model weights therefore does not mean that the model can be used without reviewing its licensing conditions.
Hardware Requirements Still Matter
Although MiniMax provides multiple deployment paths, generating high-quality music locally can still require significant computing resources.
Audio generation over several minutes is computationally demanding.
Users with more powerful GPUs can generally expect greater flexibility in running the model, while lower-memory configurations may require techniques such as offloading.
This means that local deployment is becoming more accessible, but it is not yet equivalent to running a lightweight desktop application.
AI Music Could Lower Production Costs
One of the biggest commercial opportunities for models such as Music3 is reducing the cost and time required to produce certain types of music.
A company creating a promotional video, for example, could generate several original background tracks and select the most suitable option.
Independent creators could also produce demos, soundtracks and complete songs without hiring a large production team.
For small businesses and creators, this could make professional-sounding audio more accessible.
Advertising Is a Major Use Case
Advertising agencies increasingly need music that can be adapted to different campaigns and audiences.
AI-generated music can potentially provide customized tracks for specific products, moods and campaign lengths.
A brand could generate multiple versions of a song for different advertisements or platforms.
The technology could therefore become part of a broader AI-assisted advertising workflow that already includes image, video and voice generation.
Independent Musicians Could Use AI as a Production Tool
Music3 could also appeal to independent musicians who want to experiment with arrangements or generate early versions of songs.
The technology does not necessarily have to replace human musicians.
It can instead serve as a creative assistant for exploring ideas before a final human-produced version is recorded.
However, musicians will need to consider licensing, originality and ownership issues when incorporating AI-generated material into commercial releases.
Copyright Questions Remain Important
AI-generated music continues to raise questions about copyright and the training data used to build generative models.
The industry is still developing legal and commercial frameworks around AI-generated works.
Creators and businesses using Music3 will therefore need to understand the applicable licensing terms and the legal status of AI-generated content in their markets.
The availability of open weights does not eliminate these broader legal questions.
AI Music Is Moving Toward Full Production
The development of Music3 illustrates how generative AI is moving from producing isolated creative elements toward complete production workflows.
Earlier systems often focused on generating a short melody, instrumental clip or voice.
Newer systems are attempting to coordinate lyrics, vocals, instruments, structure and mastering-like processes within a single generation.
This could eventually turn AI music models into production environments rather than simple generation tools.
What Developers Will Watch
Developers evaluating Music3 will likely focus on several areas:
- Audio quality
- Long-form consistency
- Vocal realism
- Prompt controllability
- Hardware requirements
- Inference speed
- Licensing conditions
- Local deployment
- Commercial usability
- Integration with creative software
The model’s performance in real-world workflows will determine whether it can compete effectively with established commercial platforms.
Key Facts at a Glance
| Metric | MiniMax-Music3 |
|---|---|
| Developer | MiniMax |
| Category | AI music generation |
| Release | August 2026 |
| Model type | Open-weights |
| Maximum generation length | Up to 5 minutes |
| Audio quality | 32 kHz, 16-bit stereo |
| Output format | WAV |
| Global LLM | 8B parameters |
| Local LLM | 0.6B parameters |
| Flow matching component | 2.4B parameters |
| Flow-VAE | 123M parameters |
| Commercial use | Permitted with license conditions |
| Revenue threshold requiring authorization | $20 million annually |
Infographic: How MiniMax-Music3 Creates a Song
USER INPUT
↓
LYRICS
+
SECTION TAGS
+
DETAILED MUSIC DESCRIPTION
↓
MINIMAX-MUSIC3
↓
8B GLOBAL LLM
+
0.6B LOCAL LLM
↓
CONTINUOUS MUSIC SYNTHESIS
↓
FLOW MATCHING
+
FLOW-VAE
↓
VOCALS
+
INSTRUMENTS
+
ARRANGEMENT
+
SONG STRUCTURE
↓
COMPLETE TRACK
↓
UP TO 5 MINUTES
↓
32 kHz
16-BIT STEREO WAV
The Bigger Picture
MiniMax-Music3 represents another step in the rapid evolution of generative AI for music production. Its ability to create complete songs of up to five minutes from lyrics and a structured music description places the model closer to an end-to-end production tool than a simple audio generator. The open-weights release is particularly significant because developers can experiment with the technology locally and integrate it into their own creative workflows rather than relying exclusively on a hosted platform.
The release also highlights how competition in AI music is moving toward longer-form generation, better vocals, greater control and more accessible deployment. MiniMax is entering a market that already includes established commercial platforms, but its combination of open weights, local deployment options and complete-song generation could appeal to developers and independent creators. At the same time, licensing requirements, hardware costs and unresolved copyright questions remain important considerations for commercial adoption.
Looking Ahead
MiniMax’s next challenge will be turning Music3’s technical capabilities into a broader ecosystem of tools for creators and developers. Improvements in generation speed, audio quality, editing, controllability and hardware efficiency could make open AI music models increasingly practical for professional production. Integration with video, image and agentic AI systems could also allow creators to generate complete multimedia projects from a single workflow.
Over the longer term, models such as Music3 could change how music is created, particularly for advertising, social media, gaming, independent production and other areas where large volumes of customized audio are required. Human musicians and producers are unlikely to disappear from the process, but AI could increasingly handle experimentation, arrangement and production tasks. MiniMax’s decision to release Music3 with open weights could accelerate that transition by giving developers the foundation to build new AI-native music applications.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.


