Searching for MiniMax voice clone can lead to three different workflows: the browser-based MiniMax Audio tool, a Speech text-to-speech API that returns a reusable voice ID, or MiniMax H3’s reference-audio video workflow. They are not interchangeable. For repeatable narration, use Speech TTS; for a character inside a generated video, use H3; and for an application, choose an API route before recording your sample.
What “MiniMax voice clone” actually means
MiniMax voice cloning is primarily a text-to-speech workflow: you provide a speaker sample, create a custom voice, and pass the resulting voice identifier into later TTS requests. It is not automatically voice conversion, which preserves an original performance while changing only its timbre.
MiniMax H3 is a separate case. Its reference-audio workflow can guide a generated character’s voice, but community reports describe different failure modes from ordinary TTS cloning, including repeated reference lines and missing background audio.
| Workflow | Input | Output | Best fit |
|---|---|---|---|
| MiniMax Speech TTS | Reference audio + text | Reusable voice ID and synthesized speech | Narration, assistants, scripted audio |
| Hosted clone endpoint | Reference audio, usually a public URL or upload | voice_id / custom_voice_id and optional preview | Fast API experiments |
| MiniMax H3 reference audio | Audio reference + video prompt/workflow | Generated video with referenced voice | Character video and audiovisual scenes |
Pick the access path before you record
The access path changes the practical limits, billing, and lifecycle of the cloned voice. The table below separates facts documented by each provider rather than treating every “MiniMax voice clone” endpoint as the same service.
| Route | Useful documented facts | Cost or lifecycle caveat |
|---|---|---|
| MiniMax Audio voice cloning | Official interface; its public description calls for roughly 10–60 seconds of clean audio and a consent confirmation | Check the current MiniMax account credits and terms |
| Replicate MiniMax voice cloning | MP3, M4A, or WAV; schema lists 10 seconds to 5 minutes and under 20 MB; returns voice_id and a preview | The page lists $3 per clone output and says data is sent to MiniMax; its README’s 5-second claim conflicts with the schema’s 10-second minimum |
| fal MiniMax voice clone | At least 10 seconds; accepts MP3, OGG, WAV, M4A, and AAC; returns custom_voice_id; optional preview text is capped at 1,000 characters | $1.50 per clone plus $0.30 per 1,000 preview characters; fal documents a 7-day use requirement to retain a voice permanently |
| Direct MiniMax API via a Go example | Supports a file ID, local upload, or audio URL; default endpoint shown is https://api.minimax.io | The example warns that URL cloning may not work in the China region |
For production, do not copy a hosted provider’s voice_id into a different provider’s TTS request. The identifier belongs to the service that created it.
Prepare a reference clip that survives the API
Record one speaker in a quiet, dry room, keep the microphone distance stable, and avoid music, overlapping speech, and aggressive noise reduction.
Use 10 seconds as the safe minimum for API integrations. Replicate’s page contains a shorter marketing statement, but its input schema requires 10 seconds; fal and the public CLI documentation also use 10 seconds. A clean 20–40-second sample is a more useful starting point than a noisy clip that merely passes validation.
Before uploading, check:
- One speaker is present throughout the clip.
- The beginning and end are not clipped.
- Room echo and background music are minimal.
- The speaker uses the language and accent you need to synthesize.
- The file stays within the selected provider’s duration and size limits.
- You have the speaker’s permission and have checked the applicable commercial-use terms.
Clone once, synthesize many times
Clone once, store the returned voice ID securely, and reuse it for later text-to-speech calls instead of repeating the cloning operation.
A hosted endpoint such as fal exposes the same basic shape: send audio_url, optionally request preview text, then save custom_voice_id. Its API supports queue states such as IN_QUEUE, IN_PROGRESS, and COMPLETED, so a production integration should handle asynchronous completion instead of assuming an immediate response.
A direct MiniMax implementation commonly has these stages:
- Upload the audio and receive a file ID.
- Bind that file ID to a custom voice ID.
- Call the TTS endpoint with the voice ID and new text.
- Save the audio output and log the provider, model, and voice ID.
The public Go example documents three source-input modes: existing file ID, local upload, and audio URL. The community minimax-voice CLI describes the same upload → clone → TTS chain and recommends cloning once before repeated synthesis.
Keep fal and MiniMax API keys server-side; fal specifically warns against exposing FAL_KEY in browser or mobile code.
The failure modes to test before production
Test short and long sentences, numbers, names, punctuation, and the target language before committing to a workflow.
TTS cloning can sound right but still drift
A reference clip with room echo, multiple speakers, or an accent absent from the target script can cause timbre drift or pronunciation errors. Hosted schemas expose noise reduction and volume normalization, but treat those controls as experiments rather than guaranteed quality switches.
H3 reference audio can mix identity and content
Real-user reports for H3 describe failures that ordinary TTS documentation does not cover:
“sometimes I still get this audio ‘gibberish’ … the cloned voice just says something random in between the actual dialogue lines.” — u/nemesew, Reddit
In a separate H3 discussion, users reported that reference mode preserved the voice but removed background music and footsteps, while another mode kept more environmental audio but matched the voice less closely. If H3 repeats the sample’s original words, shorten the reference, remove distinctive spoken instructions, and follow the official reference-audio syntax. This is a workflow trade-off, not evidence that standard Speech TTS is broken.
Cost, retention, and consent: the operational gate
The exact bill depends on the route, and a clone fee is separate from later speech generation unless the provider says otherwise.
| Item | Documented value | What to verify before launch |
|---|---|---|
| Replicate clone operation | $3 per output | Separate Speech 02 generation pricing and current terms |
| fal clone request | $1.50 | Failed-request billing, TTS rate, and current retention policy |
| fal preview text | $0.30 per 1,000 characters | Whether preview is necessary for your workflow |
| fal voice retention | Use with TTS within 7 days for permanent retention | What “permanent” means under the current service terms |
Voice cloning should be limited to voices you are authorized to use. MiniMax’s public interface includes a rights/consent confirmation, but that UI check is not a substitute for a release, contract, or local legal review when the voice belongs to another person. Do not use a clone to impersonate someone or to mislead listeners about who is speaking.
MiniMax voice clone FAQ
How long should the sample be?
Use at least 10 seconds for API routes; a clean 20–40-second clip is a practical starting point, while some services accept up to 5 minutes.
Is MiniMax voice cloning the same as voice conversion?
No. TTS cloning generates new speech from text in a learned voice. Voice conversion aims to preserve the original performance while changing the speaker identity; the sources reviewed here do not establish that MiniMax’s standard clone endpoint provides that full workflow.
Can I reuse the cloned voice through an API?
Yes, when the selected provider returns a reusable voice identifier. Store the ID with its provider and model because a Replicate or fal identifier is not automatically portable to the direct MiniMax API.
Why does the clone repeat the source audio?
This is mainly reported in H3 reference-audio workflows, where the model can treat reference speech as content. Use a clean, short reference and test with new text before using the result in a long video.
Can I clone someone else’s voice?
Only with authorization and in compliance with applicable privacy, publicity, copyright, impersonation, and platform rules. A technically successful clone is not proof that its use is permitted.
Make the decision by workflow, not by the demo
Choose MiniMax Audio if you want the least technical browser workflow. Choose fal when you want a documented queue-based API, a $1.50 clone operation, and a custom_voice_id; plan around its seven-day retention rule. Choose Replicate when its model-page privacy and API conventions fit your stack, but resolve the 5-second README versus 10-second schema discrepancy by following the schema. Choose a direct MiniMax integration when you control the upload, voice registry, billing, and TTS pipeline yourself.
For H3 video, judge the output on two axes: voice identity and the rest of the scene’s audio. If reference mode fixes the voice but removes Foley or introduces gibberish, keep the workflow for character-voice experiments rather than treating it as a drop-in voice-conversion system.
Related MiniMax coverage: MiniMax H3, MiniMax H3 prompt guide, MiniMax H3 Max API pricing, and MiniMax H3 vs LTX 2.3.