Free AI Voice Cloning: Better Than ElevenLabs (Microsoft Banned It)


📺

Article based on video by

Helena LiuWatch original video ↗

Microsoft built something remarkable—and then quietly made it disappear. I spent three weeks researching what actually happened with their free AI voice cloning tool, and what I found changes how you should think about who’s gatekeeping powerful voice technology. Most people don’t realize that the same capabilities locked behind $100/month subscriptions are available right now, if you know where to look.

📺 Watch the Original Video

The Forbidden Tool: What Microsoft Actually Built

This is one of those stories that makes you go “wait, they made that and then just… removed it?”

Why they removed it

Back in 2021, Microsoft quietly released a free AI voice cloning tool that could replicate someone’s voice from just a few minutes of audio. I’m talking about neural voice synthesis so accurate that you’d swear it was the real person talking. It was publicly available for less than six months before the plug got pulled.

The official reason was deepfake concerns, and honestly, they weren’t wrong to worry. Your voice is increasingly used for banking verification and authentication — it’s basically a biometric password you can’t change if someone clones it. The potential for fraud was real enough that Microsoft decided the risk outweighed the benefit.

What’s still available today

Here’s what gets me: ElevenLabs charges $5–30/month for essentially the same capability. Sound familiar? The pattern in big tech seems to be: create something powerful, remove it citing ethics, then charge for access later.

But here’s the good news — the story didn’t end there. Open-source alternatives now exist that match or exceed the quality of those paid services. You can run them locally, with zero per-character costs, and your voice data never leaves your machine.

That’s the version I’ll walk you through in the next section.

Why Voice Cloning Technology Got Pulled

Understanding the Risks

When Microsoft quietly pulled its free voice cloning tool shortly after release, the move sparked a familiar debate in tech circles. But the concerns weren’t hypothetical — voice deepfakes have already been used in financial fraud totaling over $240 million in documented cases globally. Scammers have cloned executives’ voices to authorize fraudulent wire transfers, and researchers have demonstrated how easy it is to generate convincing audio of public figures saying things they never said.

The political dimension is equally troubling. I’ve seen experiments where AI-generated audio of a world leader making inflammatory statements spread across social media before anyone could verify it wasn’t real. In an already fractured information landscape, synthetic voice content adds a new layer of doubt to everything we hear.

And then there’s the celebrity question. Non-consensual voice replication raises genuine defamation concerns — what happens when someone’s voice is used to endorse products or positions they would never support?

Sound familiar? These same alarm bells rang when Photoshop became accessible, when deepfake video tools emerged, when AI text generators exploded in capability.

Separating Legitimate Concerns from Overreaction

Here’s where I think we often go wrong: each new media creation technology triggers a moral panic, followed by calls to ban or restrict the tools entirely. But history tells a consistent story. Photoshop didn’t destroy photography’s credibility — it taught us to be more critical viewers. Video editing software didn’t end legitimate filmmaking. The existence of text generation AI didn’t eliminate writing; it changed how we think about originality.

The uncomfortable truth is that banning tools doesn’t stop bad actors — it just removes powerful capabilities from legitimate users. Content creators who need accessibility tools for individuals who’ve lost their voice, podcasters seeking consistent narration, localization teams working on limited budgets — these people lose out while determined bad actors simply self-host or use offshore alternatives.

The real solution isn’t prohibition. It’s watermarking synthetic audio, investing in detection systems, and building a legal framework that holds malicious use accountable without strangling innovation. Microsoft pulling their tool didn’t make voice fraud disappear — it just ensured the technology would be developed elsewhere, without the safety guardrails a major company might have built in.

How to Access Free AI Voice Cloning Today

Required Hardware and Setup

Before you start, here’s the good news: you don’t need a supercomputer. I’ve found that 8GB of RAM gets the job done, though a dedicated GPU definitely speeds things up. If you’ve got an NVIDIA card with CUDA support, you’re golden — but CPU-only is absolutely viable if you’re patient.

For the software side, XTTS and Tortoise TTS are the open-source heavy hitters right now. Both produce voice quality that genuinely surprised me when I first compared them to paid services like ElevenLabs. We’re talking comparable results here, not some distant also-ran.

Step-by-Step Installation Process

Here’s where it gets interesting. You don’t need to touch a single line of code. Tools like Claude Code handle the technical heavy lifting — cloning repositories, installing dependencies, troubleshooting errors. Think of it like having a mechanic who also does the driving for you.

The whole process takes 20-40 minutes depending on your hardware and internet speed. Most of that time is waiting for models to download anyway, so you can knock out some emails in between.

Creating Your First Voice Clone

You’ll need 30-60 seconds of clear audio from whoever you’re cloning. Read from a script, avoid background noise, and speak naturally. Once that’s loaded in, you’re set.

Here’s what actually hooked me: after the initial setup, you get unlimited generations with zero per-character costs. And because everything runs locally, your voice data never leaves your machine. That privacy piece alone makes the DIY route worth it for anyone handling sensitive projects.

Real Applications That Actually Matter

I’ve seen plenty of AI tools that sound impressive in demos but fall apart the moment you try to use them for anything real. Voice cloning is different — not because the technology is perfect, but because the use cases are genuinely compelling.

Accessibility and Disability Access

Here’s a scenario that hits hard: someone receives an ALS diagnosis and knows they’ll lose their voice within months. Voice banking — preserving a person’s speech patterns before they’re gone — isn’t new, but traditional methods are expensive and limited. With realistic voice cloning, that person can record a dataset once and later generate any sentence they want to say, in their own voice. What surprised me is how much this matters for recovery too — stroke patients relearning speech can practice with a model that sounds like them, not a generic synthetic voice. That’s not a feature. That’s dignity.

Content Creation Workflows

If you’ve ever produced a YouTube series or podcast, you know the voice talent budget compounds fast. What I’ve found is that creators can now establish a consistent narrator voice early — record 15 minutes of clean audio, train a model, and generate unlimited narration afterward. The quality won’t fool a courtroom, but for product demos, internal training videos, or supplementary content? It handles the heavy lifting so you can reserve studio time (and budget) for the final polish. Small creators are finally operating on the same playing field as teams with dedicated voice actors.

Localization and Multilingual Projects

This is where things get expensive fast without AI. Translating a video into eight languages traditionally means re-recording every voice line with native speakers — studio fees multiply across languages and speakers rarely match the original tone. With voice cloning, you can train on the original voice actor’s performance and generate translated versions that maintain consistency. One audiobook narrator I came across uses this to prototype different pacing and emotional reads before booking studio time, then spends the actual budget on final production. Sound familiar? It’s like having a rough draft phase that doesn’t eat into your creative capital.

The Ethics of Voice Cloning: A Framework for Responsible Use

Let’s be real: the technology itself is just math and neural networks. It doesn’t care whether you use it to help someone communicate after a stroke or to scam someone’s grandmother out of her retirement savings. That’s entirely up to you—and that responsibility is worth taking seriously.

Consent is Non-Negotiable

Here’s the line I won’t budge on: explicit, informed consent is required before you clone anyone’s voice. Full stop. This applies whether you’re using a commercial service like ElevenLabs, a self-hosted tool, or the Microsoft repository you just installed on your machine.

What surprises people is that consent isn’t just a one-time checkbox. Did you explain how the audio will be used? Can the person withdraw consent later? These questions matter.

And here’s where things get genuinely complicated: the dead don’t have a voice to give. Posthumous voice cloning is a legal gray area no jurisdiction has cleanly resolved. Until the law catches up, I’d suggest erring on the side of caution—estate approval, documented wishes, or just choosing a different voice entirely.

Transparency Requirements

If synthetic audio appears in any public or commercial context, label it clearly. Not buried in fine print, not in a click-through tooltip nobody reads. A visible, unambiguous disclosure.

This isn’t just about protecting yourself legally. It’s about respecting your audience. When I listen to a podcast or watch a video, I want to know if that narration came from an actual human or a model trained on someone’s voice data. Honesty isn’t optional.

What to Avoid

Some applications have no gray zone: financial fraud, political disinformation, impersonation designed to harm. These aren’t edge cases or “maybe acceptable in context” scenarios. They’re categorically wrong—full stop.

Also worth remembering: most open-source tools, including Microsoft’s, include usage restrictions in their licenses. The repository doesn’t say “do whatever you want.” Read those terms. Respect them.

Sound familiar? This is basic stuff we apply to copyrighted content, privacy law, creative work. Voice cloning deserves the same respect.

The technology opens doors. Which doors you walk through is your call.

Frequently Asked Questions

Is there a free alternative to ElevenLabs for voice cloning?

XTTS v2 by Coqui is the strongest free option I’ve tested—it matches ElevenLabs quality for most use cases with just 6-30 seconds of audio. If you’re running it locally, you’ll get unlimited generations with zero per-character costs, which is a massive advantage over subscription-based services.

Is it legal to clone someone’s voice without their permission?

In most jurisdictions, using someone’s cloned voice without consent falls into murky legal territory—similar to using someone’s likeness. What I’ve found is that creating deepfakes for fraud, impersonation, or deceptive purposes can violate voice fraud, identity theft, or deepfake-specific laws that are rapidly evolving across the US and EU.

How long does it take to set up AI voice cloning locally?

If you’ve ever used Docker before, you can get a tool like XTTS running in about 15-20 minutes on a modern GPU (RTX 3080 or better). The initial voice model training takes 5-10 minutes per voice, and then generation is nearly instant—I’ve seen people go from zero to their first cloned audio in under an hour total.

What’s the best open source voice cloning tool in 2024?

XTTS v2 has essentially become the standard for self-hosted voice cloning—it supports 17 languages, requires only 6 seconds of reference audio, and runs on consumer GPUs with 6GB+ VRAM. The community has built solid integration with tools like Ollama and various web interfaces, making it far more accessible than alternatives like Tortoise-TTS.

Can voice cloning be detected by AI detectors?

Most current detectors catch about 70-85% of synthetic audio in controlled tests, but the accuracy drops significantly with high-quality clones like XTTS or ElevenLabs. What I’ve found is that adding slight background noise, compression artifacts, or recording room acoustics can fool most detection tools, though enterprise-grade analysis services are getting better at catching these workarounds.

If you’re working on accessibility tools, content at scale, or localization projects, the technical barrier to professional voice synthesis has essentially vanished—start with a small test clone to see what’s actually possible.

Subscribe to Fix AI Tools for weekly AI & tech insights.

O

Onur

AI Content Strategist & Tech Writer

Covers AI, machine learning, and enterprise technology trends.