VoxLoom bundles OmniVoice, Qwen3-TTS and Chatterbox into one local CLI + Web UI. Generate and clone voices on hardware you control — no subscriptions, no rate limits, no vendor lock-in.
The open-source alternative to ElevenLabs that you actually own.
We got rate-limited by a commercial TTS API one time too many. So we aggregated the best open-source speech models into one studio you run yourself.
MIT / Apache-2.0. Audit it, fork it, own it. Your voice data never leaves your machine.
Your GPU, your rules. Not someone else's API tier or monthly quota.
OmniVoice covers 646 languages; Qwen3-TTS gives studio-grade Chinese cloning.
3-second reference audio, or text-prompt a voice — no training rig required.
pip install voxloom and point it at your GPU. No account, no key.
Need more horsepower or a premium voice? Our hosted API is there when you do.
One CLI, three best-in-class open models. Pick per job.
| Engine | Languages | Cloning | License | Min VRAM |
|---|---|---|---|---|
| OmniVoice | 646 | 3s ref / text-prompt | Apache-2.0 | 6 GB |
| Qwen3-TTS | 10+ | yes | Open | 8 GB |
| Chatterbox | 23 | few-sec ref | MIT | 8 GB |
Quick start: pip install voxloom && voxloom serve --engine omni-voice --device cuda
Start free and self-hosted. Pay only when you need cloud horsepower, a premium voice, or a commercial SLA.
Be first to get premium voices and hosted GPU. No spam — one launch email.
Opens your email app and sends straight to our inbox. Most reliable — no middleman.
Note: the form routes through a 3rd-party relay and may be delayed or filtered.
Yes — the repo is MIT and every bundled engine is open (OmniVoice Apache-2.0, Chatterbox MIT). Run it forever at no cost.
OmniVoice runs on 6 GB VRAM; Qwen3-TTS / Chatterbox on 8 GB. CPU works but slower. Cloud API covers the rest.
Self-hosted: yes under each engine's license. Cloud API plans include a commercial license.
You own the stack, your audio stays local, and there are no rate limits or subscriptions. Cloud is optional.