Boson AI
Open audio foundation models — the Higgs family — for speech synthesis, recognition and audio understanding.
Highlights
- Higgs Audio v2: expressive text-to-speech foundation model pretrained on 10M+ hours of audio
- Higgs Audio 3.0: speech-to-text (ASR) model supporting 94 languages with language detection
- Higgs Audio 2.5: production-focused voice generation with reduced latency
- Audio understanding with real-time reasoning over sentiment and semantics
- Open weights published on Hugging Face and code on GitHub
- Multi-speaker and emotion-controllable speech synthesis
- Published the EmergentTTS-Eval benchmark (NeurIPS 2025) for expressive TTS
- Aligned foundation models and custom solutions for enterprise via direct contact
External link — opens boson.ai in a new tab. Boson AI is a third-party product; we are not affiliated with it.
About Boson AI
What it is
Boson AI builds the Higgs Audio family: an expressive text-to-speech foundation model pretrained on over ten million hours of audio, a speech-to-text model covering 94 languages with automatic language detection, and production-focused voice generation. It was founded in 2023 by Alex Smola and Mu Li, both former AWS AI leaders and co-authors of the textbook Dive into Deep Learning.
Why it's different
The distinguishing choice is treating audio as a foundation-model problem rather than as separate synthesis and recognition products. Models pretrained across that much audio carry prosody and expressiveness that older pipeline systems reconstruct awkwardly, and openness means teams can run them on their own infrastructure — which matters when the audio is customer calls or medical dictation that cannot leave your estate. The trade against ElevenLabs is the usual one: the commercial incumbent has better tooling, better voice cloning and a far easier onboarding, and the open models take engineering effort to deploy well.
How people use it
It is used by developers building speech into products who want control over deployment and cost rather than per-character pricing — voice interfaces, transcription pipelines, dubbing at volume. The 94-language recognition model is the standout for anyone whose users are not English-speaking, where the mainstream options degrade sharply. Self-hosting is the main reason to choose it, so the realistic comparison is against Whisper, not against a hosted API.
Written by the n3os team. We are not affiliated with Boson AI.
This listing was written from public information, without Boson AI’s involvement. If you own it and something here is wrong — or you would rather not be listed at all — email us and we will correct or remove it.
Get the ones worth knowing about
We write one of these for every tool worth the trouble. Get the new ones, plus what we have found genuinely useful lately.
Your address goes to Buttondown, who send the emails on our behalf. One click unsubscribes, and the list is never sold or shared.