Dia 1.6B
Open-source TTS for ultra-realistic two-speaker dialogue.
Upvotes
2
Pricing
Open Source
Category
Text-to-Speech
In the database since
Oct 2026
Overview
About Dia 1.6B
What Is Dia 1.6B?
Dia 1.6B is an open-source text-to-speech model created by Nari Labs that generates ultra-realistic dialogue in a single pass. Instead of voicing one line at a time, you hand it a full transcript with speaker tags, and it renders a complete two-person conversation — pacing, interruptions, and all — as one continuous audio file.
What sets Dia apart from conventional TTS is expressiveness. The model produces non-verbal sounds like laughter, coughing, and sighing, and it can be conditioned on an audio sample to control emotion, tone, and voice identity. It is released under the Apache 2.0 license with public checkpoints and inference code, and it has gathered a large open-source community since its release in April 2025.
Dia is a research-grade tool rather than a consumer app: you run it locally on a GPU, through a Gradio web interface, a command-line script, or the Hugging Face Transformers library. The model currently supports English generation only.
Key Features
1. Single-pass dialogue generation
Generates an entire two-speaker conversation from one transcript in a single model pass, keeping timing and delivery coherent across the whole exchange.
2. Speaker tags
Direct who-says-what with [S1] and [S2] tags, alternating between speakers for natural back-and-forth dialogue.
3. Non-verbal sounds
Inserts lifelike non-verbals such as (laughs), (coughs), (sighs), (gasps), and (chuckle) directly into the spoken output.
4. Voice cloning from a short sample
Clones a voice from 5–10 seconds of audio plus its transcript, then speaks your script in that voice.
5. Emotion and tone control
Condition generation on a reference audio clip to steer the emotional delivery and speaking style.
6. Transformers and Hugging Face integration
Runs through the Hugging Face Transformers library via DiaForConditionalGeneration, with weights hosted on the Hub.
7. Gradio UI and command-line interface
Generate from a browser-based Gradio app or automate batch jobs with the included CLI script.
Frequently Asked Questions
How do I write a script for two speakers in Dia?
Start your text with [S1] and alternate between [S1] and [S2] for each speaker's lines — the model voices each tag as a distinct speaker in the conversation.
Can Dia clone my voice?
Yes. Provide 5–10 seconds of reference audio with its transcript in the same [S1]/[S2] format, and Dia will speak your script in that voice; cloning real people without their permission is strictly forbidden by the project's usage terms.
What non-verbal sounds can Dia produce?
Dia recognizes tags like (laughs), (clears throat), (sighs), (gasps), (coughs), (sings), (mumbles), (groans), (claps), (sneezes), (chuckle), and (whistles) — use them sparingly, as overuse can cause audio artifacts.
What hardware do I need to run Dia 1.6B?
A CUDA-capable GPU is required — CPU support is not available yet; on an RTX 4090 the model needs about 4.4 GB of VRAM in half precision and generates roughly twice as fast as real time.
Does Dia support languages other than English?
No. The model only supports English generation at the moment, though the project is open source and community fine-tunes may extend this.
Why does the voice sound different every time I run it?
Dia was not fine-tuned on any specific voice, so each run samples a new one; pin the random seed or supply an audio prompt to keep the voice consistent.
Conclusion
Dia 1.6B is the open-source reference for expressive dialogue synthesis: realistic two-speaker conversations, non-verbal sounds, and voice cloning, all from a freely available 1.6-billion-parameter model. The trade-offs are honest ones — English only, GPU required, and voices that vary unless you anchor them. For developers, podcasters, and researchers who want full control over AI dialogue without a commercial API, it is one of the most capable open options available.
Keep exploring
Similar AI tools
Ranked by the job they do, their capabilities, and how closely they fit this workflow.