Poe API
Gemini-2.5-Flash-TTS
Gemini‑2.5‑Flash‑TTS is Google’s low‐latency text‑to‑speech model that converts text input into audio output, supporting both single‑ and multi‑speaker voices with controllable style, accent, and expressive tone — ideal for applications like podcasts, audiobooks, and conversational voice systems.
Notes:
- Text and style prompt limited to 4,000 bytes each (8,000 bytes combined)
- Max output duration: approximately 10 minutes
- Multi-speaker requires SpeakerName: text format (example: Alice: Hi! Bob: Hello, must be on new lines)
- The model auto-detects the input language. The Language setting is a hint to help choose the right voice/accent, the model may override it if the text is in a different language.
This bot supports optional parameters for additional customization.
Powered by a server managed by @empiriolabsai. Learn more
Build with Gemini-2.5-Flash-TTS using the Poe API
Start by creating an API key, for use with any bot on Poe:
See the full documentation for comprehensive guidance on getting started.
More from EmpirioLabs AI
New
Qwen3.8-Flash-EL
New
GLM-5.3-Flash-EL
New
Wan-3.0
GLM-5.3-EL
DeepSeek-V4-Pro-0813-EL
Qwen3.8-27B-EL
Grok-Imagine-Image-2
Muse-Glimmer-30B-EL
Qwen-Image-3.0-EL
Seedance-2.5-EL