
Unlimited text-to-speech and voice cloning, flat-priced per stream with WebSocket and HTTP APIs
Gandr sells unlimited text-to-speech for everything with a voice, and the pricing page headline explains the model in five words: flat per stream. No per-character billing, no per-minute meter, no counting voices or requests. You pay for concurrency and generate as much audio as your streams can carry, which is a structure audio-heavy applications will recognize as either liberation or a trap depending on their traffic shape.
The integration list on the homepage names names: LiveKit, Pipecat, Vapi, Retell, Daily, and Whisper. That is the voice-agent ecosystem in one row, and dropping into what you already run is the explicit pitch. Agents connect over WebSocket, narration and dubbing over HTTP, and the console demo, type a line, hear it back, is the fastest possible proof the pipeline works.
The toolify listing adds the cloning specifics: about ten seconds of reference audio clones a voice instantly, 23 languages are supported, and every output carries a watermark. That last detail matters more than it seems, since it is the difference between a cloning platform that takes abuse seriously and one waiting for a takedown notice.
Who is this for? Voice agent builders, narration and dubbing pipelines, and anyone whose audio bill scales painfully with usage. Contact is public at contact@gandr.ai. The evaluation is simple: price your current per-character spend against a flat stream and see which one lets you sleep.
What it is
Unlimited text-to-speech and voice cloning, flat-priced per stream with WebSocket and HTTP APIs
Pricing
Paid
Primary category
Audio
Source code
Closed source / hosted
Gandr is one of many tools in its category. Before committing to it, weigh a few practical points that tend to decide whether a tool actually fits your work:
We describe Gandr and its alternatives honestly so you can compare on substance. Browse the similar tools below to see how it stacks up against other options in the same category.