1 comment

[ 2.7 ms ] story [ 4.5 ms ] thread
Hey HN, Toby from Nari Labs here.

We're launching our realtime TTS endpoints for Qwen3-TTS 1.7B today. > "Fast" endpoint: client-side p90 time-to-first-audio (TTFA) at 50ms. > "Standard" endpoint: dirt cheap at $5 per 1M characters, still faster than industry avg at 200 ms TTFA.

For comparison, 11Labs V3 is $100 for 1M and Cartesia Sonic 3.6 is $50.

TTFA is measured using our OSS benchmarking tool from us-west-1 and us-east-1 (https://github.com/nari-labs/benchmarks.git).

Quality: link to TTS samples (https://narilabs.com/blog/introducing-nari-qwen3-tts/). Based on preliminary benchmarks by a 3rd-party (Speko AI), we scored above leading TTS providers such as Grok and Rime (https://narilabs.com/blog/qwen3-tts-speko-benchmark/)!

We are working hard on further lowering the price of our endpoints, as well as adding more voices and support for input streaming.

We believe open-source will win - not just in LLMs but also in multimodal :) During public beta, all endpoints are free. Try it out and let me know how it goes!