8 comments

[ 0.81 ms ] story [ 6.7 ms ] thread
No weights, not very interesting.
> Word-level timestamps. Each word has precise start/end times and confidence scores.

Neat. This is missing from many models.

(comment deleted)
I don't understand how these benchmarks get things so wrong. Nova-3 is by far the best model on the market in WER. Azure's STT is atrocious. And yet every benchmark rates Azure's STT higher than Nova-3's.
Mai 2 transcribe is giving me better results than nova
$0.10/hour is a good price vs $0.30/hour for Gemini Transcribe 3.5.