Hindi text-to-speech comparisons often use clean Devanagari prose. Real messages are less tidy. Product names, addresses, acronyms and whole clauses arrive in English; Hindi may be typed in Roman letters; one sentence may contain both scripts. That is why a demo winner may fail in a Hinglish workflow.
Google lists its Hindi models in the Cloud TTS catalogue, while Microsoft documents Swara, Madhur and other choices in its speech catalogue. Nkomo exposes released voices instead of treating one female voice as the language itself.
Three tests, not one
Evaluate formal Devanagari Hindi, conversational mixed-script Hinglish and Romanized Hindi separately. Include Indian names from several regions, rupee amounts, phone numbers, URLs and common technology terms. Preserve the original spelling during comparison; “fixing” inputs differently for each provider makes the result meaningless.
What online opinion gets right
Listener reports commonly describe Google’s neural Hindi as clear and steady and Microsoft’s Swara as warm and expressive. Those descriptions are useful hypotheses, not a benchmark. Speaking rate, sentence type and code-switch density can change the preference.
Nkomo’s current route
Google Neural2 remains the first route for Hindi while the evidence is mixed. If it fails, Nkomo tries the selected Microsoft voice through Edge and then Azure. The two Microsoft transports preserve the selected voice; Azure provides the supported fallback when the Edge endpoint is unavailable.
The next routing change should come from blind native review across all three writing patterns, not from a single literary paragraph.
Comments
No comments yet.