Nkomo
← All posts

What Happens When Speech Recognition Meets a Twi-English Switch?

5 min read · Published August 31, 2026
Audience recording a live performance with a smartphone, showcasing modern technology use.

Photo by DANIEL INTUS on Pexels

Recognition quality can change inside one breath because speech systems may handle English and Twi with different levels of confidence, then struggle at the point where one language becomes the other. A clean English transcript can therefore turn uncertain the moment a speaker code-switches, even when the full sentence sounds completely natural to a human listener.

In 1999, NASA’s Mars Climate Orbiter approached Mars after a journey of hundreds of millions of kilometres. One part of the navigation process used pound-force seconds. Another expected newton-seconds. Each side could produce valid-looking numbers, but the handoff between them carried the wrong assumption. The spacecraft entered the planet’s atmosphere far lower than planned and was lost.

NASA’s Mars Climate Orbiter Mishap Investigation Board documented the failure. The lesson reaches beyond spacecraft: a system can perform smoothly within one frame, then fail where two frames meet.

That boundary is exactly where a Twi-English utterance can become difficult for speech recognition.

One sentence can contain several recognition problems

Picture someone recording a message in Accra. The first words are in Ghanaian English, and the transcript appears correctly. Halfway through the same breath, the speaker moves into Twi because that is the clearest, warmest or most natural way to finish the thought.

The transcript begins confidently. Then the first Twi phrase becomes an English word that sounds vaguely similar. The next phrase loses its spacing. A familiar name gets pulled into the confusion. By the end, the written version no longer carries what the speaker said.

The speaker did not make an error. They used language as they ordinarily use it.

A speech system has several jobs happening at once. It must identify sounds, separate words, account for accent and background noise, and decide which language patterns best explain what it heard. During intra-sentence switching, those decisions can change within seconds.

A 2025 study of spontaneous intra-sentence language switching examined this problem from model- and data-centric perspectives. It identified both language switching and accent bias as sources of confusion in automatic speech recognition. That matters because adding a language label to a product does not prove that mixed-language speech will be handled well. The difficult part may live at the boundary.

Why the English can look better than the Twi

Recognition systems learn from data. When their training and evaluation material contains far more English, more accent varieties represented as English, or too few examples of natural Twi-English switching, the system may favour the patterns it knows best.

That preference can surface as guessing. A Twi word is forced into a plausible English spelling. A switch is treated as noise. The system stays committed to English for too long because the opening words established that context.

This resembles the Mars Climate Orbiter failure at the level that matters: the trouble appears during interpretation across a boundary. Valid signals enter the system, but an assumption at the handoff changes what they mean.

Accent adds another layer. Ghanaian English has its own rhythms and sound patterns. Twi pronunciation can also vary by speaker and context. Recognition quality may shift with pace, microphone quality, surrounding noise, and how quickly the speaker moves between languages. A polished result from one sentence cannot guarantee the next one.

For a closer look at what a language switch changes inside a family message, read What Does a Twi to English Switch Change in a Family Voice Note?.

What better handling looks like in practice

The useful test is ordinary speech. Ask someone to say what they would actually send to a sibling, parent or colleague, without slowing down for the machine. Keep the natural pauses, names, restarts and changes of language.

Then inspect the boundary:

  • Does the first Twi phrase survive after an English opening?
  • Are names preserved when the surrounding language changes?
  • Does the transcript recover after one uncertain word?
  • Does the product show an error when it cannot complete the request?

That last point matters. A wrong transcript presented with confidence can travel into a reply, reminder or family message. Nkomo does not fail silently. When an error occurs, it is shown rather than swallowed.

Nkomo supports typed and spoken conversation in natural Twi and Ghanaian English. In voice mode, you can hold to speak, interrupt naturally and hear the reply. The goal is to let people speak in the mix they already use, while staying exact about what the product can do today. Recognition still deserves testing against real voices and real code-switching, sentence by sentence.

Privacy deserves the same attention. Before sharing personal speech, Nkomo lets you choose whether cloud access is never allowed, requested each time or allowed for the session. History stays on the device under your control, and turning history off purges it immediately. The distinction between language handling and data handling is explored further in AI Voice Note Privacy: How Efua Understood Twi and English, Then Purged the Chat.

Test the handoff, not only the languages

A language list tells you which languages a product claims to support. It says little about what happens when both arrive in one breath.

Try the sentence you would actually say. Begin in English, move into Twi, include the family name or local expression that matters, then check the transcript before relying on the reply. Repeat the test with another speaker and a different switching point.

NASA’s 1999 investigation showed how costly an unchecked boundary assumption can become. The stakes here are smaller, but the practical discipline is the same: inspect the handoff. That is where a smooth transcript can start guessing, and where honest evaluation should begin.

Nkomo

A private, natural Twi and Ghanaian English voice-and-text companion — talk or type, in the mix of languages people actually speak, with clear control over what stays on the device versus what reaches the cloud.

Try Nkomo

Comments

No comments yet.