NkomoNkomo
← All posts

English–Twi Speech Recognition: What Akua’s Twi Correction Reveals About Voice Control

A dedicated English–Twi speech-recognition model can make mixed-language speech more likely to be transcribed as one natural utterance, rather than treated as a mistake at the point of the switch. Product teams still need to evaluate real conversations, voice control, meaning, error handling, and privacy before describing the experience as natural.

At 6:40 p.m. in Kumasi, Akua is standing outside a shop with a crumpled list in one hand and her phone in the other. Her mother has asked her to send a voice message to a cousin: “Tell him the documents no, ɔmfa nka ho. The appointment is tomorrow.” Akua tries twice. Each time, the English arrives, then the Twi phrase becomes nonsense on screen.

The risk is small in length and large in consequence. If the message changes meaning, her cousin could arrive with the wrong papers, and the appointment could be lost. She cannot stop to translate herself into one language before she speaks. This is how she talks.

A speech-recognition model built for English–Twi switching changes the starting point. It gives the system a better chance to hear the language mix as speech people actually use. That matters. It does not complete the work.

Recognition must survive the switch itself

A general speech system may perform well when someone speaks English all the way through, or Twi all the way through. The difficult moment often sits between them: a name, an instruction, a question, a correction, or an English word that belongs naturally inside a Twi sentence.

Teams should test the moments people rarely rehearse. Ask for a short request spoken quickly. Ask for a message that includes a family name, a place, a number, and a correction halfway through. Include Ghanaian English phrasing rather than treating British or American English as the default.

The transcript needs to preserve intent, not merely produce words that look plausible. A wrong word after a switch can reverse an instruction. A missed negative can turn “don’t add it” into “add it.” As explored in What Happens When an AI Rejects an English Word in a Twi Question?, the failure can begin when the system decides that one part of a mixed sentence does not belong.

Evaluation should therefore include examples from the conversations the product expects to support. Product teams can build a consented test set, label the intended meaning, and review errors by type: words lost at a language switch, names altered, short Twi responses misunderstood, or English terms forced into the wrong spelling.

A good transcript can still create an awkward conversation

Speech recognition is only the first turn. A voice experience feels natural when the person can speak, pause, change their mind, and hear a response without fighting the interface.

Akua does not speak in neat blocks. She starts again: “Tell him the documents no…” Then she catches herself. “No, wait, tell him sɛ ɔmfa nka ho.” A useful voice experience should let her interrupt, correct the message, and continue. If she has to wait through a long reply or begin from scratch, the conversation becomes work.

That is why teams should test the whole interaction, including:

  • How reliably the system knows when someone has finished speaking.
  • Whether a person can interrupt a reply naturally.
  • Whether spoken responses stay concise enough for the situation.
  • Whether the product makes uncertainty visible instead of quietly returning a confident-looking error.

Nkomo supports typed and spoken conversation in Twi and Ghanaian English, including hands-free voice mode where a person holds to speak, can interrupt naturally, and hears the reply. Those interaction details deserve the same attention as transcription accuracy. A model can recognise the sentence correctly and still leave the speaker feeling unheard if turn-taking fails.

Meaning needs human review, not a single accuracy score

A single score can hide the errors that matter most. If a model performs well overall but repeatedly misses a switch into Twi during instructions, reminders, health questions, or money-related language, the average offers false comfort.

Use scenario-based reviews alongside technical measures. Give bilingual reviewers the original audio, the transcript, and the system response. Ask a direct question: could a listener act correctly on this? Then ask whether the response sounds appropriate to the language mix, rather than formal, flattened, or strangely translated.

The point is not to demand that every speaker use identical Twi or identical English. Natural conversation includes variation. The product needs clear boundaries around what it has evaluated, where it remains uncertain, and what it does when it cannot understand.

That last part builds trust. Silence makes people guess whether their words were heard. Clear errors give them a chance to repeat, type, or try again. Nkomo is designed to show errors rather than swallow them.

Privacy belongs in the voice experience

Voice can feel personal before it feels technical. Akua is sending a family message, and she should be able to understand what happens to the audio and conversation history without digging through vague settings.

Teams need to test the privacy flow with the same care they give the microphone flow. Can people choose whether cloud processing is allowed? Do they understand the choice at the moment it matters? Can they turn off on-device history, and does it actually disappear when they do?

Nkomo gives people explicit cloud-consent choices: never, ask each time, or this session. Its on-device history can be turned off and purged immediately, and people can export their data or delete their account in one tap.

Later that evening, Akua checks the message before sending it. The Twi correction remains where she put it. She sends it, then turns off history because the message is private. That is the standard worth testing toward: the person speaks naturally, sees what happened, and stays in control of what remains.

Nkomo

A private, natural Twi and Ghanaian English voice-and-text companion — talk or type, in the mix of languages people actually speak, with clear control over what stays on the device versus what reaches the cloud.

Try Nkomo

Comments

No comments yet.