NkomoNkomo
← All posts

Voice AI misunderstandings: What Adwoa’s medicine request taught about repair

a woman holding a cell phone in front of a pink background

Photo by Ahmed Nasiru on Unsplash

A voice AI earns trust by making misunderstandings visible, clear, and easy to correct. When it guesses wrong, it should say so plainly instead of carrying on as though it heard you.

At 6:42 p.m., Adwoa stood beside a trotro stop in Kumasi with her shopping bag pressed against her knee and her phone close to her mouth. Her younger brother had sent a voice note about their auntie’s medicine, mixing Twi and English as naturally as he spoke at home. Adwoa tried a voice assistant for a quick summary before she called the pharmacy.

The reply came back confident and wrong. It had turned a question about collecting medicine into a request about booking a pickup.

She tried again, slower. Same answer, with different wording.

“Medaase, I’ll explain it myself,” she said, ending the call.

That sentence sounds polite. It carries frustration. She had already spent time repeating herself, and the wrong answer had made the next step less clear. The pharmacy could close before she reached it. Her auntie might miss what she needed that evening. The bot had not merely failed to help. It had hidden the fact that it was confused.

A wrong answer needs to look wrong

Voice AI will sometimes misunderstand people. Accents vary. Speech gets interrupted by traffic, children, poor signal, or another person calling from across the room. Twi and Ghanaian English often sit in the same sentence. A system that pretends otherwise asks the speaker to carry all the cost of its mistake.

Visible misunderstanding is the minimum respect a voice AI owes the person speaking.

That means showing an error when something fails instead of swallowing it. It means giving the speaker a chance to interrupt, rephrase, or type the detail that was missed. It means avoiding the polished but risky performance of an answer that sounds certain while heading in the wrong direction.

For Adwoa, the useful response would not need to be dramatic. It could say it did not catch the medicine name, then invite her to repeat that part. She could correct one detail rather than start the whole story again.

A confident wrong answer creates a particular kind of pressure. People often assume they spoke badly, mixed languages too much, or used the wrong word. The burden shifts onto them. That is a poor bargain when the task involves health, money, family plans, or a message someone needs to send before the evening changes shape.

Code-switching should not become a test you have to pass

People do not pause everyday speech to choose a language setting. A family conversation can move from Twi to English and back within a few seconds. The point is understood by the people in the conversation because they share context, tone, and habit.

A voice tool should meet that reality with humility.

Nkomo is built for natural Twi and Ghanaian English conversation by voice or text, including the mix of languages people actually speak. You can hold to speak, hear a reply, and interrupt naturally when the answer has gone off course. If speech is unclear or something goes wrong, the product shows the error rather than acting as if everything worked.

That design choice matters because correction is part of conversation. A helpful companion leaves room for, “No, I meant the other one,” or “Meka sɛ…” without making the speaker begin from zero.

The same principle appears in what happens when “unless” changes a Twi voice note. Small words can change the whole request. A voice AI needs to treat uncertainty carefully, especially when a sentence moves between languages.

Control matters after the conversation too

Adwoa eventually typed the medicine name herself and got a clearer way to frame her question for the pharmacy. She still had to make the call. The useful part was smaller: she could see where the system had failed, correct it, and decide what to do next.

That same clarity should extend to the conversation record.

Nkomo gives people explicit choices about cloud use: never, ask each time, or this session. Its history stays on the device unless you choose otherwise, and turning history off purges it immediately. You can export your data or delete your account in one tap.

These controls do not promise that every voice exchange will be perfect. They give people a plain view of what is happening when they speak, where their data goes, and what remains after the chat ends.

For someone asking a personal question in Twi, that clarity can shape whether they speak freely at all. Privacy is not one feature because consent, local history, export, and deletion each answer a different concern.

Build for repair, not performance

The most respectful voice experience has a repair path built into it. Let people interrupt. Show failures. Ask for the missing detail. Keep a typed option close by. Give people control over the record their conversation leaves behind.

Adwoa did not need a bot to sound clever at the trotro stop. She needed it to be honest when it had missed her meaning.

That is the standard worth holding: when a voice AI does not understand, it should make the next move easier for the person who does.

Sources (1)
  1. reddit.comGhanaian users voice skepticism about frustrating AI voice-agent experiences

Nkomo

A private, natural Twi and Ghanaian English voice-and-text companion — talk or type, in the mix of languages people actually speak, with clear control over what stays on the device versus what reaches the cloud.

Try Nkomo

Comments

No comments yet.