NkomoNkomo
← All posts

Twi voice assistant: What Kwabena’s Self-Corrections Teach About Natural Speech

A useful Twi voice assistant must handle the speech people produce after careful recording speech gives way to ordinary conversation: pauses, repairs, English words, changed minds, and all. Those moments reveal meaning, intent, and uncertainty that disappear when every turn is spoken like a script.

At 6:40 p.m., Kwabena is halfway through a volunteer recording session at his kitchen table in Kumasi. His phone is propped against a blue mug. He has repeated the same prompt twice because he wants every word to sound clear.

Then he hears a message from his sister in the next room and answers without thinking.

“Ɛhe na mekae no? The appointment no, I mean, tomorrow morning... anaa wait, Friday no.”

He stops, laughs quietly, and tries again.

The recording task asks for casual phone conversation, yet the careful voice has already disappeared. What comes out instead is familiar: Twi carrying the shape of the thought, English arriving for a word that feels faster or more exact, then a repair because the first version was not quite right.

For an assistant, that repair matters. Kwabena has not failed to speak clearly. He is doing what people do when they are working out a thought in real time.

Casual speech contains the parts clean recordings miss

A polished sentence can teach a system pronunciation. A real conversation has more to teach.

People pause to remember a name. They restart a sentence because the listener has misunderstood. They say “the meeting” in English, then switch back to Twi for the part that carries the feeling. They interrupt themselves when a better word arrives. Sometimes the important information comes after “wait,” “I mean,” or “anaa.”

Kwabena’s sentence is short, but it contains several decisions an assistant has to make. The assistant needs to keep listening through the pause. It needs to recognise that “tomorrow morning” was corrected to Friday. It needs to treat the English word “appointment” as part of the conversation, rather than as a reason to reject the turn or force a language switch.

This is why natural speech data matters. An assistant cannot serve a Ghanaian conversation well by expecting one language per turn, one tidy sentence per request, or a speaker who has rehearsed before touching the microphone.

An August 2026 recruitment post for Twi speakers sought roughly 20 hours of casual phone conversation. The request points to a practical need: conversational speech has to sound like conversation. The hesitations and corrections are part of the material, not noise to remove before anyone can learn from it.

A pause can carry the real instruction

Imagine Kwabena trying to send a quick update while he is packing food into containers for the next day. He holds down the microphone and says, “Send a message to Auntie, sɛ mebɛba no... no, tell her I may come late.”

The first phrasing could create the wrong message. If an assistant acts too quickly, Auntie receives certainty when Kwabena meant uncertainty. If it waits for the repair, it has a better chance of preserving what he actually wanted to say.

That is the small risk inside ordinary speech. A pause may look empty in a transcript, but it can be where the speaker changes the plan.

Nkomo is built for natural Twi and Ghanaian English conversation by voice or text, including the mix of languages people actually speak. In hands-free voice mode, you hold to speak, can interrupt naturally, and hear the reply. The goal is a conversation where correcting yourself feels normal, because it is.

That same principle shaped the language picker Nkomo removed. A person should not have to stop mid-thought and choose a language before they can continue a sentence that already contains both.

Good voice interaction leaves room for repair

Speech tools often create pressure to get the first try right. That pressure changes how people talk. They slow down, avoid Twi words they expect a system to miss, or turn a simple request into stiff English because it feels safer.

Kwabena started his recording session that way. By the time he reached the later prompts, the effort of sounding perfect had faded. He was speaking as he would to a cousin on the phone: beginning in Twi, borrowing an English word, backing up when the detail changed.

That is the voice an assistant needs to meet.

The practical standard is simple. A person should be able to say a request, revise it before the thought is finished, and see when something has gone wrong. Nkomo does not swallow errors silently. When an error occurs, it is shown, giving the person a chance to decide what to do next instead of wondering whether the assistant heard them.

The same care applies to what happens around the conversation. Voice can feel more personal than typed text, especially when it includes family updates, money questions, or a message someone wants to rehearse before sending. Nkomo gives clear cloud-consent choices: never, ask each time, or this session. Its on-device history can be turned off and is then purged immediately. Data export and account deletion are available in one tap.

Those controls do not make a sentence easier to understand. They make it easier to choose how the sentence is handled.

Record the living language, then design for it

The lesson for anyone collecting speech data or building voice experiences is to protect the moment when the recording voice disappears. Ask for natural conversation. Keep the false starts. Mark the corrections. Notice where English enters without ceremony. Treat interruptions as part of the exchange.

Kwabena finishes his session with a sentence he would never have planned at the start: “Medaase, but can you make it shorter? M’ani nnye this one.”

He has given a correction, a preference, a switch in language, and a reason. Tomorrow, when he reaches for his phone to shape a family message, he should be able to speak that way from the first word.

Nkomo

A private, natural Twi and Ghanaian English voice-and-text companion — talk or type, in the mix of languages people actually speak, with clear control over what stays on the device versus what reaches the cloud.

Try Nkomo

Comments

No comments yet.