Nearly a million translated rows can help a system learn patterns, but they cannot capture every way Twi and English meet around a kitchen table. Colloquial phrases, regional variation, tone, and mid-sentence switches carry meaning that translation pairs alone may miss.
A sentence can change before the words end
At 8:12 p.m., Ama stood in her mother’s kitchen in Kumasi, holding her phone above a bowl of peeled kontomire. Her younger brother had sent a voice note about money for a school matter, and her mother wanted help replying before he made a decision that would be hard to undo.
Ama began in Twi, then moved into English for the detail: “Meka kyerɛ no sɛ, wait small, because the amount no, we need to check first.” She stopped. The phrase was not a neat translation exercise. “Wait small” carried urgency without panic. “The amount no” pointed back to the exact figure both people already knew. The Twi opening softened the instruction because this was family.
A system trained mostly on clean, one-language-at-a-time rows could pull those pieces apart. It might translate words correctly and still miss the relationship, the caution, and the reason Ama chose that phrasing. If the reply sounded too firm, her brother might send the money before the family had checked. If it sounded too vague, he might ignore it.
The stakes sat in that small voice note. A kitchen-table phrase had to carry the right meaning before someone acted.
Translation data teaches patterns, conversation tests them
Parallel datasets are valuable. They give language technology examples of how one expression can map to another. The Pristine Twi–English Parallel Dataset represents serious work toward making Twi more visible in language resources.
Yet a translated row usually has a tidy job: connect one written sentence to its counterpart. Everyday conversation is messier in useful ways.
A speaker may begin, “Medaase, but…” because gratitude makes room for disagreement. They may use English for a work term, a school subject, a number, or a name, then return to Twi before the thought is finished. A word that sounds ordinary in one home may land differently in another. Someone speaking quickly into a phone may restart halfway through a sentence, correct themselves, or trail off because the listener already understands the rest.
Those are not errors around the edges of language. They are part of how people speak.
This is why a conversational AI needs evaluation beyond the number of rows in its training material. Can it follow a speaker who changes languages mid-thought? Can it recognise when “yɛbɛhwɛ” means a careful “we will see,” rather than a firm promise? Can it respond in a way that fits the language mix the person chose?
What Happens When English Leads Into a Twi Clause? explores one version of that problem: the meaning can depend on where the switch happens, not only on the words around it.
Dialect and context keep language alive
Twi is spoken by people with different families, rhythms, local expressions, and habits of emphasis. A useful system should treat that variation with care.
That does not mean pretending every phrase has one universal interpretation. It means showing uncertainty when the meaning is unclear, listening to how the person speaks, and avoiding the false confidence that makes a reply sound fluent while getting the point wrong.
Ama eventually records a shorter message. “Medaase for telling us. Wait small, yɛnhwɛ amount no first.” Her mother listens once, nods, and asks her to send it.
The message works because it sounds like the family’s own way of speaking. The English detail stays clear. The Twi carries the warmth and caution. Nothing has been flattened into a classroom exercise.
For a voice-and-text companion, this is the standard worth aiming for. Nkomo is built for natural Twi and Ghanaian English conversation, typed or spoken, including the ordinary code-switching people use when one language alone does not quite do the job. That work calls for continual, careful evaluation. Dataset size is one input. Real conversations reveal what still needs to be understood.
Better language tools start with better listening
People should be able to speak naturally without first selecting a language for every turn. They should also be able to correct a voice assistant mid-sentence, hear a response, and know when something has gone wrong instead of receiving silence.
Privacy matters in these conversations too. A family message, health concern, or money question can be personal. Clear choices about cloud consent and local history give people a say in what leaves the device and what remains available afterward.
Ama puts the phone down beside the kontomire. Her brother receives a reply that sounds familiar enough to pause. The family gets time to check the amount together.
That is the practical test for language technology: can it help someone say what they mean when the meaning lives between Twi, English, context, and trust?
Comments
No comments yet.