NkomoNkomo
← All posts

The Three Jobs Behind Hold, and What Cloud Consent Can Cost You

woman in black tank top holding silver and black trophy

Photo by Ato Aikins on Unsplash

Pressing and holding to speak starts three separate jobs: turning your speech into text, preparing a conversational reply, and checking whether cloud use has your permission. In Nkomo, each job has a different purpose, so each deserves a different kind of control.

In 1970, Apollo 13’s crew had a carbon dioxide problem after the oxygen-tank explosion changed the mission. Engineers in Mission Control in Houston had square lithium hydroxide canisters available from the command module, while the lunar module used round receptacles. The parts could help, but they could not connect as they were.

Ed Smylie led the team that developed an adapter procedure using materials available on the spacecraft. NASA’s Apollo 13 Flight Journal documents the effort. The crew followed the procedure, the improvised adapter worked, and the filters could do their job.

The lesson is practical: a system can contain several useful parts, yet each part still needs a clear connection and a clear role. When you press Hold in a voice assistant, speech recognition, response generation, and cloud consent are connected, but they are not the same job.

Speech recognition turns your voice into words

The first job begins when you speak. You might say, “Medaase, but this message to my landlord, can you make it softer?” Or you may start in English, switch into Twi for the part that carries the feeling, then return to English for the practical detail.

Speech recognition listens for the words you said and turns them into text the conversation can work with. Its job is accuracy and context at the level of what was spoken. A name, payment reference, place, or language switch can change the meaning of a request, so this step matters before any answer appears.

Speech recognition does not decide whether the reply is helpful. It does not decide whether the request should reach the cloud. It is the listening-and-transcribing part of the chain.

That distinction helps when something looks wrong. If your Twi phrase was heard incorrectly, the problem may be in the words captured from your voice. Nkomo shows errors rather than swallowing them, so a failed request does not pretend it succeeded. You can try again, rephrase, or type the part that needs to be exact.

For a closer look at why names and mixed-language speech need care, read Code-Switched Speech Translation: Why Accurate Names Matter in Mixed-Language Speech.

The conversational response works from what was understood

Once your speech has been captured, the next job is conversation. This is where Nkomo works with the meaning of your request and prepares a reply in natural Twi, Ghanaian English, or the mix you actually use.

A good response depends on more than recognising individual words. “Make it softer” can mean change the tone of a message. “Mepa wo kyɛw” can be part of the message itself, or it can signal how you want the whole message to sound. The conversational job is to keep those cues together and reply in a way that follows the thread.

This is also why interruption matters. In hands-free voice mode, you hold to speak, hear the reply, and can interrupt naturally. If the response is heading in the wrong direction, you do not have to wait through a long answer to correct it. Speak again and steer the conversation.

The Apollo 13 adapter did not remove carbon dioxide by itself. It made it possible for the canister designed to do that job to connect where it was needed. In the same way, a clean transcript is necessary, but it is only one input. The conversational response still has its own work to do.

Cloud consent is the permission check. It answers a different question from “Did Nkomo hear me?” and “Can Nkomo respond?” It asks whether this interaction may use the cloud under the choice you made.

Nkomo makes that choice explicit: never, ask each time, or this session. Choose “never” when you do not want cloud use. Choose “ask each time” when the context changes from one request to the next. Choose “this session” when you want that permission to last for the current session, rather than treating every spoken turn as the same kind of request.

That matters because a voice prompt can move quickly from ordinary to personal. You might begin by asking for help rewriting a note, then add a health detail, a family matter, or a payment reference. Consent should be something you can see and choose, not something buried behind a vague assumption.

Your conversation history is a separate control again. Nkomo keeps history on your device, and you can turn it off. When you do, it is purged immediately. You can also export your data or delete your account in one tap. Those controls concern what remains after an interaction, while cloud consent concerns whether cloud use is permitted for it.

The Cloud Choice Nkomo Asks You to Make Before You Speak explores that decision in more detail.

Use the control that matches the concern

When you press Hold, start with the question you actually need answered. If you are worried that Nkomo heard a name or Twi phrase incorrectly, check the speech and repeat or type it. If the words are right but the answer misses your meaning, interrupt and clarify the conversational request. If the content is sensitive, set cloud consent before you speak.

Apollo 13’s crew did not need one heroic object that solved every problem. They needed the right connection between separate parts, under pressure, with the limits visible. Voice conversations deserve the same clarity: know what is listening, know what is answering, and know what permission you have given before your words go further.

Nkomo

A private, natural Twi and Ghanaian English voice-and-text companion — talk or type, in the mix of languages people actually speak, with clear control over what stays on the device versus what reaches the cloud.

Try Nkomo

Comments

No comments yet.