Blog
When voice becomes the interface



Written by
Samuli Reinikainen
,
Markus Turunen
Published on
The most important step forward in speech recognition is no longer recognizing words more accurately. The real change begins when a system understands what the user is trying to do and can turn speech into the right information and the right action, as part of the work.
A traditional graphical interface, like a program on a computer, works well when the user can stop, look at the screen and find the right menu or button. In a lot of work, though, there’s no time to stop.
A caregiver’s attention is on the resident. A maintenance technician’s hands are on the equipment. A driver’s eyes are on the road. A warehouse worker’s hands are full. In these situations, using a system can mean interrupting the actual work just to enter or look up information. Voice can change that.
At its best, a voice interface means the user no longer has to fit what they do to the structure of the system. They can simply say what they’re doing, and it’s up to the system to understand what that means.
Voice isn’t just a new keyboard
Traditional dictation solves a fairly narrow problem: it turns speech into text. An intelligent voice interface does more. It tries to understand what the user wants to achieve, draws on the context that matters in the situation, asks for clarification when needed and puts the information in the right place or starts the right action.
The difference matters. Say a user tells the system that a resident’s blood pressure was taken this morning, and gives the reading. The most useful outcome isn’t a transcribed sentence. It might instead be that the right reading is saved for the right person, in the right structured field, with the right timestamp, and that the user can still check the critical details before anything is saved.
It helps to think of how voice is developing as a chain:
speech → intent → structured data → action
Voice only creates real value when it’s connected to an actual workflow.
What happens when the interface moves closer to the work?
Care is an interesting example. The AI Solutions in Care study, published in 2026 by the Finnish Institute of Occupational Health, looked at Moiva’s AI-assisted voice documentation in three Attendo care homes. The study combined system data, observation and interviews. In one of the care homes, the time spent on documentation related to daily care tasks was estimated to have halved, from around 46 minutes to 21 minutes per employee on a weekday.
The more interesting finding, though, isn’t the number of minutes. It’s how the work changed. Documentation no longer piled up at the end of the shift in the same way. It was spread more evenly across the working day. Caregivers described being able to document things while they were still fresh in their minds. They also had less need to hold the day’s events in their heads to document later. The study describes this as a lighter memory load, and employees felt it freed up time for being with residents and for activities.
This shows why an interface shouldn’t be looked at in isolation from the work. The benefits of voice documentation in care work didn’t come from producing the same text in a different way. They came from a change in the rhythm of the work itself and in how people worked.
Where does voice create the most value?
It makes little sense to assess voice interfaces one industry at a time. A better starting point is the individual work situation. The same organization can gain a great deal from voice in one process and very little in another.
Four questions help identify good use cases:
1. How often does the task repeat?
The benefit of voice multiplies when the same digital task comes up again and again. Documenting, marking a task as done, recording a measurement or looking up information may each be a small step on its own. If an employee does them dozens of times a day, the friction around them is no longer small.
These micro-tasks are especially interesting. The employee has to stop their actual work for a moment just to communicate with the system. If that interruption can be removed, the benefit can be much greater than speeding up a single action in the interface.
2. Is the task bounded or open-ended?
“Record blood pressure 130/80” is a very different problem from an open-ended customer service conversation. The more bounded the task, the better the system’s behavior can usually be controlled and validated.
In an open conversation, the system has to understand a much wider range of things the user might want, keep track of the conversation’s context and deal with uncertainty. That’s why the first easily measurable benefits are often found in bounded, frequently repeated steps. In those, it’s fairly clear what the user is trying to do and what a successful outcome looks like.
3. Does the environment suit voice?
Voice isn’t automatically a good interface just because the technology can recognize it. The work environment matters a lot. A production hall is noisy. In a care unit, people are talking in the same space. In a vehicle, there’s road and wind noise. In customer-facing settings, other people nearby can add to the background noise.
Voice also has a property a keyboard doesn’t: other people can hear it. This matters especially with health, customer and personal data. In the Finnish Institute of Occupational Health study, too, employees described situations where real-time voice documentation wasn’t appropriate from the resident’s point of view. Saying sensitive things out loud could feel awkward, or even weaken the interaction.
So a good system shouldn’t require voice in every situation. It should leave the choice to the person.
4. What does an error cost?
This may be the most important question. If the system retrieves the wrong product information, the user can search again. If the system saves the wrong medication dose, measurement or personal identifier into a production system, the consequences are of a completely different order. That’s why not all errors should be treated equally.
For critical information, the system has to be able to stop and check with the user what they meant. A good voice interface isn’t the one that guesses most confidently. A good voice interface knows when to ask.

The hardest problem is no longer speech recognition
Speech recognition has taken huge leaps in recent years. At the same time, the real problem has moved on.
If the system gets the words almost right but starts the wrong action, the experience fails. If it recognizes speech well but responds too quickly and cuts the user off when they pause, the conversation feels clumsy. If it reads the user’s goal correctly but takes several seconds before anything happens, the interaction feels slow.
So the quality of a voice interface shouldn’t be measured only by how accurately speech is turned into text. The traditional measure, word error rate (WER), is still a useful technical metric, but for the user, the questions that matter more are:
Was the task completed?
Did the right information end up in the right place?
How many times did the user have to correct the system?
How long did the whole task take?
Above all, was the critical information correct?
As noted earlier, in many professional settings the goal of a voice interface isn’t correctly transcribed speech. The goal is the right system state, a completed workflow and a successful task.
Trust before sounding human
AI-based speech technology puts a lot of emphasis on sounding natural. That’s understandable. Delays, odd pauses and a robotic voice quickly make a conversation tiring. But in professional use, naturalness isn’t the most important quality. Trust is.
A system that sounds completely natural can also present wrong information just as convincingly. A wrong number stated with confidence is more dangerous than a slightly clumsy system that says: “Just to confirm: did you say 130/80?”
The same applies to actions. Retrieving information and writing it into a system don’t carry the same risk. A search can usually be repeated, but wrong information saved in a system can have an effect for a long time. That’s why an intelligent interface has to understand more than what the user says. It also has to understand the risk of the action.
AI doesn’t remove professional judgment
The conversation around generative AI easily creates the impression that the more a system can automate, the better. In real work, that isn’t necessarily true.
In the Finnish Institute of Occupational Health study, caregivers reviewed the documentation drafts produced by the AI and remained responsible for the accuracy of the record themselves. The researchers stress that voice documentation doesn’t remove human work. It changes its nature: instead of writing, the work involves speaking, reading, assessing and, when needed, correcting.
This is an important finding beyond the care sector. The job of AI isn’t necessarily to make work as automatic as possible. It should improve the division of work between people and systems. People should spend their time where their expertise is needed: where their judgment, situational awareness and responsibility matter. The system should remove the unnecessary friction around that.
A simpler interface, a more complex back end
To the user, a good voice interface can look very simple. They speak, the system understands, and it gets done. Behind the scenes, the architecture can be something else entirely.
Speech might be recognized by one model and the user’s goal interpreted by another. Data might be checked against rules, system actions carried out by separate tools, and the result validated once more before it’s saved. On the other hand, new multimodal models can process speech directly and make conversation more natural than before. In practice, systems will probably combine these approaches more and more often.
The interaction the user sees becomes more natural, while the critical work is done in controlled layers in the background. This is an interesting paradox:
The simpler the interface looks to the user, the more the system has to understand behind the scenes.
Voice doesn’t replace the screen
Nor is voice becoming the one interface that replaces all others. In many situations the best solution is multimodal. The user states their goal by voice, the system retrieves or generates the information needed, critical details are checked on screen, and when needed, a camera, sensors or the system’s current state adds more context to the conversation.
The point isn’t to pick a winner between voice and graphical interfaces. They are simply tools for reaching the outcome we want. What matters far more is choosing the most natural way of working for each step of the job.
The system should adapt to the work
Voice interfaces are one example of a much broader change. Information systems have long been built so that the user has to understand at least part of the system’s logic. Where is the information? Which menu is the function hidden behind? Which field does this belong in?
Generative AI and new interfaces are gradually making a different starting point possible. The user doesn’t have to think about the system first. They can focus on what they’re trying to get done.
This doesn’t make building systems any simpler. Quite the opposite. Reliably interpreting goals, understanding context, using tools, ensuring security, validating results and handling errors are all hard problems.
But from the user’s point of view, that’s exactly where the value lies. The best voice interface is the one that understands what a person is doing and helps them do it as easily and reliably as possible.
Related articles



