Talking to robots should make life easier

I recently spoke about how we communicate with robots. As often happens when I am nervous in front of a camera, some of the thoughts I wanted to share only came back afterwards. So I decided to write them down.

Two ideas were behind what I wanted to say about communicating with robots. The first is simple:

Communicating with a robot and teaching it should take less effort than the help it gives us is worth.

The second matters even more to me:

The point of the help is to give us more of our own time and more control over how we spend it, not to turn us into full-time operators of the machines meant to help us.

A robot might be able to perform an impressive task. But if explaining and demonstrating that task, correcting mistakes and supervising every step takes more effort than doing it ourselves, its usefulness is limited. That matters especially in homes and small businesses, where objects, routines and needs keep changing.

Words are only part of the conversation

Imagine teaching someone how to load your dishwasher. You might say, “Take this plate and put it here, like this,” while pointing and demonstrating the movement.

The words alone leave much of the instruction unspecified. Which plate? Where is “here”? What does “like this” mean? That’s why it is not simple and straightforward to just apply a large language model to robotics and to the real world.

We communicate through a combination of language, gestures, demonstrations and observed context of our surroundings. Each contributes something different: words can explain the goal, a gesture can identify an object, and a demonstration can show a movement that would be difficult to describe, all in the context of the current scene, task and interaction.

This is the kind of interaction I want to make possible with robots. A person should be able to explain what they need using familiar ways of communicating. The robot should connect those signals to the objects and actions in front of it.

Understanding my words is not the same as understanding my intention

“Clean the table” sounds straightforward. But perhaps I want the dishes removed while leaving my open book exactly where it is.

Even a person visiting my home for the first time would not know all my preferences. I would explain them. A useful robot needs a similar opportunity to learn.

Misunderstandings can also be very small. I point towards two cups and say, “Give me that one.” The robot correctly recognises both cups and every word, but selects the wrong cup.

This is why communication needs to work in both directions. The robot should be able to indicate what it has selected and ask, “Do you mean this one?”

There are prompting techniques for LLMs (e.g. #Grill-me) where the model interviews you before answering, building shared understanding through questions rather than guessing at your intent. I would like a similar exchange to be possible with robots.

In robotics, that exchange can also be visual. The robot might indicate an object or show a preview of its intended movement, allowing us to check whether its interpretation matches ours. The important question is how much dialogue or demonstration is useful, and when. A preview can help reveal a misunderstanding, although it cannot by itself guarantee that the physical action will succeed.

Sometimes the useful response is a question

Generative models give us powerful ways to produce language, plans and robot actions. But producing an output is not the same as recognising that more information is needed.

A system designed to map an instruction and an image directly to an action may generate an action even when the request is ambiguous or cannot be fulfilled. If I ask it to pick up a banana and there is no banana, I want it to notice the mismatch.

Generative models are not inherently unable to ask questions or seek information. But we should not assume that a model will reliably pause simply because it lacks something important. The useful next step might be to ask a question, look from another angle or check whether the object is present. The challenge is deciding when to do that. A robot that asks about every tiny detail can become just as burdensome as one that acts too confidently.

Trust should reflect what the robot can actually do

People readily interpret robot behaviour in human terms. In our experiments, a robot simply failing to recognise a gesture has prompted comments like “it doesn’t like me.” I do not think this proves emotional attachment. But it shows how quickly we reach for intentions to explain a machine, and how a robot that talks fluently will be assumed to understand more than it does.

That is a reason to be careful about fluency. A robot that talks smoothly should be equally good at saying what it cannot do and why it is unsure. Trust that matches the robot’s actual abilities is more durable than trust built on a pleasant manner — and it is also what prevents the opposite reaction, the fear that the machine is acting on its own.

A correction should help beyond this moment

There is another point I find important: responding to a correction now is different from learning something useful for tomorrow.

If I say, “Please leave that book on the table,” the robot might change its current action. But what should it remember? Should it leave every book on every table? Only this book? Only while I am reading it?

Similarly, “Don’t make me tea today” should not automatically become “I never want tea.” Yet if I have explained that I do not usually drink tea in the morning, I should not have to repeat that every day. Perhaps the preference is more specific: I enjoy morning tea, but not when I am in a hurry. Can the robot recognise that situation? What clues should it use, and when should it ask rather than assume?

This is what makes personalisation challenging: learning the conditions under which a preference applies, while leaving room for exceptions and changes of mind.

In our PersonalRobot project, we are interested in how robots can learn through observation, explanation and interactive dialogue. That includes deciding what to remember, when a previous experience applies and when to ask again.

Conversation needs a capable body and context

Explaining an action well does not guarantee being able to perform it.

A model may describe how to open a jar, while a robot still needs to detect slipping, adjust its grip and respond to resistance from the lid. Its reach, strength and sensors matter. So does continuous feedback from the physical world. Recent progress in language and reasoning can help robots with high-level planning, but those capabilities still need to connect to perception, control and physical experience. We cannot simply plug a large language model into a robotic body and expect it to know how to act.

The robot also needs to know what the task itself requires. In PersonalRobot we are investigating more explicit statistical and structural representations: which interpretation the available evidence supports, how uncertain the robot should be about it, and which steps and object relationships matter for success. Loading a dishwasher requires knowing that the door must be open before anything goes in, and that the plate has to be held at an angle that allows it. A robot that represents this explicitly can also say what it is unsure about, instead of failing silently.

Human demonstrations and corrections help build the connection between language, context and action in the real world. Yet teaching robots remains demanding. Many approaches rely on large datasets collected through teleoperation, with people guiding robots through tasks. How much additional training a new skill requires depends on what the robot already knows and how different the task is. Even a seemingly simple new task can require substantial effort to make it work reliably.

For practical use, that effort matters: teaching a robot must eventually take less effort than the work it saves us. The ability to perform a task is an important step. Making that ability easy to teach, adapt and use in everyday life is the next challenge.

The kind of helper I want

I want robots to become useful and pleasant helpers for particular people. I want robots that people can teach through ordinary communication, that adapt to their particular needs, and that give them more freedom to live as they choose.

A robot arriving from a factory might have basic skills, but it would still need to learn where things belong, which routines matter and which tasks its user actually wants help with. It should also learn when help is welcome and when we prefer to act independently. A craftsperson may want help with repetitive work while keeping the parts that require their personal touch. Someone receiving assistance at home may want support with tasks they find difficult while continuing to do the things they enjoy.

This brings me back to the thought at the beginning: robots should expand our possibilities and support our independence. We should have time to pursue our own activities, without becoming full-time operators of the machines meant to help us. My grandmother lives alone, and some everyday tasks are becoming harder for her. A robot that could bring something down from a high shelf, or open a jar that has been closed too tightly, could help her stay in her own home on her own terms. But only if it also recognises the mornings when she would rather manage by herself — and only if using it does not become another difficulty in her day.

That is what natural communication is for. It lets us explain, demonstrate and correct in ways we already know, and it leaves us deciding what to hand over and what to keep — making our lives our own.

The question I want us to keep asking is: after all the explaining, teaching and supervising, has the robot actually made this person’s life easier and more enjoyable?

Leave a Reply

Your email address will not be published.