Fluidity in simulated human-robot interaction with speech interfaces
Lead Research Organisation:
SWANSEA UNIVERSITY
Department Name: College of Science
Abstract
The need for interactive robots which can collaborate successfully with human beings is becoming important in the UK considering some of the biggest challenges we now face, including the need for high-value manufacturing exports to compete economically internationally, robots which can handle dangerous waste and navigate hazardous environments, and robotics solutions for social care and medical assistance to meet our demographic challenges.
A key problem for human-robot interaction (HRI) with speech which limits the wider use of such robots is lack of fluidity. Although there have been significant recent advances in robot vision, motion, manipulation and automatic speech recognition, state-of-the-art HRI is slow, laboured and fragile. The contrast with the speed, fluency and error tolerance of human-human interaction is substantial. The FLUIDITY project will develop technology to monitor, control and increase the interaction fluidity of robots with speech understanding capabilities, such that they become more natural and efficient to interact with. The project will also address the difficulty of developing HRI models due to the time, logistics and cost of working with real-world robots by developing a toolkit for building and testing interactive robot models in a simulated Virtual Reality (VR) environment, making scalable HRI experiments for the wider robotics, HRI and natural language processing (NLP) communities possible.
The project focusses on pick-and-place robots which manipulate household objects in view where users will utter commands (e.g. "put the remote control on the table") and issue confirmations and corrections and repairs of the robot's current actions appropriately (e.g. "no, the other table"), allowing rapid, natural responses from both a human confederate teleoperating the robot model and automatic systems. Crucially, appropriate overlap of human speech and robot motion will be permitted to allow more human-like transitions. The project will put interaction fluidity and the rapid recovery from misunderstanding with appropriate repair mechanisms at the heart of interactive robots, which will lead to improved user experience.
The means for achieving fluid interaction will firstly be adaptation of Spoken Language Understanding (SLU) algorithms which are not only word-by-word incremental but go beyond that for more human-like real-time measures of confidence the robot has in its interpretation of the user's speech. For the basis of these algorithms, mediated Wizard-of-Oz data will be collected from pairs of human participants, with one participant confederate 'wizard' controlling the robot model and one user. From the visual, audio and motion data collected, SLU algorithms will be built which return the most accurate user intention incrementally word-by-word, but also a continuous measure of confidence corresponding as closely as possible to the reaction times of the human confederate.
The project will also address user perception of the robot's intention from the robot's motion by experimenting with different models of motion legibility. The hypothesis is that the more accurately the legibility of the robot's motion can be modelled in real time, the greater the fluidity of interaction possible, as user repairs and confirmations can be interpreted appropriately earlier in the robot's motion.
The SLU and legibility algorithms will be integrated in an end-to-end system where interaction fluidity can be controlled, with evaluation in both the VR environment and a comparison to a real-world robot model. The project will provide an abstract theoretical framework for interaction fluidity and practical outcomes of a VR environment, an HRI dataset collected in the environment which will be made publicly available for benchmarking, and software which will be open-source and adaptable for other robot models.
A key problem for human-robot interaction (HRI) with speech which limits the wider use of such robots is lack of fluidity. Although there have been significant recent advances in robot vision, motion, manipulation and automatic speech recognition, state-of-the-art HRI is slow, laboured and fragile. The contrast with the speed, fluency and error tolerance of human-human interaction is substantial. The FLUIDITY project will develop technology to monitor, control and increase the interaction fluidity of robots with speech understanding capabilities, such that they become more natural and efficient to interact with. The project will also address the difficulty of developing HRI models due to the time, logistics and cost of working with real-world robots by developing a toolkit for building and testing interactive robot models in a simulated Virtual Reality (VR) environment, making scalable HRI experiments for the wider robotics, HRI and natural language processing (NLP) communities possible.
The project focusses on pick-and-place robots which manipulate household objects in view where users will utter commands (e.g. "put the remote control on the table") and issue confirmations and corrections and repairs of the robot's current actions appropriately (e.g. "no, the other table"), allowing rapid, natural responses from both a human confederate teleoperating the robot model and automatic systems. Crucially, appropriate overlap of human speech and robot motion will be permitted to allow more human-like transitions. The project will put interaction fluidity and the rapid recovery from misunderstanding with appropriate repair mechanisms at the heart of interactive robots, which will lead to improved user experience.
The means for achieving fluid interaction will firstly be adaptation of Spoken Language Understanding (SLU) algorithms which are not only word-by-word incremental but go beyond that for more human-like real-time measures of confidence the robot has in its interpretation of the user's speech. For the basis of these algorithms, mediated Wizard-of-Oz data will be collected from pairs of human participants, with one participant confederate 'wizard' controlling the robot model and one user. From the visual, audio and motion data collected, SLU algorithms will be built which return the most accurate user intention incrementally word-by-word, but also a continuous measure of confidence corresponding as closely as possible to the reaction times of the human confederate.
The project will also address user perception of the robot's intention from the robot's motion by experimenting with different models of motion legibility. The hypothesis is that the more accurately the legibility of the robot's motion can be modelled in real time, the greater the fluidity of interaction possible, as user repairs and confirmations can be interpreted appropriately earlier in the robot's motion.
The SLU and legibility algorithms will be integrated in an end-to-end system where interaction fluidity can be controlled, with evaluation in both the VR environment and a comparison to a real-world robot model. The project will provide an abstract theoretical framework for interaction fluidity and practical outcomes of a VR environment, an HRI dataset collected in the environment which will be made publicly available for benchmarking, and software which will be open-source and adaptable for other robot models.
Publications
Baptista De Lima C
(2026)
Achieving Interaction Fluidity in a Wizard-of-Oz Robotic System: A Prototype for Fluid Error-Correction
Foster M
(2025)
Aye, Robot: What Happens When Robots Speak Like Real People?
Förster F
(2025)
Editorial: Failures and repairs in human-robot communication.
in Frontiers in robotics and AI
Förster F
(2023)
Working with roubles and failures in conversation between humans and robots: workshop report.
in Frontiers in robotics and AI
Healey P
(2025)
Negotiating shared understanding: Coding repair in social interaction
in Research on Language and Social Interaction
| Description | The main achievements of the research and activity from the award are, firstly, a detailed exploration of HRI work leading to discovering the key technical requirements for achieving fluid interaction with robots (generally, but more specifically those which understand speech commands), and the development of an open-source simulation system which meets those requirements. A novel simulation system for the pick-and-place robot Fetch was developed using a Unity-based visualisation for users in a VR environment, and a robotic control system that was capable of the four key requirements for fluid interaction discovered by the project, namely: interruptibility and correction, pollability, latency measurement and optimisation and time-accurate reproducibility of actions from logging data. Other research achievements included the development of a Natural Language Understanding model of Conceptual Pact building of how conversational partners come to agree on the naming conventions on objects in a shared space over time, and the transfer of that model to an in-robot setting. Work on the role of errors and corrections within the in-robot system was also taken on in more depth than previous robotics research. In terms of partnerships and collaborations, new connections were made with Heriot-Watt University, the University of Surrey, the University of Louvain and other institutions. These colleagues were interested in improving Human-Robot Interaction through the methods pioneered in the project. Work was shared with small companies like Cereproc and larger ones like Honda at workshops and conferences, and the hosting of a workshop on the topic at HAI 2024. A summer school course on the speech-based control of the system developed was taught at the ESSLLI summer school. The Intelligent Robotics group at Swansea has continued to develop through the project, in terms of specialist PhD students in Natural Language Understanding and robot motion focussed on increasing the fluidity of the interactions. |
| Exploitation Route | In future, a more specific focus on the use of robots with the capacity for fluid interaction within the care sector would be a great opportunity for deployment-based research. For long-term assistant robots for people living alone, particularly those with dementia and other ageing-related conditions, would be a great opportunity with the right partnership and governmental support. Interactive robots are becoming part of everyday life, but for them to become truly usable, they have to have the capacity for fluid interaction. The FLUIDITY project has defined these criteria and implemented a simulation system - deploying this in real-world settings over a longer-term project to test the potential benefit in the real-world would be extremely valuable. |
| Sectors | Digital/Communication/Information Technologies (including Software) Healthcare Culture Heritage Museums and Collections Transport |
| URL | https://fluidity-project.github.io/ |
| Description | Swansea Intelligent Robotics Infrastructure |
| Amount | £300,000 (GBP) |
| Organisation | Higher Education Funding Council for Wales (HEFCW) |
| Sector | Public |
| Country | United Kingdom |
| Start | 12/2024 |
| Title | Conceptual Pact Models for Reference Resolution using Dynamic Small Language Models |
| Description | A reference resolution dataset and model for objects based on small language models which are dynamically constructed during conversations between participants, derived from the PENTOREF dataset. |
| Type Of Material | Computer model/algorithm |
| Year Produced | 2024 |
| Provided To Others? | Yes |
| Impact | The dataset and model are used by the EPSRC FLUIDITY project and by the EPSRC ARCIDUCA project as the basis for models of reference to objects in situated conversations. |
| URL | https://github.com/julianhough/conceptualpacts |
| Description | University of Hertfordshire collaboration |
| Organisation | University of Hertfordshire |
| Country | United Kingdom |
| Sector | Academic/University |
| PI Contribution | Leading the FLUIDITY project experiments and writing of papers. Co-hosting workshops and special sessions. |
| Collaborator Contribution | The time, expertise and use of the Hertfordshire Robot House. |
| Impact | See https://fluidity-project.github.io/publications for publications |
| Start Year | 2023 |
| Title | Python code for Conceptual Pact Models |
| Description | Python code for the 2024 LREC-COLING paper Hough et al., "Conceptual Pacts for Reference Resolution using Small, Dynamically Constructed Language Models: A Study in Puzzle Building Dialogues". |
| Type Of Technology | Software |
| Year Produced | 2024 |
| Open Source License? | Yes |
| Impact | Code is used by the EPSRC FLUIDITY project and by the EPSRC ARCIDUCA project as the technical implementation for models of reference to objects in situated conversations. |
| URL | https://github.com/julianhough/conceptualpacts/ |
| Description | Hosting of the workshop Fluidity in Human-Agent Interaction at the 12th International Conference on Human-Agent Interaction (HAI 2024) |
| Form Of Engagement Activity | Participation in an activity, workshop or similar |
| Part Of Official Scheme? | No |
| Geographic Reach | International |
| Primary Audience | Professional Practitioners |
| Results and Impact | Approximately 30 participants took part from academic and industrial research on the topic of fluidity in human-robot interaction. |
| Year(s) Of Engagement Activity | 2024 |
| URL | https://fluidity-project.github.io/fluidityhaiworkshop |