📣 Help Shape the Future of UKRI's Gateway to Research (GtR)

We're improving UKRI's Gateway to Research and are seeking your input! If you would be interested in being interviewed about the improvements we're making and to have your say about how we can make GtR more user-friendly, impactful, and effective for the Research and Innovation community, please email gateway@ukri.org.

ACAS: Analysis of Complex Acoustic Scenes

Lead Research Organisation: University of Sheffield
Department Name: Computer Science

Abstract

For speech recognition technology to work reliably in noisy everyday environments, solutions are needed to the more general Machine Listening problems of sound source separation and acoustic scene understanding. This proposal seeks a four week visit to LabRosa at Columbia University, one of the internationally leading research groups working in these areas, with the aim of exploiting the synergy between their research and the statistical speech recognition research conducted by the proposer's host institution. The collaboration will focus on a dataset constructed from recordings made in a noisy domestic living environment that have been recorded as part of an existing EPSRC project, CHiME, for which the proposer is PI. The proposal outlines a number of immediate research lines that will be pursued tackling the problems of voice detection, source separation and binaural source localization. Possibilities for a longer term collaboration in the development of a model-driven acoustic scene analysis framework will also be discussed during the visit.

Planned Impact

Machine Listening is an under-developed research area which despite recent growth (e.g. EPSRC projects EP/D055466/1, EP/G007144/1, EP/G039046/1) still lags a long way behind its sister field, Computer Vision. The planned visit and the ensuing collaboration described in this proposal, will exploit the complementary expertise of the visitor and the host's institutions to push the machine listening field a necessary step forward in a number of directions. The immediate research goals focus on the problem of processing speech in noisy environments -- which is motivated by the enormous importance of robust speech recognition as a machine listening technology. However, this initial research is expected to foster a longer term collaboration the scope of which embraces the more general problem of extracting meaning from complex acoustic scenes. The project anticipates machine listening maturing to the point where the list of viable applications is limited only by the imagination. The ultimate beneficiaries of this research will then be the general public who will eventually be using Machine Listening applications as part of their everyday life. More broadly speaking, the UK as a whole will also benefit through the economic advantage arising from the control and/or deployment of Machine Listening technology. Machine Listening technology has the potential to enhance quality of life in a wide variety of ways. Application areas include: * Audio surveillance, i.e. machines able to listen to live audio recordings and act on the basis of their 'understanding' of the acoustic scene. Consider the security benefits, of audio-based burglar-detectors able to detect intruders in poor light; systems able to monitor the well-being of the elderly or infirm in their own homes, responding to the sound of a fall or a cry of distress, for example; public surveillance systems which use audio to supplement CCTV video and which can detect unexpected acoustic events and relay data to manned command centres. * Intelligent hearing aids able to suppress background noise without removing sound sources that might be significant to the listener. * Social robotics, i.e. robots that form a human-like model of the world to enable them to act and interact in an intuitive and natural manner. * Audio information indexing and retrieval -- e.g. systems that process audio or audio-visual archives (such as YouTube) and automatically add consistent semantic search tags. Such automated tagging will be a key technology in the expansion of the semantic web. Although some Machine Listening technology is finding its way into applications available today, the benefits of the mature technology will only be realised by concerted research funding over a long period. However, the field is sufficiently young that well-targeted small efforts have the potential to make a difference. Accordingly, the current proposal is targeting areas where the complementary expertise of the proposer and the host institution can have significant effect - source separation, environmental sound source modelling, model combination and model adaptation. Given the small size of this funding requested, every effort will be taken to maximise the impact of the project. Crucially, the proposer is currently engaged in a larger EPSRC-funded project, CHiME, that is committed to building an automatic speech recognition system built on a computational hearing framework. The current project is aimed to focus on CHiME data sets and the CHiME application scenario but to introduce complementary research -- that could not be pursued without international cooperations -- and that extends the scope of the CHiME proposal. This link to the CHiME project will enable the research outputs to feed directly into the instruments promoting CHiME, e.g. the CHiME demonstration system, the website, the CHiME corpus, the planned ASR competition and CHiME workshop.

Publications

10 25 50