<?xml version="1.0" encoding="UTF-8"?><ns2:project xmlns:ns1="http://gtr.rcuk.ac.uk/gtr/api" xmlns:ns2="http://gtr.rcuk.ac.uk/gtr/api/project" xmlns:ns3="http://gtr.rcuk.ac.uk/gtr/api/fund" xmlns:ns4="http://gtr.rcuk.ac.uk/gtr/api/person" xmlns:ns5="http://gtr.rcuk.ac.uk/gtr/api/project/outcome" xmlns:ns6="http://gtr.rcuk.ac.uk/gtr/api/organisation" ns1:created="2026-08-26T13:36:10Z" ns1:href="http://gtr.ukri.org/gtr/api/projects/7FF3049B-D131-4573-95E4-7801D8A6251C" ns1:id="7FF3049B-D131-4573-95E4-7801D8A6251C"><ns1:links><ns1:link ns1:href="http://gtr.ukri.org/gtr/api/persons/8B1BCC2C-D0E8-4C59-9013-506A96CBEA2D" ns1:rel="PM_PER"/><ns1:link ns1:href="http://gtr.ukri.org/gtr/api/organisations/6B6D9EF7-83F2-41DD-A995-9F2BAC6A5DCD" ns1:rel="LEAD_ORG"/><ns1:link ns1:href="http://gtr.ukri.org/gtr/api/organisations/6B6D9EF7-83F2-41DD-A995-9F2BAC6A5DCD" ns1:rel="PARTICIPANT_ORG"/><ns1:link ns1:end="2019-02-28T00:00:00Z" ns1:href="http://gtr.ukri.org/gtr/api/funds/181E43A9-2741-401B-9244-C031E4371749" ns1:rel="FUND" ns1:start="2018-03-01T00:00:00Z"/></ns1:links><ns2:identifiers><ns2:identifier ns2:type="RCUK">133525</ns2:identifier></ns2:identifiers><ns2:title>Visual understanding of faces in motion</ns2:title><ns2:status>Closed</ns2:status><ns2:grantCategory>Feasibility Studies</ns2:grantCategory><ns2:leadFunder>Innovate UK</ns2:leadFunder><ns2:abstractText>Recent days have seen an explosion of markerless facial performance capture methods that can reconstruct the dense geometry of the face from multi-view or even a single video. Unfortunately, even the best of these methods fail to capture the extensive range of shapes and deformations displayed by lips in motion, which results in a lack of expression in video synthesis of faces in speech. Lips exhibit agile and highly deformable motions which are extremely hard to capture -- at the same time, subtle lip motions are often the main vehicle for humans to convey emotions. **Capturing 3D lip motion with high accuracy** is therefore crucial if we are to design systems that can **synthesise realistic human speech** to convey human-like emotions.

**The key goal and main innovation in VISIM will be to construct a low-cost, class-leading, three-dimensional (3D) morphable model (3DMM) of the face that focuses on representing the detailed dynamics of the mouth in speech to achieve unprecedented levels of realism in video synthesis of human speech.**

We will achieve this by constructing a lightweight, multi-view, high frame-rate, video capture system to reconstruct the motion of the mouth and perform further lip shape refinement using photometric consistency and exploiting the redundancy in multiple views. We will use this to capture a range of subjects and train a new, richer 3DMM that encodes the detailed 3D dynamics of different people in speech.

The **business value for Synthesia** is clear: this new, highly expressive 3DMM model will be instrumental in allowing us to **automate** the accurate capture of detailed 3D lip motion just from **monocular video** using an analysis-by-synthesis approach. We will then apply this to high-end facial video synthesis for media and entertainment industries. VISIM will create impact by bringing the high-levels of realism at the affordable price of a low-cost, video-based, automated capture system. This will pave the way to democratising high-end facial performance capture, allowing it to be adopted in new markets such as language dubbing for the TV, movie and advertising industries, where photorealistic results are critical but production budgets have so far restricted access to high-end content creation. This technology will enable Synthesia, a newly formed UK SME, to establish a new market in 'natural' language dubbing with photorealistic synthesis of the face.

Keywords: Facial Performance Capture, Video-Based Tracking, Digital Content Creation.</ns2:abstractText></ns2:project>