Robotics: Science and Systems (RSS 2026)

Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

Zongzheng Zhang*1,2, Zi Lin*1, Jiawen Yang1, Ziqiao Peng1, Junyan Lao1
Lin Cheng4, Huazhe Xu3, Hang Zhao3, Hao Zhao†1,2

1 Institute for AI Industry Research (AIR), Tsinghua University   2 Beijing Academy of Artificial Intelligence (BAAI)   3 Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University   4 Beihang University
* Equal contribution   Corresponding author

Overview

Overview of the end-to-end physical conversational face system across diverse interactions
Fig. 1: We demonstrate our end-to-end physical conversational face system across diverse multi-round interactions: (a) reenactments of Star Wars (Yoda-Luke Skywalker) and (b) Titanic (Rose-Jack); (c) a dialogue between a real elf and a virtual elf; (d) a three-character encounter from the Chinese folktale Legend of the White Snake; and (e) human-robot dialogue.

Abstract

Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making large-scale personalization prohibitively slow and costly. In this work, we pursue automated and scalable mechanical face synthesis, aiming to rapidly generate a physically realizable facial mechanism for a wide range of facial geometries. We introduce a parametric, linkage-driven mechanical face template whose topology and actuator layout are explicitly parameterized to support systematic scaling and retargeting across diverse facial morphologies. Building on this template, we propose a hierarchical automatic design algorithm that takes a single 2D portrait as input, reconstructs a target 3D face, and synthesizes a collision-free, manufacturable internal mechanism. The algorithm combines anatomy-guided feasible motion volumes, Action Unit (AU)-derived trajectory-based expressiveness objectives, and a collision-driven outer-loop refinement strategy. Beyond hardware synthesis, we argue that future mechanical faces deployed at scale must engage in bidirectional, multi-turn conversation, rather than functioning solely as speaking or listening heads. To this end, we develop a dual-identity conversational facial motion synthesis framework that jointly models speaking and listening behaviors from audio, producing temporally coherent 3D facial motion suitable for physical execution. We validate our system through extensive experiments, including (i) quantitative evaluation of automatic mechanism synthesis across diverse facial geometries, (ii) comparisons against manual mechanical design, (iii) benchmarks on conversational facial motion synthesis and real-time deployment, and (iv) perceptual user studies.

Hardware Template

The system starts from a reusable, linkage-driven face template whose modules can be scaled and retargeted to new facial geometries.

Mechanical face template with full assembly and modular mechanisms
Fig. 2: Mechanical face template. (a) Full assembly of the linkage-driven robotic face with a soft skin and a 3-DoF neck module. (b) Exploded view of the four modular facial mechanisms: eyebrow, eyes, mouth, and jaw.

Automatic Mechanical Design

Given a portrait, the pipeline reconstructs the target face, initializes module placement, optimizes local mechanisms, and resolves collisions in the final CAD assembly.

Hierarchical automatic mechanical design pipeline
Fig. 3: Overview of the hierarchical automatic design pipeline. (a) From a 2D portrait, we reconstruct a 3D head mesh, semantic landmarks, and initial module base poses. (b) The inner loop performs module-wise kinematic synthesis under anatomy-guided feasible volumes and AU-derived trajectories. (c) The outer loop assembles the global CAD model, detects interferences, and applies MTV-guided updates until a collision-free assembly is obtained.

Diverse Facial Mechanism Synthesis

The same template generalizes across human, stylized, and non-human faces, producing valid internal mechanisms at relative physical scale.

Automated mechanism synthesis across diverse facial morphologies
Fig. 4: Automated mechanism synthesis across a spectrum of facial morphologies. The optimized internal CAD assemblies span cinematic characters, a stylized elf, and characters from the Chinese folktale Legend of the White Snake, confirming that the pipeline generates collision-free linkage assemblies across diverse target faces.

Interaction Synthesis and Control

Beyond hardware generation, the system synthesizes conversational facial motion and maps it to physically executable motor commands.

Overview of interaction synthesis and control
Fig. 5: Overview of interaction synthesis and control. (a) Given dual-speaker audio, the talking-head model predicts multi-round facial motion for both speaking and listening. (b) The predicted coefficients are mapped to robot motor commands via region-wise regressors and neck IK, enabling physically feasible multi-round expressions on the mechanical face.

End-to-End Demonstrations

These supplementary demonstrations show the complete system in live human-robot dialogue, robot-robot reenactment, and mixed physical-digital interaction scenarios.

Real-Time Human-Robot Interaction

A live dialogue with the robotic head, demonstrating real-time inference, conversational turn-taking, lip synchronization, and affective facial motion during a semi-structured exchange.

Dyadic Robot-Robot Dialogue: Titanic

A physical reenactment between Rose and Jack, emphasizing soft emotional expression, responsive listening cues, and coordinated multi-round robot-robot interaction.

Dyadic Robot-Robot Dialogue: Star Wars

Luke Skywalker and Yoda are embodied as two physical faces, testing distinct speaking styles, listener reactions, and expressive timing across an iconic cinematic exchange.

Hybrid Avatar-Robot Interaction

A mixed-reality conversation between a digital avatar and the physical robot, showing how the motion synthesis framework transfers across heterogeneous embodiments.

BibTeX

@inproceedings{zhang2026automated,
  title     = {Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots},
  author    = {Zhang, Zongzheng and Lin, Zi and Yang, Jiawen and Peng, Ziqiao and Lao, Junyan and Cheng, Lin and Xu, Huazhe and Zhao, Hang and Zhao, Hao},
  journal   = {arXiv preprint arXiv:2607.11688},
  year      = {2026}
}