Self-Driven and Cross-Identity Animation

The method reconstructs fine facial structure, eye blinks, lip motion, and expression-dependent details while following either the subject's own motion or a different driving identity.
Recent head-avatar methods can reproduce facial motion but often miss identity-specific details. DipGuava introduces a structured two-stage representation: the first stage learns a geometry-driven base appearance, and the second predicts personalized residuals such as wrinkles and subtle skin deformation. Dynamic appearance fusion integrates these residuals after geometric deformation, maintaining spatial and semantic alignment. The disentangled design produces photorealistic avatars with stronger identity preservation and expression fidelity than prior approaches.

The method reconstructs fine facial structure, eye blinks, lip motion, and expression-dependent details while following either the subject's own motion or a different driving identity.

Component studies show that base appearance, personalized residuals, geometric deformation, and dynamic appearance fusion play complementary roles in preserving both global structure and high-frequency detail.

Personalized residuals adapt wrinkles and other local features to expression and identity, instead of treating them as fixed textures.
@inproceedings{lee2026dipguava,
title={{DipGuava}: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video},
author={Lee, Jeonghaeng and Choi, Seok Keun and Li, Zhixuan and Lin, Weisi and Lee, Sanghoon},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
year={2026}
}