DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video

Jeonghaeng Lee1,Seok Keun Choi1,Zhixuan Li2,Weisi Lin2,Sanghoon Lee1
1Yonsei University, 2Nanyang Technological University
AAAI Conference on Artificial Intelligence, 2026
Overview of the two-stage DipGuava framework.

DipGuava separates stable geometry-driven appearance from personalized residual details, then fuses both after deformation to create expressive, identity-preserving 3D Gaussian head avatars from monocular video.

Abstract

Recent head-avatar methods can reproduce facial motion but often miss identity-specific details. DipGuava introduces a structured two-stage representation: the first stage learns a geometry-driven base appearance, and the second predicts personalized residuals such as wrinkles and subtle skin deformation. Dynamic appearance fusion integrates these residuals after geometric deformation, maintaining spatial and semantic alignment. The disentangled design produces photorealistic avatars with stronger identity preservation and expression fidelity than prior approaches.

Self-Driven and Cross-Identity Animation

Qualitative DipGuava results for self-driven animation and cross-identity reenactment.

The method reconstructs fine facial structure, eye blinks, lip motion, and expression-dependent details while following either the subject's own motion or a different driving identity.

Disentangled Personalized Details

Ablations of DipGuava components and dynamic appearance fusion.

Component studies show that base appearance, personalized residuals, geometric deformation, and dynamic appearance fusion play complementary roles in preserving both global structure and high-frequency detail.

Expression-Aware Dynamics

Expression-aware wrinkles and identity-specific residuals generated by DipGuava.

Personalized residuals adapt wrinkles and other local features to expression and identity, instead of treating them as fixed textures.

BibTeX

@inproceedings{lee2026dipguava,
  title={{DipGuava}: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video},
  author={Lee, Jeonghaeng and Choi, Seok Keun and Li, Zhixuan and Lin, Weisi and Lee, Sanghoon},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  year={2026}
}