OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

Zhisheng Han1, Shiyao Wu1, Jiayan Qiu1, Yakun Ju1, Lu Liu2, Le Zhang3, Pengfei Feng4, Huiyu Zhou1, Zheheng Jiang1,*
1University of Leicester, Leicester, UK 2University of Exeter, Exeter, UK 3University of Birmingham, Birmingham, UK 4China University of Geoscience, Wuhan, China
*Corresponding author
OASIS teaser

Abstract

Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occlusion and the complex pose-dependent deformation of highly articulated hands. Existing methods predominantly rely on implicit NeRF-style representations, whose volumetric fitting is computationally expensive and often struggles to preserve fine-grained hand details. In this work, we present OASIS, a tailored 3D Gaussian Splatting framework for single-image hand avatar reconstruction. To faithfully encode sparse image-specific appearance cues in single-view reconstruction, we construct geometry-aligned visual evidence tokens by explicitly aligning input image observations with 3D hand geometry and context-adaptively tokenizing the resulting visual evidence. Since severe self-occlusion makes the reliability of image evidence inherently visibility-dependent, we introduce a visibility-conditioned point-image attention to reliably transfer visual evidence to geometric tokens, yielding occlusion-aware Gaussian features for faithful and robust reconstruction. To further capture non-rigid deformation of articulated hands, we introduce a Feature-on-Mesh representation to enable Gaussian deformation to be guided by local surface stretching. Under this framework, we adopt a one-shot adaptation scheme that learns a shared hand prior from multi-identity training data and then fits it to a target image for target-specific reconstruction. Extensive experiments show that OASIS outperforms existing baselines in both visual fidelity and efficiency across challenging poses and in-the-wild scenarios, and further demonstrates strong versatility in downstream applications such as text-to-avatar generation and texture editing. Code will be released upon paper acceptance.

Network Architecture

Network Architecture

Video Comparisons

Input Image 1 Input Image
Ours Preview 1
OHTA Preview 1
Ours OHTA Reconstruction results
Ours Animation
OHTA Animation
Input Image 2 Input Image
OHTA Preview 2
Ours Preview 2
Ours OHTA Reconstruction results
Ours Animation
OHTA Animation
Input Image 3 Input Image
Ours Preview 3
OHTA Preview 3
Ours OHTA Reconstruction results
Ours Animation
OHTA Animation
Input Image 3 Input Image
Ours Preview 3
OHTA Preview 3
Ours OHTA Reconstruction results
Ours Animation
OHTA Animation
Input Image 4 Input Image
Ours Preview 4
OHTA Preview 4
Ours OHTA Reconstruction results
Ours Animation
OHTA Animation

Experimental Results

Qualitative Comparison on InterHand2.6M

Qualitative comparison on InterHand2.6M

In the Wild Comparison from HanCo, COCO-Hand, and WHIM Dataset

In the wild comparison

Applications

Applications

BibTeX

@inproceedings{oasis2026,
  title={OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting},
  author={Zhisheng Han and Shiyao Wu and Jiayan Qiu and Yakun Ju and Lu Liu and Le Zhang and Pengfei Feng and Huiyu Zhou and Zheheng Jiang},
  booktitle={Proceedings of the 34th ACM International Conference on Multimedia},
  year={2026}
}