The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
ViDiHand satisfies the target properties of occlusion robustness, accuracy, and temporal smoothness for 4D hand recovery.
Click a case below, then switch between our results and method comparison.
The VACE branch is finetuned with hand-overlay rendering while the base DiT remains frozen, yielding a hand-aware video diffusion model.
A lightweight dual-branch decoder extracts MANO pose, 2D joints, and translation from a single intermediate VACE feature.
At inference, the same feature is decoded in a single VACE pass.
Comparison on three egocentric benchmarks. ARCTIC and HOT3D are in-distribution; HOI4D is a held-out cross-dataset comparison fair to all methods.
| Method | Detection | 3D Pose | Orient. & Position | Temporal | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| FAcc ↑ | Recall ↑ | F1 ↑ | MPJPE-p ↓ | PA-MPJPE-p ↓ | EPE-p ↓ | GO-p ↓ | CT-p ↓ | Jitter ↓ | ||
| ARCTIC | InterWild | 0.878 | 0.943 | 0.959 | 30.817 | 15.952 | 53.888 | 25.386 | 0.097 | 46.577 |
| HaMeR | 0.875 | 0.943 | 0.957 | 29.197 | 14.596 | 65.289 | 24.907 | 0.095 | 18.279 | |
| Hamba | 0.833 | 0.912 | 0.941 | 31.233 | 17.168 | 87.047 | 27.822 | 0.110 | 15.357 | |
| WildHands | 0.879 | 0.946 | 0.960 | 25.704 | 13.941 | 50.517 | 22.320 | 0.058 | 12.972 | |
| WiLoR | 0.919 | 0.951 | 0.974 | 22.012 | 11.873 | 71.527 | 17.358 | 0.075 | 24.091 | |
| Dyn-HaMR | 0.842 | 0.918 | 0.951 | 27.904 | 17.017 | 85.723 | 25.951 | 0.121 | 12.840 | |
| HaWoR | 0.700 | 0.817 | 0.895 | 45.357 | 26.375 | 158.062 | 43.325 | 0.149 | 19.789 | |
| OmniHands | 0.866 | 0.949 | 0.954 | 29.674 | 14.203 | 51.505 | 24.580 | 0.087 | 45.312 | |
| ViDiHand (Ours) | 0.997 | 0.999 | 0.999 | 21.668 | 9.821 | 12.407 | 14.642 | 0.047 | 3.183 | |
| HOT3D | InterWild | 0.669 | 0.881 | 0.868 | 77.168 | 24.811 | 71.482 | 58.501 | 0.213 | 101.164 |
| HaMeR | 0.692 | 0.904 | 0.883 | 68.314 | 21.455 | 59.077 | 49.636 | 0.102 | 23.632 | |
| Hamba | 0.632 | 0.828 | 0.853 | 71.732 | 29.620 | 107.625 | 56.525 | 0.128 | 18.507 | |
| WildHands | 0.655 | 0.863 | 0.844 | 52.791 | 28.946 | 111.438 | 53.933 | 0.157 | 22.885 | |
| WiLoR | 0.827 | 0.897 | 0.937 | 30.966 | 19.980 | 72.978 | 25.746 | 0.098 | 17.976 | |
| Dyn-HaMR | 0.614 | 0.811 | 0.802 | 74.214 | 38.201 | 171.617 | 43.851 | 0.571 | 44.942 | |
| HaWoR | 0.348 | 0.499 | 0.654 | 71.396 | 66.031 | 327.294 | 79.350 | 0.262 | 23.872 | |
| OmniHands | 0.649 | 0.895 | 0.868 | 63.281 | 22.682 | 68.437 | 49.120 | 0.133 | 69.510 | |
| ViDiHand (Ours) | 0.948 | 0.974 | 0.983 | 21.514 | 11.383 | 14.953 | 15.829 | 0.040 | 3.741 | |
| HOI4D | InterWild | 0.731 | 0.922 | 0.864 | 53.072 | 22.909 | 80.549 | 41.743 | 0.228 | 98.866 |
| HaMeR | 0.731 | 0.923 | 0.864 | 44.481 | 21.580 | 79.494 | 33.557 | 0.187 | 20.068 | |
| Hamba | 0.710 | 0.885 | 0.849 | 47.161 | 25.924 | 115.793 | 37.390 | 0.204 | 21.556 | |
| WildHands | 0.730 | 0.924 | 0.864 | 45.623 | 23.601 | 82.246 | 45.654 | 0.159 | 18.615 | |
| WiLoR | 0.962 | 0.966 | 0.972 | 33.710 | 14.903 | 41.579 | 25.527 | 0.115 | 17.449 | |
| Dyn-HaMR | 0.750 | 0.863 | 0.845 | 45.097 | 29.259 | 144.643 | 40.176 | 0.258 | 17.947 | |
| HaWoR | 0.869 | 0.864 | 0.919 | 47.329 | 28.851 | 135.748 | 43.091 | 0.139 | 28.376 | |
| OmniHands | 0.655 | 0.937 | 0.834 | 44.255 | 18.689 | 70.662 | 34.392 | 0.108 | 24.212 | |
| ViDiHand (Ours) | 0.984 | 0.991 | 0.990 | 30.090 | 13.960 | 24.460 | 23.420 | 0.117 | 4.010 | |
If you find ViDiHand useful for your research, please consider citing our paper.
@misc{wang2026surprisingeffectivenessvideodiffusion,
title={The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction},
author={Yuxi Wang and Chengkai Jin and Yufei Liu and Wenqi Ouyang and Tianyi Wei and Zhiwei Zeng and Siyuan Huang and Zhiqi Shen and Xingang Pan},
year={2026},
eprint={2606.30308},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.30308},
}