Dexterous Tactile World Model

Predict future egocentric human manipulation videos with whole-hand touch

Ziyao Zeng1, Xiatao Sun1, Hao Wang2, Yueyang Pan3, Zhengxiang Yu4, Fengyu Yang1, Tianyu Liu1, Zhiwen Fan2, Daniel Rakita1 1Yale University  2Texas A&M University  3University of California, Los Angeles  4University of Washington

observed frame 0
left gloveright glove

pressure finger flexion

observed: RGB, hand pose, touch predicted: 36 frames at once 0 12 13 28 44 48 chunk 1 chunk 2 3

Predict the future with touch

World models predict the future frames of physical interaction.

They can generate training data, initialize policy models, and serve as simulators for embodied AI.

Force is invisible.

How much force the hand applies, and how it changes over time, decides how an interaction unfolds.

Dexterous touch measures them directly.

Whole-hand tactile gloves now record where and how firmly both hands press.

Can world models use dexterous touch to better predict the future?

Dexterous Tactile World Model

We embed each hand's tactile reading and add it to the video tokens at that hand's location, through a zero-initialized residual in a pretrained video diffusion transformer.

Observations t ≤ K
Observed frame
frames
Rendered hand skeleton
hand pose
touch

Task instruction

Video diffusion transformer, 30 blocks

tactile residual added in 5 of the 30 blocks

Embedeth = Whsth + frame(t) + hand(h)
PlaceGaussian footprint at the handfootprint g at the hand
Add as a residualu ← u + proj(LN(e)) g
Predicted frames t > K
Predicted frame
TrainK ∈ {4, …, 44}, flow-matching loss to ground truth
InferK = 12, predict 36 frames at once

Better visual quality, hand location and motion

tactile field, frame 12
Observation at frame 12
observation, frame 12
Conditioning footprint
conditioning footprint
Ground truth at frame 44
ground truth, frame 44
Vision-only baseline at frame 44
vision-only baseline
DTWM at frame 44
DTWM (ours)
TouchWorldFeelWorldVT-WMDTWM (ours)
Ground truth
DTWM (ours)
Vision-only baseline
TouchWorld
VT-WM
FeelWorld
observed frame 0

First predicted chunk, frames 13 to 28. Dashed boxes: ground-truth hand regions.

Better generalization to unseen objects & tasks

Ground truth
DTWM (ours)
Vision-only baseline
TouchWorld
VT-WM
FeelWorld
observed frame 0

Why touch helps?

Touch history correlates with future manipulation. The observed frames look alike, but the observed touch evolves differently, and so do the manipulations that follow.

force holds steady

Observed frame 12, force holds steadyobserved frame 12
Future frame 44: the grasp holdsgrasp holdsfuture frame 44

force decays

Observed frame 12, force decaysobserved frame 12
Future frame 44: the hands releasehands releasefuture frame 44
left handright hand

Changing touch → the interaction is about to change

A decaying touch leads to a release 2.5× as often (16%), a rising touch 1.8× as often (12%).

A rising force usually marks a short push, press or pinch in progress, which ends within a second and the hand lets go.

Training with touch helps, even without it at inference

Training with touch teaches the model to tell visually similar frames with different contact states. Even without touch at inference, it applies this learned representation to estimate current contact, which also reduces its uncertainty. An uncertain model averages over possible futures and produces blur.

Vision-only baselineTouch in training onlyTouch in training and inference (DTWM)

The benefit increases over the prediction horizon

Vision-only baselineDTWM (ours)

BibTeX

@article{zeng2026dtwm,
  title   = {Dexterous Tactile World Model},
  author  = {Zeng, Ziyao and Sun, Xiatao and Wang, Hao and Pan, Yueyang and Yu, Zhengxiang
             and Yang, Fengyu and Liu, Tianyu and Fan, Zhiwen and Rakita, Daniel},
  journal = {arXiv preprint arXiv:2609.34286},
  year    = {2026}
}