Splat2Act

From Captured Scenes to
Personalized VLA Manipulation

Xiaomeng Ye1Krishna Jaganathan1Gyubeum Lim1Gunjan Chhablani2Muhammad Zubair Irshad3Zsolt Kira1
1 2 3
See it in action
A captured indoor lounge flows into a Gaussian-splat simulation of Stretch, then the real robot picking up a personalized object in the same scene.
01Your scene & objectsA handheld scan. A few photos.
02Personalized simulationAutomated tasks and policy training.
03Real-world actionDeploy in the space you captured.
Explore the research

THE IDEA

Your space. Your objects.
A policy that learns both.

A generalist robot policy has seen a lot of the world. It still needs to learn your world.

Splat2Act turns a phone scan and object images into a photorealistic digital twin, generates successful pickup demonstrations, and personalizes a vision-language-action policy for the captured environment. All training demonstrations come from simulation.

Human input ends at scene and object capture.

101captured indoor scenesBuilding-scale environments
852discovered receptaclesTables, shelves, and beyond
68,021pickup demonstrationsGenerated automatically
63.9%real-world pickup successOffice · 12 tasks, 36 attempts
Read the abstract

A user who brings a generalist manipulation policy into their home or office wants it to pick up their objects from their shelves, counters, and tables. Yet even a strong generalist like π0.5 achieves 0% pickup success in our captured deployment scenes without adaptation, and collecting real demonstrations in every user's environment is prohibitively expensive. We present Splat2Act, an automated real-to-sim-to-real pipeline that personalizes a pretrained vision-language-action (VLA) policy to a specific scene and object set using only a handheld phone scan and images of the user's objects. Unlike prior real-to-sim work that reconstructs individual tabletops, Splat2Act operates on building-scale indoor environments: it builds a photorealistic digital twin with 3D Gaussian Splatting and collision geometry, automatically discovers tabletop and non-tabletop receptacles, constructs feasible pickup tasks, and generates successful demonstrations with a privileged simulation oracle, with only scene and object capture required from a human. Applied to 101 scanned campus environments with a median navigable area of 104.87 m² and 7 receptacles, the pipeline yields 852 receptacles and over 68,000 pickup demonstrations. We fine-tune π0.5 in two stages: Stage 1 adapts to hundreds of public object assets, drawn either from the deployment scene alone or from all 101 scenes; Stage 2 personalizes to the user's own objects on the deployment scene's receptacles. Our strongest personalized policy reaches 62.5% aggregate pickup success in the target scene. Deployed on a real Stretch 3, the same policy achieves 63.9% success on a 12-task subset, closely tracking its 66.7% simulated success on those tasks. We find that Splat2Act is able to adapt to the target scene and objects from only simulation oracle data, suggesting an alternative way for personalizing VLA manipulation training using an automated real-to-sim-to-real approach.

THE REAL-TO-SIM-TO-REAL PIPELINE

Capture once.
Let simulation do the work.

From reconstruction to demonstrations to adaptation,
one continuous pipeline connects the captured world to robot action.

01 CAPTURE & RECONSTRUCTION

A familiar space,
made simulation-ready.

A 20–30 minute handheld RGB-D scan captures a building-scale environment. Gaussian splats preserve its appearance; an extracted mesh provides the geometry for collision and task planning.

3D Gaussian splatsPhotorealistic observationsCollision meshGeometry for interaction

DN-Splatter reconstruction · SuGaR mesh extraction

From a handheld scanRGB · depth · normals

Your objects, reconstructed.

One photo and measured dimensions become a movable 3D asset.

02 AUTOMATED TASK GENERATION

A building becomes
a training ground.

The pipeline discovers usable surfaces throughout the scene, including non-tabletop receptacles, and constructs feasible pickup tasks.

  1. 01

    Find the navigable space

    Build a navigation mesh around the robot’s geometry.

  2. 02

    Discover support surfaces

    Cluster upward-facing geometry into stable, reachable receptacles.

  3. 03

    Assemble feasible tasks

    Sample object placements, robot poses, and distractors; filter for reachability.

  4. 04

    Generate successful demonstrations

    A privileged geometric oracle plans, executes, and retries pickup attempts.

Fully automated after capture

03 DATA AT SCALE

Thousands of demonstrations.
No teleoperation.

Across 101 campus environments, Splat2Act generates 68,021 successful pickup trajectories over 852 automatically discovered receptacles.

273.46hours

of recorded control data
from automated simulation

Lounges, kitchens, libraries, offices, and more.
Median navigable area: 104.87 m² per scene.

04 TWO-STAGE POLICY ADAPTATION

Learn the task.
Then learn your world.

We fine-tune π0.5 in two stages: first with diverse public objects, then with captured user objects on the deployment scene’s receptacles.

Broad adaptation provides experience with the task and embodiment. Personalization grounds that experience in the scene and object inventory that matter at deployment.

π0.5Pretrained VLA
STAGE 1

Broad adaptation

1,024 demonstrations
Hundreds of public objects

Office1korMulti1k
Deployment scene or all 101 scenes
STAGE 2

Personalization

500–504 demonstrations
Captured objects + target scene

Your object inventory
Your scene’s receptacles
Scene-specific supervised fine-tuning
Personalized
robot policy

ON THE REAL ROBOT

Learn in simulation.
Act in the real world.

Simulation-trained policies run on Hello Robot Stretch 3,
using head and wrist RGB, proprioception, and a language prompt.

CLOSING THE LOOP

The same world,
on both sides.

Gaussian-splat observations preserve the visual character of the deployment scene. The personalized policy transfers from this reconstructed environment to the physical robot.

66.7%Simulation
63.9%Real world

Office1k + Office personalization
Same 12-task subset · 12 simulated episodes · 36 real attempts

SIMULATION
REAL WORLD
Before adaptation: the pretrained policy

Broad pretraining is a starting point.

The pretrained π0.5 policy scores 0% in both captured task banks, without prior exposure to the Stretch embodiment or this action space. Splat2Act supplies simulation experience for adaptation and scene-specific personalization.

WHAT THE EXPERIMENTS SHOW

Personalization matters.
So does where you learn.

Explore the effects of broad adaptation, scene alignment,
and the gap between simulation and deployment.

Pickup success in Office

40 configurations · easy + hard

FINDING 01
+17.5pp

A better starting point for personalization.

On Office, broad adaptation in the deployment scene followed by personalization reaches 62.5% success — 17.5 percentage points above direct personalization from the pretrained model.

The strongest adaptation source depends on the deployment scene.

View the original paper figures

Personalization is local.

These experiments study receptacle-local pickup. Success drops substantially on unseen lab tables, reinforcing the value of adapting to the actual deployment scene. Longer-horizon mobile manipulation remains a direction for future work.

THE PAPER

From captured scenes
to personalized action.

Splat2Act: From Captured Scenes to
Personalized VLA Manipulation

Dataset release is planned upon acceptance.

BibTeX
@misc{ye2026splat2act,
  title = {Splat2Act: From Captured Scenes to
           Personalized VLA Manipulation},
  author = {Ye, Xiaomeng and Jaganathan, Krishna
            and Lim, Gyubeum and Chhablani, Gunjan
            and Irshad, Muhammad Zubair and Kira, Zsolt},
  year = {2026}
}
Download .bib