Physical AI · Data & deployments

Data and deployments infrastructure for Physical AI. Capture. Annotate. Deploy.

DeepBuild collects egocentric, multimodal data of humans doing real work, annotates it densely, and runs the infrastructure that puts models in the field and keeps them learning.

✦ Example session · Kitchen prep ● Delivered
OP Operator · Kitchen Stereo · IMU

“Dinner service prep: wash, chop, portion, plate. Bimanual, tool use, wet and deformable objects.”

  • ✓Captured — stereo cameras + 6-axis IMU, syncedCapture
  • ✓SLAM — 6-DoF head & hand trajectoriesProcess
  • ✓Segmented — atomic action steps, verbs + objectsLabel
  • ▮▮QC — second-annotator review, flagged steps re-labeledReview
  • ✓Released — aligned streams, robot-retargeted hand posesShip
QC passed1600×1200 · 30 fpsRetargeted
Aligned · labeled · reviewed
Modalities
Stereo RGBsynced pair, 1600×1200 · 30 fps Stereo depthcomputed from the stereo pair IMU6-axis gyro + accel, synced to video Hand & head posederived, 6-DoF
01/Capture

Real work, recorded from the first person.

Kitchens, warehouses, labs, homes and factories. Our operators wear a synchronized rig and do the job — no scripted tabletop tasks, no lab-only lighting.

  • Head-mounted stereo cameras and IMU on one clock
  • Hand and head pose recovered from the egocentric streams — no external rig
  • Long-horizon tasks with tool use, bimanual manipulation and deformables
✦ Capture rigSynchronized
Stereo cameras — synced pair, 1600×1200, 30 fps, 180° FOVLive
Head IMU — 6-axis, gyro + accelerometerLive
Hand tracking — 21 keypoints per hand, from videoDerived
Stereo depth — computed from the stereo pairDerived
Head & hand pose — from SLAM + hand trackingDerived
All streamsHardware-synced
✦ Annotation pipelinePer session
LayerOutputStatus
Action steps — verb + objectPer step● Verified
6-DoF hand & head posePer frame● Verified
Object masks & contact eventsPer step● Verified
Robot retargetingYour embodimentQC review
Quality controlTwo annotators + automated checks
02/Annotate

Dense labels, with a second pair of eyes on every step.

Every session is segmented into atomic action steps and aligned across modalities. Two annotators plus automated checks, so you train on labels, not noise.

  • Atomic verb–object steps, contact events and success signals per step
  • Our SLAM stack recovers full 6-DoF trajectories through fast, difficult motion
  • Hand poses retargeted to your robot's embodiment on request
03/Train

Data mixtures that move the number.

We run controlled ablations on our own data so you don't have to. Same model, same fine-tune, same eval — only the pre-training mixture changes.

  • Off-the-shelf datasets by environment, task family and modality
  • Custom collection scoped to your embodiment and target tasks
  • Held-out evaluation suites and baseline comparisons delivered with the data
✦ Ablation — pre-training mix onlyModel fixed
Model architectureHeld fixed
Fine-tune recipeHeld fixed
Held-out evaluation tasksHeld fixed
Pre-training mixture: DeepBuild vs public datasetsVaried
ResultReported with every dataset
✦ Fleet dashboardLive
Robot health & uptimeLive
Autonomous episodes loggedContinuous
Operator interventions capturedOn-site
Episodes back to training setNightly
Policy rolloutsCanary
Task successTracked per policy
04/Deploy

Put it on robots. Keep the data coming back.

Deployment infrastructure for the field: fleet telemetry, intervention capture, and a nightly loop that turns every episode into training data.

  • Fleet monitoring, episode logging and intervention capture on-site
  • Human-in-the-loop intervention capture when the policy asks for help
  • Data flywheel: field episodes annotated and returned to your training set
Why DeepBuild

Hours aren't the metric. Task success is.

Public manipulation datasets are large, but scripted, single-modality and lab-bound. Models trained on them stall the moment they meet a real kitchen.

Public datasets
  • ✕Tabletop tasks under lab lighting, seconds long
  • ✕Single RGB stream, sparse or crowd-sourced labels
  • ✕No contact events, no hand pose you can retarget
  • ✕Collection ends at release; nothing flows back from the field
DeepBuild
  • ✓Real work in real environments, minutes to hours per task
  • ✓Aligned multimodal streams, two annotators per step
  • ✓Contact events and 6-DoF hand poses retargeted to your robot
  • ✓Deployment loop returns field episodes to training nightly
Get in touch

Researchers: get the dashboard and samples.

Tell us your embodiment, target tasks and modalities. We'll send dashboard access, sample sessions and a quote within two business days.

On the shelfStereo RGB, stereo depth, IMU and hand & head pose data
FormatsAligned streams, per-step labels, retargeted poses; MCAP, HDF5 or LeRobot
CustomScoped collection from 50 hours; your sites or ours
LicensingResearch and commercial, per-dataset or subscription
Physical AI

We still barely understand how humans move, handle and work. The data is how we find out.

DeepBuild's datasets and deployment infrastructure are the foundation for automating manual labor and for understanding cognition, spatial computing and prosthetics.

EgocentricMultimodalDensely annotatedField deployed