Axis Robotics has launched Axis Sim Dataset V1, one of many largest open-source simulation datasets for Franka arm manipulation, with the complete dataset, coaching code, and benchmarks publicly obtainable. V1 is constructed from greater than 50,000 human-teleoperated simulation trajectories throughout 207 manipulation duties and 60,000+ scene variants on a simulated Franka Analysis 3 arm.
This dataset drew over 160,000 downloads, making it probably the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with each consequence open and verifiable.

Axis Robotics is constructing the last word compounding information engine for Bodily AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The corporate raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Community Ventures, 10K Ventures, and angel traders.
A Guess In opposition to “Clear Knowledge Solely”
A typical assumption in robotics is that demonstrations should be near-optimal to start with — filter all the way down to skilled trajectories, standardize the setup, and discard something noisy earlier than it’s secure to mimic. Axis’s thesis runs the opposite means: information high quality lives on the distribution stage, not the only trajectory. When a big and numerous sufficient crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives throughout coaching.
Axis Sim Dataset V1 places that thesis to a public take a look at. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and gear use, all collected by way of Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd moderately than a single skilled crew. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the College of Michigan, and different establishments.
Outcomes That Scale
On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Efficiency improves constantly as pretraining information scales from 25% to 100% of the dataset, with no saturation in sight, proof that the positive factors come from range and protection moderately than a one-off bump. The most important enhancements seem below digital camera, sensor-noise, and format perturbations, the precise axes Axis randomizes throughout technology.

The crew says V2 is already underway, scaling to 1.2 million trajectories throughout 1,200 duties, with cross-embodiment generalization and outcomes throughout a number of VLA fashions exhibiting that suboptimal simulation information trains strong insurance policies.
The Engine Behind the Dataset
The dataset is one output of a bigger, actively compounding information engine. The place a standard information vendor collects to a hard and fast spec and stops, Axis makes use of mannequin efficiency and failure instances to find out what must be collected subsequent, so each coaching spherical informs the subsequent. That engine runs on a hybrid technique throughout 4 information traces, and all 4 now run at scale:
- Simulation: over 200,000 distributed contributors on Axis Hub, a top-3 dApp on Base, producing 4.7M+ trajectories throughout 13 embodiments.
- Selfish: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in actual properties and companies throughout 14 industries: 200,000+ hours already banked and rising by 4,000+ hours daily, with Vicon-verified hand pose.
- Loco-manipulation: 500+ hours combining mobility and dexterity on actual humanoids (Unitree G1, Booster T2) by way of hardware-agnostic teleoperation.
- Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge instances.
Each activity and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.
From Open Knowledge to Business Deployment
Past open-sourcing simulation information, Axis works straight with robotic embodiment firms to construct personalized, embodiment-specific information pipelines and mannequin priors.
As Booster Robotics’ first sim-data accomplice, Axis rebuilt Booster’s actual workspace as a task-aligned digital twin, had distributed contributors acquire 42,000+ simulation episodes on it, and distilled them right into a Booster-specific mannequin prior. With simply 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box π0.5, matching π0.5 utilizing half the real-world demonstrations.
Different companions span embodiment firms (Feagine Robotics), mannequin firms (Manycore Tech, Dexmal) and industrial automation (Lotus Automobiles, Geely Auto). Axis additionally provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Bodily AI’s Knowledge Basis
“The way forward for Bodily AI isn’t a static dataset you obtain as soon as,” mentioned Chris Feng, founding father of Axis Robotics. “It’s an engine that retains producing the info the mannequin wants subsequent. Scale will get you broad protection. Range retains the noise unbiased. The closed loop turns each failure into progress. That’s what compounds.”
Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who’ve scaled client platforms to over 30 million customers. Its analysis is suggested by Jiachen Li, Assistant Professor at Georgia Tech.
Paper Hyperlink: https://arxiv.org/abs/2607.21588
Venture Web page: https://axisaiorg.github.io/AXIS-V1/
Dataset Hyperlink: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Coaching
