You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI models are becoming the backbone of robotics planning and manipulation, with increasily technical advances and capabilities. The class of world models are specially interesting considering the fact that they let a robot imagine/predict the consequences of its actions before executing them. But one drawback is that today's world models are optimized for capability instead of a balance with safety aspects.
This project develops open-source methods, benchmarks, and research that let the community explore the unsafe behaviours of action-conditioned world models applied in robotics, before they reach the real physical environments world.
One of our goals is demonstrate that safety mechanisms can be trained into action-conditioned world models without sacrificing task performance.
And our initial exploration will be in action conditioned world models, but extending future directions to vison-language-action models (VLA) and world-action models (WAM).
We have already released a JAX-based library for action-conditioned world models, called xwm, and published a survey of predictive safety in embodied AI exploring 236 papers that maps where the field's blind spots are related with safety aspects. This project turns that roadmap into working methods.
The project will be hosted in Kamara Lab
Goal 1 : We will integrate safety constraints methods directly into action-conditioned world models and benchmark them against leading unconstrained methods, exploring if the trade-off safety vs accuracy is true for this class of models. The target is SOTA task efficacy with safety, demonstrating that safety constraints do not have necessarily to degrade performance.
Goal 2 : our library xwm already provides the tools and model families (JEPA, TD-MPC2, MuZero) on a common interface, plus a Franka FR3 environment in Nvidia Newton physics and readers for datasets used in robotics learning. We will extend it with safety-evaluation suites with sim-to-real pipelines, and keep it documented, tested, and open.
Goal 3 : Starting with a robotic arm (matching the Franka environment already in xwm), our goal is evaluate the proposed methods on physical hardware and report the sim-to-real gap for safety metrics. As resources and partnerships allow, we will extend evaluation to further embodiments (like quadrupeds, humanoids, and drones).
Goal 4 : We aim to publish at leading AI and robotics venues (NeurIPS, ICLR, CoRL, RSS, ICRA, ICWM), run an open competition on safety in world models for robotics to find contributors and baselines, and propose a Safety in World Models for Robotics workshop at a major conference in 2027/2028.
A research-grade manipulator with gripper, cameras, and safety enclosure. This is the single largest item and the one that makes our third goal possible: without physical hardware, safety claims stay in simulation.
GPU hours for training world models and running large-scale simulation benchmarks (5,000 A100/H100-hours across the project).
Prize pool, participant compute credits, and organisation costs for an open Safety in World Models for Robotics competition built on xwm. Participants submit safety mechanisms evaluated on a shared benchmark. Winning entries are merged into the library as open baselines. This grows the contributor base and enhance trust in research and industrial communities.
No funds go to salaries. All research time is contributed by the team. And the grant covers hardware, compute, the competition, and dissemination only.
Minimum viable funding: With $15,000 (compute only) we can deliver the goals 1 and 2, simulation results and the library.
Kleyton da Costa: Ph.D. Student in Computer Science @ University College London and Research Scientist @ Holistic AI.
Advisors: Prof. Dimitrios Kanoulas (UCL, robotics) and Prof. Philip Treleaven (UCL) advise the underlying Ph.D. research.
We will also recruit new collaborators for the project based on the resources.
5K in computational resources in the last 2 months to partially fund the initial results