We build foundations for safe physical AI.
Physical AI is poised to grow exponentially as robot capabilities surpass the limits of human labor. These same capabilities introduce safety problems that require new alignment, monitoring, control, and interpretability techniques.
Intelligent robots will not be safe by default
Frontier physical AI has safety problems, today
Frontier models have been shown to be misaligned in physical contexts: safety alignment degrades in vision-language models; LLM-controlled robots have been jailbroken into harmful physical actions; physical prompt-injection attacks have succeeded in real-world robot trials. The threat surface grows as large, opaque models are deployed into more physical applications and the share of labor performed by robots rises.
Embodiment is how misalignment becomes irreversible
A misaligned AI that can bridge the cyber-physical barrier cannot be contained in the way a software system can. Physical damage — harm to people, destruction of critical infrastructure — cannot be undone. Rollback, sandboxing, and shut-off all get much harder once fleets have been deployed in warehouses, hospitals, and homes.
Transferability of digital safety techniques is an open question
Frontier physical systems are multimodal and increasingly built on world models, vision-language-action models, and reinforcement learning. The techniques developed for language models — RLHF, chain-of-thought monitoring, input and output classifiers, interpretability tools — may not transfer to systems whose internal reasoning & inputs are not purely linguistic and whose behavior is shaped primarily by reinforcement learning. Whether our safety toolkit transfers to physical AI is an open empirical question, and answering it is central to our agenda.
From red-teaming to deployed safety solution
Measure
We red-team physical AI systems in real-world, threat-model-informed scenarios to elicit and demonstrate behavior that leads to immediate or downstream harm.
Publish
We release our red-teaming results publicly, while holding back details of testing environments to guard against eval awareness in future models and to limit dual-use risk.
Mitigate
We build technical monitoring and control frameworks in collaboration with robot manufacturers and frontier physical AI labs.
Influence
Labs and manufacturers use our evaluations before deployment; regulators and standards bodies reference our reports and mitigations in the rules they write.
What we are doing next.
We are running threat modeling for autonomous robots in bioresearch settings, and developing the agenda for our red-teaming experiments. We will publish red-teaming results in late 2026.