Mission

We build foundations for safe physical AI.

Physical AI is poised to grow exponentially as robot capabilities surpass the limits of human labor. These same capabilities introduce safety problems that require new alignment, monitoring, control, and interpretability techniques.

The problem

Intelligent robots will not be safe by default

Frontier physical AI has safety problems, today

Frontier models have been shown to be misaligned in physical contexts: safety alignment degrades in vision-language models; LLM-controlled robots have been jailbroken into harmful physical actions; physical prompt-injection attacks have succeeded in real-world robot trials. The threat surface grows as large, opaque models are deployed into more physical applications and the share of labor performed by robots rises.

Embodiment is how misalignment becomes irreversible

A misaligned AI that can bridge the cyber-physical barrier cannot be contained in the way a software system can. Physical damage — harm to people, destruction of critical infrastructure — cannot be undone. Rollback, sandboxing, and shut-off all get much harder once fleets have been deployed in warehouses, hospitals, and homes.

Transferability of digital safety techniques is an open question

Frontier physical systems are multimodal and increasingly built on world models, vision-language-action models, and reinforcement learning. The techniques developed for language models — RLHF, chain-of-thought monitoring, input and output classifiers, interpretability tools — may not transfer to systems whose internal reasoning & inputs are not purely linguistic and whose behavior is shaped primarily by reinforcement learning. Whether our safety toolkit transfers to physical AI is an open empirical question, and answering it is central to our agenda.


Theory of change

From red-teaming to deployed safety solution

Measure

We red-team physical AI systems in real-world, threat-model-informed scenarios to elicit and demonstrate behavior that leads to immediate or downstream harm.

Publish

We release our red-teaming results publicly, while holding back details of testing environments to guard against eval awareness in future models and to limit dual-use risk.

Mitigate

We build technical monitoring and control frameworks in collaboration with robot manufacturers and frontier physical AI labs.

Influence

Labs and manufacturers use our evaluations before deployment; regulators and standards bodies reference our reports and mitigations in the rules they write.

Roadmap

What we are doing next.

We are running threat modeling for autonomous robots in bioresearch settings, and developing the agenda for our red-teaming experiments. We will publish red-teaming results in late 2026.