WorldEvals¶
The Inspect Evals for robotics.
A curated catalog of physical-AI / VLA benchmarks built on Inspect Robots. Each benchmark lives in its own repository (so it owns its release cadence, dependencies, and hardware notes); WorldEvals is the index that ties them together.
Two views of the collection
inspect-robots list tells you what Inspect Robots tasks are installed.
worldevals list tells you what benchmarks exist and how to get them.
Benchmarks¶
3 benchmarks · 16 tasks
-
A one-task smoke benchmark: lay a detached clapstick on a clapperboard centered 19 inches in front of the bimanual arm line. Exists to prove an eval stack works end to end before a full benchmark run.
1 tasks · bimanual ·
smoke-testbimanualmanipulationRepo ·
alpha -
10 bimanual kitchen-manipulation tasks: pick-place, stacking, slotted insertion, granular pour & tool-scoop, lid open/seal, cloth folding, a two-arm handover, and a multi-instance cutlery sort.
10 tasks · bimanual ·
kitchenbimanualmanipulationRepo ·
alpha -
5 bimanual desk-stationery tasks, one fixed setup each: uncap a marker, open a box, take out an eraser, and close the lid, pull the middle pad from a sticky-note stack, pour paper clips into a lifted bowl, and hand a ruler between arms cup to cup.
5 tasks · bimanual ·
stationerybimanualmanipulationRepo ·
alpha
Browse from the command line¶
pip install worldevals
worldevals list # all benchmarks
worldevals list --tag bimanual # filter by tag
worldevals info kitchenbench # repo, install command, task keys
worldevals tasks # Inspect Robots tasks installed locally, by benchmark
For LLMs: llms.txt ·
llms-full.txt.