CPIBench-0 is our first benchmark for evaluating how frontier models and agent harnesses perform in operations that rely heavily on cyber-physical systems.
Built from real projects across mechatronics, manufacturing, materials, and energy, CPIBench-0 tests models and agent harnesses in multimodal workflows where cyber-physical signals, tools, and real-world consequences intersect.
CPIBench-0 measures performance across connected cyber-physical workflows, from making reliable operational decisions to progressing safely toward real-world action.
The public dataset repo contains a selection of replayable example eval tasks and rollouts. Please contact us to perform the full eval suite.