We met at Google, and we both worked on Search there. Evaluation there is a whole engineering discipline, and we took it for granted for years, until we found out that it is not a universally solved problem. Physical AI does not have one yet, so we are building it.
We met at Google and we both worked on Search ranking. Evaluation there is a whole engineering discipline, and we took it for granted for years — until we found that physical AI has nothing like it. So we are building it.

Sergey Arkhangelskiy
Co-founder and CEO
Ten years at Google, including search ranking. Co-founded WANNA, an augmented-reality try-on company with Gucci and Louis Vuitton among its clients, and sold it to Farfetch in 2022.

Sergey Arkhangelskiy
Co-founder and CEO
Ten years at Google, including search ranking. Co-founded WANNA, an augmented-reality try-on company with Gucci and Louis Vuitton among its clients, and sold it to Farfetch in 2022.
Nothing of ours is on the leaderboard, so no result of ours is a result about us. A private eval runs the same way: the same rigs, the same tasks and the same scoring as the public board, on your checkpoint, and the numbers go to you alone.
The rigs are in the EU (Cyprus), with an operator who resets the scene after every attempt. Every run is recorded, and the scoring is fixed before it runs: the method is in the PhAIL paper, the harness is on GitHub. On the leaderboard those recordings are public. On your eval they are yours.
Nothing of ours is on the leaderboard, so no result of ours is a result about us. We are not owned by a cloud provider, a robot maker or a model lab, and nobody pays us for an outcome.
The robots are in the EU (Cyprus) — fixed stations, each with an operator who resets the scene after every attempt. Every run is recorded, and the scoring is fixed before it runs: the method is in the PhAIL paper, the harness is on GitHub. On the leaderboard those recordings are public. On your eval they are yours.
Our pre-seed round is led by 33East, with participation from RTP, Davidovs Venture Collective, Orion VC and multiple angels.
Nebius is a founding partner of the leaderboard.
We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.
We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.
Is this checkpoint better than the last one? On real robots that question has no cheap answer.
Someone has to put the world back after every try. Operators on shift, robots to keep alive. It still buys tens of rollouts a day, not thousands.
Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.
Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.
There is no single "better". Change the robot, the simulator or the metric and the winner changes. One rig and one number cannot settle it.
Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
The best model we have tested does 64 picks an hour. A person doing the same job by hand does over 1,300.
Is this checkpoint better than the last one? On real robots that question has no cheap answer: somebody resets the scene after every attempt, which buys tens of rollouts a day, not thousands.
Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.
Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.
Checkpoint in, rollouts out. The lab, the operators and the resets are ours. An infrastructure problem becomes an API call.
Many embodiments and simulators, same tasks, scoring fixed before anyone sees a result. Both checkpoints meet the same conditions, and nobody picks the metric that wins.
More signal from each rollout. Time-to-milestone scoring, so a call that needed hundreds of runs takes tens: the ~30x trial reduction in the PhAIL paper.
A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.
Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.
Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.
Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Have a checkpoint you cannot score? Tell us what you are training, or take half an hour.
A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.
Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.
Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.
Leave an address and a line about what you are training, or take half an hour now. Either way we work out what is worth measuring on your setup, and what a first round would tell you.