From 6ba834bd925367b12a1c08b5720fb9fcd3f619b5 Mon Sep 17 00:00:00 2001 From: Sergey Arkhangelskiy Date: Thu, 3 Sep 2026 10:01:59 +0000 Subject: [PATCH 1/2] Lead with the five words, and put the customer above the argument MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The home page explained the service before it said what kind of evaluation it is, so the five words now sit under the lede: independent, continuous, commercial tasks, at scale, public and private. The Runway testimonial moves out of Protocol to just under the partner band. A customer's sentence outranks our own reasoning, so a reader meets it first. Problem opens on a published number — 64 picks an hour against a person's 1,300 — and drops from four points to two. A third way through sits between the argument and the protocol; the action was only at the top and the foot. About answers who rather than what: the founders' heading names what they did before, the independence paragraph no longer repeats the home page's service description, and 'rigs' becomes 'robots ... fixed stations', which is a word a reader outside this field already has. Requested-by: Inessa Roman Ticket: Positronic-Robotics/internal#1052 #refs --- content/pages/about.md | 6 ++-- content/pages/index.md | 10 ++++-- theme/positronic/static/css/home.css | 48 ++++++++++++++++++++++++++++ 3 files changed, 58 insertions(+), 6 deletions(-) diff --git a/content/pages/about.md b/content/pages/about.md index e3b9ec5..79a2683 100644 --- a/content/pages/about.md +++ b/content/pages/about.md @@ -5,11 +5,11 @@ URL: about.html Template: home Description: Positronic Robotics is two engineers from Google Search. We train no models and sell no robots, so no result on our rigs is a result about us. -

The best thing we can do for Physical AI is measure it properly

We met at Google, and we both worked on Search there. Evaluation there is a whole engineering discipline, and we took it for granted for years, until we found out that it is not a universally solved problem. Physical AI does not have one yet, so we are building it.

+

The best thing we can do for Physical AI is measure it properly

We met at Google and we both worked on Search ranking. Evaluation there is a whole engineering discipline, and we took it for granted for years — until we found that physical AI has nothing like it. So we are building it.

-
Founders

Both founders are engineers

Sergey Arkhangelskiy

Sergey Arkhangelskiy

Co-founder and CEO

Ten years at Google, including search ranking. Co-founded WANNA, an augmented-reality try-on company with Gucci and Louis Vuitton among its clients, and sold it to Farfetch in 2022.

Vladimir Yakunin

Vladimir Yakunin

Co-founder and CTO

Twelve years at Google, including Search. Then Snowflake, where he built the platform and performance infrastructure behind it. ICPC silver medallist.

+
Founders

Founded by two veterans of Google Search ranking

Sergey Arkhangelskiy

Sergey Arkhangelskiy

Co-founder and CEO

Ten years at Google, including search ranking. Co-founded WANNA, an augmented-reality try-on company with Gucci and Louis Vuitton among its clients, and sold it to Farfetch in 2022.

Vladimir Yakunin

Vladimir Yakunin

Co-founder and CTO

Twelve years at Google, including Search. Then Snowflake, where he built the platform and performance infrastructure behind it. ICPC silver medallist.

-
Independence

We train no models and we sell no robots

Nothing of ours is on the leaderboard, so no result of ours is a result about us. A private eval runs the same way: the same rigs, the same tasks and the same scoring as the public board, on your checkpoint, and the numbers go to you alone.

The rigs are in the EU (Cyprus), with an operator who resets the scene after every attempt. Every run is recorded, and the scoring is fixed before it runs: the method is in the PhAIL paper, the harness is on GitHub. On the leaderboard those recordings are public. On your eval they are yours.

+
Independence

We train no models and we sell no robots

Nothing of ours is on the leaderboard, so no result of ours is a result about us. We are not owned by a cloud provider, a robot maker or a model lab, and nobody pays us for an outcome.

The robots are in the EU (Cyprus) — fixed stations, each with an operator who resets the scene after every attempt. Every run is recorded, and the scoring is fixed before it runs: the method is in the PhAIL paper, the harness is on GitHub. On the leaderboard those recordings are public. On your eval they are yours.

Backing

Our pre-seed round is led by 33East, with participation from RTP, Davidovs Venture Collective, Orion VC and multiple angels.

Nebius is a founding partner of the leaderboard.

diff --git a/content/pages/index.md b/content/pages/index.md index 6f6bff4..2204e4a 100644 --- a/content/pages/index.md +++ b/content/pages/index.md @@ -4,14 +4,18 @@ Save_as: index.html URL: index.html Template: home -

The new standard of Physical AI eval

We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.

+

The new standard of Physical AI eval

We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.

Design partner
PhAILOur physical AI leaderboard
Founding partner of PhAIL
-
Problem

Every training run ends with the same question

Is this checkpoint better than the last one? On real robots that question has no cheap answer.

Someone has to put the world back after every try. Operators on shift, robots to keep alive. It still buys tens of rollouts a day, not thousands.

Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.

Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.

There is no single "better". Change the robot, the simulator or the metric and the winner changes. One rig and one number cannot settle it.

+
Design partner
Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Andy Chen, Head of Robotics, Runway
+ +
Problem

Every training run ends with the same question

The best model we have tested does 64 picks an hour. A person doing the same job by hand does over 1,300.

Is this checkpoint better than the last one? On real robots that question has no cheap answer: somebody resets the scene after every attempt, which buys tens of rollouts a day, not thousands.

Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.

Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.

Method

How teams solve it with us

Checkpoint in, rollouts out. The lab, the operators and the resets are ours. An infrastructure problem becomes an API call.

Many embodiments and simulators, same tasks, scoring fixed before anyone sees a result. Both checkpoints meet the same conditions, and nobody picks the metric that wins.

More signal from each rollout. Time-to-milestone scoring, so a call that needed hundreds of runs takes tens: the ~30x trial reduction in the PhAIL paper.

-
Protocol

What you send, what comes back

You send

A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.

We run

Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.

You get

Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.

Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Andy Chen, Head of Robotics, Runway
+

Have a checkpoint you cannot score? Tell us what you are training, or take half an hour.

+ +
Protocol

What you send, what comes back

You send

A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.

We run

Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.

You get

Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.

Talk to us

Bring us the checkpoint you cannot score

Leave an address and a line about what you are training, or take half an hour now. Either way we work out what is worth measuring on your setup, and what a first round would tell you.

diff --git a/theme/positronic/static/css/home.css b/theme/positronic/static/css/home.css index c104582..795d0e4 100644 --- a/theme/positronic/static/css/home.css +++ b/theme/positronic/static/css/home.css @@ -450,3 +450,51 @@ a.standing-mark:hover { color: var(--accent, #b1e74f); } @media (prefers-reduced-motion: reduce) { .home * { transition: none !important; } } + +/* ---- The five words, under the hero lede --------------------------------- */ +/* The claim before the explanation: what kind of evaluation, in five words a + reader takes in without reading a sentence. */ +.home .creds { + list-style: none; + margin: 26px 0 0; + padding: 0; + display: flex; + flex-direction: row; /* the global `ul` sets column; this row must opt out */ + flex-wrap: wrap; + gap: 10px 26px; + font-family: var(--mono); + font-size: 0.82rem; + letter-spacing: 0.06em; + text-transform: uppercase; + color: #aeb5a3; +} +.home .creds li { position: relative; } +.home .creds li + li::before { + content: ""; + position: absolute; + left: -14px; + top: 0.42em; + width: 3px; + height: 3px; + border-radius: 50%; + background: var(--accent, #b1e74f); +} + +/* The hook the Problem band opens on. */ +.home .lead-number { + font-size: 1.3rem; + line-height: 1.5; + color: #cdd2c4; + margin: 0 0 22px; +} +.home .lead-number b { color: var(--accent, #b1e74f); font-weight: 600; } + +/* A third way through, between the argument and the protocol. */ +.home .midcall { margin-top: 8px; } +.home .midcall .col p { font-size: 1.08rem; color: #cdd2c4; margin: 0; } +.home .midcall a { color: var(--accent, #b1e74f); } + +@media (max-width: 640px) { + .home .creds { gap: 8px 20px; font-size: 0.76rem; } + .home .lead-number { font-size: 1.15rem; } +} From 6505393c881fa8bda5c5487389f6805771c90290 Mon Sep 17 00:00:00 2001 From: Sergey Arkhangelskiy Date: Thu, 3 Sep 2026 10:06:15 +0000 Subject: [PATCH 2/2] Cut the rationale from the creds comment Ticket: Positronic-Robotics/internal#1052 #refs --- theme/positronic/static/css/home.css | 2 -- 1 file changed, 2 deletions(-) diff --git a/theme/positronic/static/css/home.css b/theme/positronic/static/css/home.css index 795d0e4..f5b8a3d 100644 --- a/theme/positronic/static/css/home.css +++ b/theme/positronic/static/css/home.css @@ -452,8 +452,6 @@ a.standing-mark:hover { color: var(--accent, #b1e74f); } } /* ---- The five words, under the hero lede --------------------------------- */ -/* The claim before the explanation: what kind of evaluation, in five words a - reader takes in without reading a sentence. */ .home .creds { list-style: none; margin: 26px 0 0;