How scoring works
· ·
“Evidence ready” is about evidence, not the robot
EVIDENCE READY FOR PUBLICATION means the evidence behind a profile meets our publication standard: every scoring-critical record is durably captured and second-reviewed, at least three score-bearing claims are publication-eligible, and evidence confidence is not Low or Very Low.
It is not a statement that the robot is mature, safe, commercially proven or recommended. A ready profile may report INSUFFICIENT DATA for capability or commercial maturity, and may carry low scores. Those are findings published with transparent insufficiencies.
Two robots in this set prove the point from opposite directions. Unitree G1 is evidence ready with no commercial score at all. Agility Digit is evidence ready with the lowest capability score in the set.
Capability and Commercial Reality are separate
Capability measures what independent evidence shows a robot can do, weighted by how much each capability matters for its category. Commercial Reality measures paying production deployment: confirmed customers, verified units in production, duration and independent confirmation.
We never combine them. A logistics robot with two paying customers and modest dexterity, and a research platform with strong laboratory results and no customers, are both real and different positions. One number would erase the distinction.
Why some robots show INSUFFICIENT DATA
When no publication-eligible evidence supports a dimension, we print INSUFFICIENT DATA rather than a zero. A zero is a measurement — it says we looked and found nothing there. INSUFFICIENT DATA says we have not established it. Reporting the first when we mean the second would be a false claim about a real company's product.
Why two claims can stay “Very Low” confidence
The confidence band is held at Very Low for any robot with fewer than three score-bearing publication-eligible claims, regardless of the calculated confidence figure. A score-bearing claim is one in a capability category that carries weight for that robot's type.
This means a robot can display a calculated confidence in the seventies and still read Very Low, because the calculation rests on too narrow an evidence base to describe as high confidence. When a third score-bearing claim arrives, the band can jump several steps at once. That discontinuity is deliberate and is a known characteristic of Engine v1 rather than a defect; it is documented here because a reader comparing two robots would otherwise find the gap inexplicable.
Evidence tiers
- T0 — Claimed
- The manufacturer asserts it. No demonstration.
- T1 — Demonstrated
- Shown, typically by the manufacturer.
- T2 — Independently Demonstrated
- Observed or measured by a party other than the maker.
- T3 — Commercially Deployed
- Independently confirmed paying production use.
- T4 — Proven at Scale
- Multiple confirmed customers, verified units, sustained duration and an established operating mode.
A claim can never exceed the ceiling its source supports. A manufacturer page cannot carry a T2 claim no matter how technical it is, because the tier describes who established the fact.
Operating mode
Every claim records whether operation was autonomous, supervised, teleoperated or unknown. Most evidence in this field does not establish autonomy, so Unknown is common and halves the credit a claim contributes. Commercial deployment is not proof of autonomy.
Source independence and durability
Sources are classified by independence — independent, partially independent, first party — and a claim's tier ceiling follows. A supplier-and-customer joint announcement is partially independent, not independent.
Every scoring source must also be durable: retrievable and verifiable later, via an external archive, a publisher-hosted artifact with a recorded hash, or our own captured snapshot with a SHA-256. A live URL alone is not evidence. Captured files are held for verification and are not redistributed.
Contracted versus verified production units
A contract for seven robots is not seven robots working. We record these as separate values and only verified production units affect any score. Digit shows seven contracted units at Toyota and zero verified production units.
Only the strongest claim per capability counts
Capability takes the maximum qualifying claim in each category, not a sum. A second independent source confirming the same ability improves confidence, not capability — because the robot did not become more capable by being watched twice.