Start with the smallest record that can disprove a claim

A first specification should not attempt to describe every robot. It should capture enough information to tell two genuinely different experiments apart and make failures diagnosable.

The record belongs in the contributor’s existing repository. A registry should index and verify it, not demand custody of the policy.

Required identity

Every result should bind to immutable revisions of the policy, manifest, benchmark, dataset where applicable, and execution environment. Human-readable version tags are useful, but content digests prevent a tag from changing underneath a result.

Required configuration

For low-cost LeRobot manipulation, the initial record should capture robot family, leader/follower configuration, end effector, motor and firmware family, calibration fingerprint, sensor arrangement, observation/action schema, control rate, compute target, and relevant dependency lockfile.

Required evaluation evidence

A result needs a named benchmark version, explicit success predicate, trial and success counts, intervention count, reset procedure, evaluator identity, timestamps, and evidence hashes. The platform should derive the displayed percentage from counts rather than accepting an unexplained percentage.

Trust must be graduated

Self-tested, independently reproduced, controlled-runner verified, and laboratory verified are different claims. Displaying them as levels lets the network grow before formal certification exists without pretending every submission has equal assurance.

The interview question behind every field

For each proposed field we will ask: did this information affect whether you could reproduce the policy or diagnose the failure? Fields that never change a decision should not burden contributors. Missing fields that repeatedly cost time become required.