Skip to content
All projects

Open source · Aug 2026 to Present

Inspect Robots Open-Source Contribution

A merged public API contribution to a robotics-evaluation framework, plus a backward-compatible metrics-layer proposal.

The problem

Benchmark authors needed one supported way to interpret human operator verdicts. The implementation lived behind a private helper, so downstream code either depended on private internals or recreated success semantics and risked drifting from the scorer.

What I built

I promoted verdict normalization to a public API, centralized the affirmative vocabulary and normalization rules, and added snapshot and regression coverage. I separately designed a metrics protocol and registry proposal so tasks could report mean, standard error, and success rate without breaking existing mean-only behavior.

Turning a private convention into a public contract

The operator scorer already knew which human verdicts counted as affirmative, but downstream benchmarks had no supported API for asking the same question. PR #415 promoted that logic into is_affirmative_verdict, made the scorer call it, and documented the helper for benchmark authors. The change is small in surface area and important in effect: there is now one contract instead of several nearly identical copies.

Tests define the compatibility boundary

The regression suite covers accepted verdict vocabulary, case and surrounding whitespace, rejected values, null input, and direct agreement between the helper and the operator scorer. An API snapshot also makes accidental public-surface changes visible in review.

The metrics layer is a proposal, not a shipped claim

I also designed a backward-compatible metrics layer with a metric protocol, its own registry namespace, and mean, standard error, and success-rate built-ins. Tasks would keep mean as the default, so existing benchmarks would not change. That design was sent to the maintainer for discussion; it is intentionally described here as proposed work rather than implemented functionality.

Outcome

  • Merged upstream as robocurve/inspect-robots PR #415
  • Public helper and operator scorer now share one verdict contract
  • Regression coverage includes case, whitespace, rejected values, null values, and scorer agreement
  • Metrics layer is explicitly a design proposal under maintainer discussion, not shipped code

Stack

PythonpytestPublic API designRobotics evaluationBenchmark metricsOpen source

Data flow

One public verdict contract, with a proposed metrics layer above it

  1. 1

    Operator verdict

    Human-written pass/fail language, including case, whitespace, and null edge cases

  2. 2

    Public normalizer

    is_affirmative_verdict centralizes the supported success vocabulary

  3. 3

    Operator scorer

    Calls the same public helper instead of maintaining separate semantics

  4. 4

    Regression suite

    API snapshot plus accepted, rejected, normalization, null, and agreement tests

  5. 5

    Proposed metric protocol

    Backward-compatible interface and independent registry namespace

  6. 6

    Proposed built-ins

    Mean, standard error, and success rate across scenes and epochs

PythonpytestPublic API designRobotics evaluationBenchmark metricsOpen source

This diagram is generated from portfolio.ts. Edit the `architecture` field to change it.

Want the deeper version of this?

I am happy to walk through the tradeoffs, the failure modes, and what I would do differently.

Email me