Project
robostats
An open-source statistics layer for robot policy evaluation: rigorous two-model and k-model comparisons with exact confidence intervals.
- Role
- Author, maintainer
- Dates
- Sep 2026 – present
- Stack
- Python, dependency-free core
Existing VLA evaluation harnesses report success rates with no statistics layer: no intervals, no paired tests, no accounting for shared episodes. robostats is an open-source Python library that adds one. It supports statistically rigorous dual-model and k-model policy comparison experiments, with exact confidence interval outputs under completely paired or partially overlapping evaluation scenarios.
It ships with benchmark adapters and a dependency-free episode recorder for LIBERO, RoboTwin, and RoboDojo outputs. The project grew out of evaluation work during my Noematrix internship, where the gap between a reported number and a defensible claim was hard to ignore.