A Qualifying Exam for Coding Agents
as Robot Learning Engineers
RLE-Bench evaluates whether general-purpose coding agents can solve robot learning engineering problems through observation, experimentation, and iterative development. The benchmark spans four aspects: interactive control, policy development, perception and estimation, and mechanical design. Agents work under task-specific time, interaction, and compute budgets, and are evaluated independently under hidden physical conditions.
RLE Index
Task scores are averaged within each workflow, then across workflows, on a 0–100 scale.
Task Breakdown
Nine tasks across four workflows. Numbering and definitions follow the research overview.
Performance and Cost
Task scores, API cost, agent time, and context length, averaged over a task's subtasks, then within each workflow, then across workflows. Cost excludes simulation, GPU compute, and other infrastructure.
Mean Task Score / API Cost
Score vs. Cost
Cost Breakdown
Scroll horizontally to view all cost metrics →