RLE-Bench: a llama operating a robot arm

A Qualifying Exam for Coding Agents
as Robot Learning Engineers

RLE-Bench evaluates whether general-purpose coding agents can solve robot learning engineering problems through observation, experimentation, and iterative development. The benchmark spans four aspects: interactive control, policy development, perception and estimation, and mechanical design. Agents work under task-specific time, interaction, and compute budgets, and are evaluated independently under hidden physical conditions.

RLE Index

Task scores are averaged within each workflow, then across workflows, on a 0–100 scale.

Task Breakdown

Nine tasks across four workflows. Numbering and definitions follow the research overview.

Performance and Cost

Task scores, API cost, agent time, and context length, averaged over a task's subtasks, then within each workflow, then across workflows. Cost excludes simulation, GPU compute, and other infrastructure.

Mean Task Score / API Cost

Score vs. Cost

Higher score · Lower cost ↖
API cost uses a logarithmic scale.

Cost Breakdown

Scroll horizontally to view all cost metrics →