Help shape the benchmark
Contribute to RLE-Bench 2.0
Bring the problems you solve as a robot learning engineer or in your daily robotics research.
What you get by contributing
- Insight into how agents perform in your workflow
- Co-authorship on the RLE-Bench 2.0 manuscript
A runnable example, author guide, submission checklist, and local checks. Start with AUTHOR_GUIDE.md after extracting the ZIP.
01What we're looking for
Tasks should exercise the judgment, experimentation, and iteration involved in real robotics work. Start with a concrete deliverable: a controller, training recipe, perception system, or mechanical design.
- Challenging
- Success requires meaningful engineering decisions. Describe the expertise involved and what a simple baseline can achieve.
- Representative
- Use the tools, constraints, and tradeoffs of a real workflow. Make the starting assets and available compute explicit.
- Verifiable
- Measure the result against clear success criteria. Define the operating conditions, physical constraints, and scoring rules.
- Reproducible
- Provide a self-contained environment, pinned dependencies, and evidence that the same deliverable earns a consistent score.
Turn an idea into an engineering task
These examples illustrate how to make a task concrete across the benchmark's four workflows.
Interactive control
Complete a manipulation goal
Give the agent observations and a bounded interaction budget. Evaluate completion across held-out physical conditions.
Policy development
Develop a motion-tracking policy
Specify training assets and a compute budget. Measure tracking error and falls on evaluation motions.
Perception and estimation
Estimate object poses
Provide sensor inputs and coordinate conventions. Score pose accuracy against private ground truth.
Mechanical design
Design a stable mobile base
State payload, footprint, and mass limits. Test stability and performance in an independent simulation.
Difficulty should come from the engineering problem. Missing instructions, unavailable data, or fragile setup make a task harder to evaluate.
02Make success measurable
Decide how you will verify success before building the environment. Choose the approach that matches the deliverable.
The agent gets the prompt and public environment. Private cases and scoring code belong in the separate verifier. Compute rewards from trusted measurements; an agent's reported score is not evidence of success.
03Prepare your task
Download and extract the template. You can develop it as a standalone Harbor task without cloning RLE-Bench or making an advance proposal.
AUTHOR_GUIDE.md- The complete authoring workflow, evaluation requirements, and validation guidance.
SUBMISSION.md- Your task summary, scoring rationale, resource requirements, and validation evidence.
task/instruction.md- The agent's brief: goal, inputs, output paths and formats, constraints, and public scoring contract.
task/task.toml- Task identity and version, artifact handoff, timeouts, and agent and verifier resources.
task/environment/- The agent image's Dockerfile, tools, and public assets. Include only what the evaluated agent may access.
task/tests/- The separate verifier image, private harness, and scoring entry point.
task/solution/- An optional reference solution for Oracle runs. If omitted, document other reproducible evidence of solvability.
author_checks/- Local regression checks for the grader. Keep them outside the runtime images.
The two-link arm example is deliberately simple. Replace its prompt, environment, metrics, reference solution, and checks with your own engineering task before packaging a contribution.
Keep both runtime environments offline by default, install dependencies at build time, and document asset provenance and licenses. The kit's author guide covers larger datasets, GPU requirements, and requested network exceptions.
04Validate locally
The supplied kit targets Python 3.12, Docker, and Harbor 0.21.0. Install Harbor in a dedicated host environment. Run these commands from the extracted kit directory to check the unmodified example.
python3 -m unittest discover -s author_checks -v
harbor run -p task -a oracle --jobs-dir /tmp/rlebench-template-oracle
harbor run -p task -a nop --jobs-dir /tmp/rlebench-template-nop
For the included example, the expected rewards are Oracle: 1.0 and Nop: 0.0. Inspect each run's result.json and verifier/reward.json; a completed command alone does not establish a valid result.
Use a fresh jobs directory for each independent check. Initial image pulls need network access. An infrastructure error does not count as an expected Nop failure.
For your own task, record evidence of:
- Solvability: a reference result under the declared resource limits, with expected and actual reward ranges.
- Correct scoring: independent metric checks, boundary cases, and repeated evaluation of the same artifact.
- Robustness: missing or malformed outputs and task-specific attempts to bypass the metrics.
- Isolation: private evaluation assets and reference solutions are unavailable to the ordinary agent.
- End-to-end execution: clean image builds, artifact transfer to the separate verifier, and a real agent trial when feasible.
Update the example checks and expected scores for your task, then record commands, results, limitations, and any checks not run in SUBMISSION.md. If you omit solution/, skip the Oracle command and document your alternative solvability evidence.
05Package your contribution
- Complete
SUBMISSION.md, including authors, engineering value, scoring, dependencies, and validation evidence. - Include the full task, build inputs, author checks, licenses, and compact evidence files. Remove credentials, caches, virtual environments, and run outputs.
- Create the archive from the kit directory. Its root should contain
task/task.tomland the author documents, without another enclosing folder.
python3 -m zipfile -c /tmp/my-task.zip \
README.md AUTHOR_GUIDE.md SUBMISSION.md task author_checks
Extract the ZIP into a fresh directory and repeat the local checks. Include any additional license or evidence files in the archive; the command above packages the starter's standard files.
Task selection guidance inspired by Agents' Last Exam. Packaging and validation follow the RLE-Bench authoring kit.