OpenRA-Bench / README.md
yxc20098's picture
Land PR #15 surgical fixes (Windows safety + human-play hints)
a04d20a
|
Raw
History Blame Contribute Delete
4.59 kB

A newer version of the Gradio SDK is available: 6.20.0

Upgrade
metadata
title: OpenRA-Bench
emoji: 🎮
colorFrom: red
colorTo: blue
sdk: gradio
sdk_version: 5.12.0
app_file: app.py
pinned: true
license: gpl-3.0

OpenRA-Bench

Standardized benchmark and leaderboard for AI agents playing Red Alert through OpenRA-RL.

Features

  • Leaderboard: Ranked agent comparison with composite scoring
  • Filtering: By agent type (Scripted/LLM/RL) and opponent difficulty
  • Evaluation harness: Automated N-game benchmarking with metrics collection
  • OpenEnv rubrics: Composable scoring (win/loss, military efficiency, economy)
  • Replay verification: Replay files linked to leaderboard entries
  • Mission Player: Static game-like website for browsing, annotating, and reviewing scenarios
  • Bilingual: English and Chinese scenario instructions generated deterministically

Quick Start

View the leaderboard

pip install -r requirements.txt
python app.py
# Opens at http://localhost:7860

Run an evaluation

# Against the HuggingFace-hosted environment (no Docker needed)
python evaluate.py \
    --agent scripted \
    --agent-name "MyBot-v1" \
    --opponent Normal \
    --games 10 \
    --server https://openra-rl-openra-rl.hf.space

# Or against a local Docker server
python evaluate.py \
    --agent scripted \
    --agent-name "MyBot-v1" \
    --opponent Normal \
    --games 10 \
    --server http://localhost:8000

Submit results

Via CLI (recommended):

pip install openra-rl
openra-rl bench submit result.json
openra-rl bench submit result.json --replay game.orarep --agent-name "MyBot" --agent-url "https://github.com/user/mybot"

Results from openra-rl play are auto-submitted after each game.

Via PR:

  1. Fork this repo
  2. Run evaluation (appends to data/results.csv)
  3. Open a PR with your results

Agent identity

Customize your leaderboard entry:

Field Description
agent_name Display name (e.g. "DeathBot-9000")
agent_type Scripted, LLM, or RL
agent_url GitHub/project URL — renders as a clickable link on the leaderboard

Replay downloads

Entries submitted with a .orarep replay file show a download link in the Replay column. Replays are stored on the Space and served at /replays/<filename>.

API endpoints

The Gradio app exposes these API endpoints (Gradio 5+ SSE protocol):

Endpoint Description
submit Submit JSON results (no replay)
submit_with_replay Submit JSON + replay file
filter_leaderboard Query/filter leaderboard data

Mission Player (Static Site)

A game-like mission selection and annotation website in site/. No framework, no build step -- a single HTML file deployable to GitHub Pages.

For players / annotators

Open site/index.html via any HTTP server:

cd site && python3 -m http.server 8765
# Open http://localhost:8765/index.html

Workflow: browse scenario cards, pick a mission, read bilingual objectives (EN/ZH toggle), switch difficulty (easy/medium/hard), annotate the map with point/region tools, tag and add notes, mark complete, navigate to next mission, export annotations as JSON.

For maintainers

Generate or refresh static data after scenario changes:

python site/generate.py            # generate scenarios.json + map thumbnails
python site/generate.py --dry-run  # print counts without writing

Map thumbnails require the Rust engine wheel (openra_train). Without it, the site works with a placeholder map area; annotations still work on the placeholder.

Deploy by copying site/index.html and site/public/ to any static host.

See docs/IMPLEMENTATION_NOTES.md for full details.

Running tests

# Data pipeline + coverage invariant tests (Python)
python -m pytest tests/test_site.py tests/test_app.py -v

# E2E DOM interaction tests (Node.js + jsdom)
npm install   # first time only
node tests/test_site_e2e.mjs

Scoring

Component Weight Description
Win Rate 50% Games won / total games
Military Efficiency 25% Kill/death cost ratio (normalized)
Economy 25% Final asset value (normalized)

Links