Status: preparing the first repeated cross-provider release; no model ranking established. MapleBench runs a real game client through ordinary login, bounded model-generated keyboard programs, logout and verified persisted XP. The next release evaluates two OpenAI and two Claude models on the same Hero fixture, four attempts each, with frozen gameplay knowledge and a declared 17-skill Hero toolkit.
The release will include original gameplay recordings, held-key visualization, and a timeline of model calls and input delivery. Results are a descriptive pilot: signed session net XP, all planned attempts, spread and evidence-completion counts. The expanded native ten-skill qualification has passed; the 16 scored attempts have not started. See the release requirements, current status and continuation handoff. The public site currently contains earlier pilot evidence, which is separate from this candidate.
MapleBench provides structured observations; this is not a vision-only benchmark. The roadmap and benchmark design cover longer horizons, class breadth and future authoritative XP-window scoring.
Moving to another computer? See the laptop handoff to reuse the existing runner without transferring its assets or credentials.
The earlier four-model server-bot batches and replay renderer remain available. Their scores and rendering provenance are separate from the new full-client recordings. See scenarios and replay provenance.
The production readiness criteria track durable isolated attempts, native save receipts, exact recording review and verified publication evidence. Repeated trials, balanced model order and longer operational acceptance remain necessary before a dependable public ranking. The new finite experiment coordinator implements predeclared repeated plans, explicit resume without API replay, and complete-plan reports. Its single-entry live acceptance passed; repeated-model and unattended operation still require acceptance.
MapleBench is an experimental benchmark for evaluating coding agents in a persistent MapleStory-like game environment, beginning with simple XP optimization and progressing toward multi-agent party-quest coordination.
The intended world implementation is a MapleStory v83-compatible open-source server such as Cosmic. The benchmark framework itself contains no Nexon game assets or WZ data.
These are proposed directions, not completed full-client protocols. The current roadmap gates longer tasks on repeatable operation and the native scoring evidence each metric requires.
- Maximize XP — give an agent a standardized character and 10 minutes; score total server-authoritative XP gained.
- Maximize XP rate — score peak sustained XP/min over a rolling 60-second window.
- Multi-agent party quests — multiple agents coordinate to complete Kerning PQ under controlled communication topologies.
Every benchmark run should also produce a gameplay recording suitable for inspection and demos.
Agents should control a real in-world character through a narrow SDK, not mutate server state.
coding agent -> MapleBench SDK/MCP -> Cosmic adapter -> Cosmic server
| |
v v
events observer
| |
verifier MP4
The upstream Cosmic bot fork is useful because it already implements real-character movement/navigation, combat, inventory, skills and party behavior. MapleBench should reuse those execution primitives while withholding its built-in autonomous policies (grind, auto-quest, etc.) from evaluated agents.
src/protocol.ts— action, observation, task, and event contracts.src/sdk.ts— thin TypeScript agent SDK + HTTP transport.src/scoring.ts— total-XP and rolling XP-rate scoring.tasks/— initial XP and XP-rate task specs/prompts.docs/COSMIC_INTEGRATION.md— server adapter architecture.docs/COSMIC_BRIDGE_V0.md— concrete Java control-plane overlay + first live smoke test.docs/RECORDING.md— gameplay video pipeline.docs/MULTIAGENT_KPQ.md— first multi-agent research design.
Requires Node 22+. The pinned TypeScript compiler is installed with npm ci.
npm ci
npm test
npm run score:demoThe demo command scores a tiny example server event stream. It is deliberately independent of Cosmic so we can lock the benchmark contract before wiring the game server.
This historical scaffold checklist describes the earlier adapter. Use the full-client roadmap for current release priorities.
- Define server-authoritative episode/event schema.
- Implement total XP and rolling XP-rate scorers.
- Define narrow TypeScript SDK contract.
- Identify concrete Cosmic movement/combat integration methods.
- Pin Cosmic bot fork + Maplewright commits and automate checkout.
- Build a zero-setup live viewer and end-to-end mock SDK plumbing.
- Prepare
observe+move_to+ requested-attack Cosmic Java bridge source. - Prepare authoritative XP hook at Cosmic's real EXP mutation point.
- Compile/boot the patched full Cosmic checkout on a machine with upstream/network access.
- Run one real end-to-end
maximize-xp-10mepisode. - Attach Maplewright observer client and emit
run.mp4. - Wrap SDK in an
execute_codeMCP tool / Harbor task. - Generalize harness to four characters.
- Implement Kerning PQ evaluation.
For the seeded Henesys combat fixture and replay commands, see the server demo guide. This short baseline experiment is separate from a standardized ten-minute benchmark episode.
scripts/run-openai-queue.py runs a bounded, serialized OpenAI Responses API
batch against the dedicated disposable server/database. Each model gets the same
reset and prompt. Model IDs returned by the API, chosen actions, latency, usage,
observations, and scores stay in ignored run directories. Credentials come only
from the runtime environment or a private runtime file.
- Cosmic: https://github.com/P0nk/Cosmic
- Cosmic bot fork: https://github.com/NDBellisario/cosmic
- Maplewright: https://github.com/Sheilem/maplewright
- RuneBench: https://github.com/MaxBittker/runebench
This repository should only contain original benchmark/framework code unless otherwise clearly marked. MapleStory names, game data, WZ files, art, audio and other proprietary assets are not distributed here. Any Cosmic-derived server patches must preserve the applicable upstream license.
The full benchmark plumbing can be exercised before Cosmic/Maplewright are installed:
npm run demo:liveThen open http://127.0.0.1:8787. A tiny mock control backend on port 8790 accepts the same MapleSDK HTTP contract intended for Cosmic, a demo agent issues movement/attack/skill actions, the backend appends authoritative JSONL events, and the viewer updates from that file over SSE.
This mock exists only to validate benchmark plumbing. It is not a game simulator and is never used for benchmark scores.
Submit configs/smoke-20.json to the persistent queue to run four OpenAI API models
five times, with automatic scoring, replay rendering and a local results gallery.
The worker resumes interrupted batches and keeps each attempt. See
automated batches, programmable control,
and scenario presets.