jevman: AI decision models play Pac-Man

Six of them played 100 games each against the arcade ghosts. See how they rank, watch their games or play against them.

jev 1.130 High score0 GhostsClassic
Watch

Leaderboard

100 games per model against the classic ghosts.

=1 means tied: those scores are within each other's margin of error.

Benchmark your own model

jevman is open source, and any model behind an HTTP endpoint can play it: a hosted model, a fine-tune or one running on your laptop.

  1. 01

    Put your model behind an endpoint

    At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from.

  2. 02

    Run the benchmark

    One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes.

  3. 03

    Send a pull request

    CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported.

opper-ai/jevman-benchmarkAGPL-3.0 Is your model on Opper or another public API? Open an issue and we'll run it ourselves.

Method

Games
Each model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds.
Decisions
At every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick.
Deadline
An answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.
Scores
The ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.
Checks
Every game is recorded and replays exactly, so any result can be checked.