jevman: AI decision models play Pac-Man
Six of them played 100 games each against the arcade ghosts. See how they rank, watch their games or play against them.
Leaderboard
100 games per model against the classic ghosts.
=1 means tied: those scores are within each other's margin of error.
Benchmark your own model
jevman is open source, and any model behind an HTTP endpoint can play it: a hosted model, a fine-tune or one running on your laptop.
-
01
Put your model behind an endpoint
At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from.
-
02
Run the benchmark
One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes.
-
03
Send a pull request
CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported.
Method
- Games
- Each model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds.
- Decisions
- At every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick.
- Deadline
- An answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.
- Scores
- The ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.
- Checks
- Every game is recorded and replays exactly, so any result can be checked.