Six AI decision models play Pac-Man, ranked live
Opper AI is an AI gateway platform that routes requests across 700+ models from 50+ providers with pay-as-you-go pricing. The page shown centers on jevman, an open-source Pac-Man benchmark built by Opper in which six AI decision models played 100 games each in real time against the classic arcade ghosts, with a public leaderboard, replays of every game, and the ability for anyone to play against the models or submit their own model to the benchmark.
At every junction the game sends the maze as JSON (including pellets and ghosts as state) and asks the model for a direction; the model returns a probability per direction and Pac-Man takes its pick. Answers must arrive within 2 seconds or a simple backup rule takes over (counted as backup moves).
Each model played 100 games against the classic scripted ghosts, playing until it lost all three lives, with games capped at 5 minutes (the longest lasted 2 minutes 24 seconds). Ranking uses mean score with a 95% margin of error (±2 standard errors); models within each other's margin are tied (=1).
Every game is recorded and replays exactly, so any submitted result can be checked. Submitted models are run under the leaderboard's rules and CI replays every game to verify its score before it joins the leaderboard as self-reported.
Any model — hosted, a fine-tune, or one running on your laptop — can play. The repo includes a 34-line example endpoint to start from, and one command runs the games under the leaderboard's rules.
The site lets you watch the models' games or play Pac-Man against them directly in the browser, including a game with the jev 1.13 model with classic ghosts.
Beyond the benchmark, Opper offers a Gateway and Control Plane with 700+ models from 50+ providers, routing and fallbacks, prompt caching, BYOK, spend limits, LLM checks (flag, redact, block, score), zero data retention (ZDR) routes, and a management API for projects, keys, members and routes from C
At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from.
One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes.
CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported. If your model is on Opper or another public API, you can open an issue and Opper will run it themselves.
Opper AI uses pay-as-you-go pricing with three plans: Gateway (3% platform fee), Control Plane (5.5% platform fee), and Enterprise (custom). Sign-up needs no credit card and free models work in the playground and API; add a card to use premium models with no minimum spend. Payment is via Stripe (credits), with invoicing for Enterprise.
3% platform fee
700+ models from 50+ providers; ZDR with usage metadata only stored by default; spend limits via prepaid balance with auto top-up; any model chosen per request; LLM checks (flag, redact, block, score); per-request fallbacks; provider-native
5.5% platform fee
Everything in Gateway plus: ZDR by default or full traces retained 1–30 days; org-wide ZDR-only enforcement via rule; model sub-processors limited to your allowlist; org budgets with caps per project, role, user or key; org allowlist with p
Custom
Custom platform fees and terms; data kept in your own cloud or custom terms; ZDR routes needing signed agreements; roles from your identity provider; per role, user or key model access; pool order under your SLA; BYOK using your cloud commi
jevman is an open-source Pac-Man benchmark for AI decision models built by Opper AI. Six models played 100 games each in real time against the classic arcade ghosts, and the results are ranked on a public leaderboard.
At every junction the game sends the maze as JSON — with pellets and ghosts as state — and asks the model one question. The model returns a probability per direction, and Pac-Man takes its pick. An answer taking longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.
Real time, 2 seconds per answer, the classic scripted ghosts, three lives, and games capped at 5 minutes so every run ends. In practice none of the games came close to the cap — the longest lasted 2 minutes 24 seconds.
Ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.
Yes. jevman is open source (AGPL-3.0) and any model behind an HTTP endpoint can play — a hosted model, a fine-tune, or one running on your laptop. The repo includes a 34-line example endpoint, one command runs the benchmark, and you submit via pull request. CI replays every game and checks its score before your model joins the leaderboard as self-reported. If your model is on Opper or another public API, you can open an issue and Opper will run it for you.
Every game is recorded and replays exactly, so any result can be checked. Submitted games are re-run by CI to verify the score.
Yes — you can watch the models' games or play Pac-Man against them directly, choosing from the six models with the classic ghosts.
Opper uses pay-as-you-go pricing with platform fees: 3% on Gateway, 5.5% on Control Plane, and custom pricing for Enterprise. You pay for the underlying models on top. Sign-up needs no credit card and free models work in the playground and API; add a card to use premium models with no minimum.