Chinese Checkers JumpStar

Chinese Checkers AI

JumpStar is the world's strongest AI engine for Chinese Checkers

Now on iPhone JumpStar Chinese Checkers Play the strongest AI offline, solve daily puzzles, get hints, and save replays.

Origin

I built JumpStar for my mom. The project began on May 15, 2026 in Indiana when I was helping my parents move out of our childhood home. While packing their things, we found our old Chinese Checkers board and started reminiscing about how fun it was to play as a family.

After my brother and I left for college, my mom had tried to find Chinese Checkers iPhone apps so she could continue to play by herself, but none of them were any good. Given the successes of AlphaGo and Stockfish, I thought that, surely, there would be many strong open source AI engines for Chinese Checkers - but there aren't. Despite its popularity and elegance, Chinese Checkers had received surprisingly little attention from the AI community. That gap is what led me to build JumpStar.

My goal was to build the first superhuman AI engine for Chinese Checkers. A strong computer player would make it possible for my mom to keep playing even when my brother and I were not nearby. A superhuman AI engine may even uncover beautiful new patterns and strategies of the game that weren't yet widely known.

That is also why the project became public-facing. A strong private engine is interesting, but it is difficult for anyone else to evaluate or improve. The public Chinese Checkers AI benchmark turns the work into something inspectable: a ruleset, a protocol, a position suite, baselines, logs, and a named JumpStar model that other systems can challenge.

Why Chinese Checkers AI is interesting

Chess and Go have decades of public AI work behind them: engines, benchmarks, rating lists, and shared game records. Chinese Checkers has very little of that, even though it's the kind of game AI is good at studying. There's no hidden information, no luck, and a clear way to win.

It also gets complicated fast. In JumpStar's two-player rules there are 14 legal first moves, 196 ways the first two moves can go, and 4,760 by the third. Later on, chained jumps and crowded lanes make good moves much harder to see than the rules suggest.

What JumpStar is

JumpStar is a self-play-trained Chinese Checkers engine built on a compact C++20 rules/search core, a neural policy/value model, and Monte Carlo tree search. Its strongest public checkpoint, JumpStar_60, is the current CCERL-2P10-v2 benchmark champion. Every JumpStar model was trained on a MacBook M4, not on the large GPU clusters behind the famous chess and Go engines.

The goal is superhuman play, and the claim is meant to be tested: JumpStar looks superhuman or close to it under the two-player rules, and CCERL lets any future engine challenge that.

The system

JumpStar is a full engine and benchmark, not just a trained model.

  • Rules and search. A compact model of the 121-hole board with strict two-player rules, fast move generation for steps and multi-jump chains, and a match runner. The rules are explicit on purpose, so an engine can't win by camping in a triangle or exploiting a quirk of the referee.
  • Self-play training. JumpStar plays itself, and those games train a policy/value model that guides the next round of search. JumpStar_60 is a 512x4 network with about 10.36M parameters, trained on 1.8M reanalyzed positions.
  • Public benchmarking. There was no public engine ladder for Chinese Checkers, so I built CCERL-2P10-v2 alongside JumpStar. It defines the rules, the engine protocol, a fixed set of starting positions, baseline engines, and how ratings are computed, and it publishes every game log so results can be checked.

How JumpStar Plays Chinese Checkers

JumpStar often doesn't just race. It slows the opponent down, keeps useful blocking marbles in place, and builds compact triangle shapes that make the middle of the board hard to jump through.

Because pieces are never captured, a defensive shape can last for many turns. It can take away a jump route, force the opponent to go around, and buy time for JumpStar's own marbles to get home.

Winning Opening Pattern

Winning opening replay An eight-move opening sequence with Player 2's final jump highlighted.
# Player Kind Move
JumpStar defensive triangle shape A production game position showing JumpStar forming a central shape that discourages jumps through the middle and pushes opposing pieces outside.
Close the middle. JumpStar often forms two compact triangle-like fronts that make the central lanes awkward to jump through, pushing opposing pieces toward the outside.
JumpStar long endgame chain A production game position showing JumpStar spacing pieces into a long jump chain toward the opposite goal.
Convert in chains. Later, those same spacing instincts create long ladders that bring trailing pieces home efficiently.

The search and training loop

At a high level, JumpStar follows the AlphaZero pattern: self-play produces MCTS visit targets, those targets train a policy/value model, and the stronger model guides the next round of search.

self-play -> MCTS visit targets -> policy/value training -> stronger search -> self-play

The practical strength came from making that loop cheap enough to run repeatedly: native C++ self-play workers, efficient move generation, batched leaf evaluation, transpositions, subtree reuse, compact records, and release builds tuned for local Apple hardware.

A major part of that compression came from using Codex with GPT-5.5 as an implementation and research partner. Codex helped inspect the codebase, profile bottlenecks, rewrite hot paths, analyze training logs, package benchmark runs, and keep multiple experiment threads moving at once. The most important gains were practical: large speed-ups in self-play and evaluation, memory reductions of roughly 98% in the training path so the work could fit on my local machine, and enough automation to keep improving the engine while I was also doing my full-time job running Edia as CEO.

The project moved unusually fast because the loop was not only self-play for the model; it was also an iterative engineering loop for the system around it. Codex made it possible to run an experiment, inspect the failure mode, optimize the code, rerun the benchmark, summarize the result, and turn the next question into a concrete patch. That feedback cycle compressed work that might otherwise have taken months of infrastructure, tooling, frontend, benchmark, and writeup time into about one week from first board discovery to public launch.

Local Codex accounting gives a rough sense of the scale of that collaboration. Across the project threads visible in the local Codex state database, recorded tokens_used totaled approximately 559.9M tokens. Those numbers include context, tool output, cached-context effects, reasoning/output accounting, and overlapping work threads, so they should be read as a process metric rather than a scientific measurement of compute.

CCERL and the public benchmark

Chinese Checkers does not have a mature public engine ladder like chess and Go, making it difficult to evaluate engine strength and progress. CCERL fills that void. CCERL is the first public benchmark for evaluating Chinese Checkers AI engine strength. Based on similar concepts from chess and Go, CCERL provides fixed rules, audited starting positions, paired side swaps, downloadable game logs, and public baselines.

Based on CCERL benchmarks, JumpStar appears to be the world's strongest publicly available AI engine for Chinese Checkers by a large margin.

I believe that JumpStar has reached near-superhuman Chinese Checkers strength under the two-player rule profile, using local MacBook-scale compute. More importantly, CCERL gives all future contributors a target: modify JumpStar, build a new model, port an existing engine more faithfully, add better positions, or submit a challenger under the same referee.

Future Work

JumpStar is not meant to be just a private bot. The project is trying to become four things:

  • A strong public AI model. JumpStar is open license, and this website is free. Anyone can play against JumpStar or try to improve it themselves. JumpStar was trained entirely on a MacBook M4. More compute will surely improve the performance.
  • Public benchmark infrastructure. Chess and Go have communities where engines can be tested, compared, and improved. CCERL brings that same infrastructure to Chinese Checkers: clear rules, reproducible matches, public logs, and a path for other people to build their own models.
  • Research into the game itself. Strong AI changes what we can see. I want to understand what good play looks like, how openings evolve, why crowded boards behave the way they do, and what happens in three, four, and six-player games when the middle of the board becomes a traffic jam of possibility.
  • Multiplayer variants. The two-player model is only the beginning. The wilder questions start when the board gets more crowded: three players, four players, six players, shifting alliances, blocked paths, and strange emergent openings. I want to find out what strong play looks like there too and on boards that are far larger than standard. The complexities and patterns that emerge will be mesmerizing.

I hope JumpStar encourages more people to think about Chinese Checkers, play it, study it, and build around it.

The code is on GitHub, and you can watch JumpStar's games move by move.