Creating a (hopefully) competitive bot using Claude Code.
-
Can you please just use names of games which can be found in TripleA, adding information on what you changed if anything? You mean that you are playing "World War II v5 1942 SE TR" with a bid for Russians or also some other changes?
-
@husky81
What is your approach doing it? -
@xxXEddieXxx well there's lots of ideas that I have and it's early stages at the moment.
The thing is entirely vibe coded, using Claude code Opus. I have the 200$/year plan so no Fable access.
I am not a software developer at all, so everything software related in entirely up to ai. I have read 0 lines of code so far, and couldn't understand most of them if I did. I studied math in undergrad, so i am used to logical thinking though.
I (meaning Opus) have built a python engine that implements the map and game's ruleset. It is also bridged to a local copy of triple A. Claude can simulate tournaments between my bots in the python engine, or use the bridge to triple a to match my bots vs the pre-exising bots in the game.
I have told Claude that I want my bot to be constructed in a modular fashion, so that modules can be swapped out or changed individually, changes can them be A/B tested by running tournaments pitting various versions of my bots against each other. I get Claude to track various statistics form these tournaments. For example, when I was building the naval logic, I had tournaments where i tracked number of amphibious assaults, ships lost, transports lost, units delivered, and a few other things. trying to get the logic right so that transports get defended, enemy fleets get zoned out and attacked, and lots of units get delivered to combat theatres.
The most important component i have not is the threat map. It uses Combat power, which is currently dice pips+2.6 times hit points, maintains an updated defense value for each territory on the board, and also maintains an overlapping projected combat power of each nation onto each territory, tracking the maximum combat power that can be brought to bear against any target, or potential target.
I started off with just a basic attack planner that takes every attack over 70% given that it doesn't leave the capital vulnerable . My next step is to try and build a competent attack planner that take other factors like stack survival into consideration in its planning. My main ideas are to go for a combination of positive TUV exchange, and restriction of the enemy's projected combat power. so for example the bot will try to block tank blitzes, because doing so reduces the enemy's projected combat power, restricting the tanks attack to fewer territories.
I will keep working and iterating through conversations with Opus. I do most of the things on my phone using Cowork. Claude code runs on my desktop building and testing the bot modules. I ask it questions, and give it information and instructions by talking into my phone.
So far i haven't actually watched any games that it plays, all of them are simulated, and only statistics are reported. This is quite fast and efficient, but once the core of my bot is built I'll see if i can look at some games to see if i detect problems that escape the q and a format i currently use. I'll at least need to finish the ground attack module, purchasing, and give it the ability to think more than one turn ahead before it's worth looking at what the bot is actually doing.
-
@husky81 hey that sounds ambitious. I don't know how deep you are in terms of how these ai models work. But I had some ideas regarding an AI aswell and made a little map of it :
│ Strategic Planner │ │ LLM / learned policy │ │ strategic objectives │ "Pressure Moscow" "Hold India" "Build naval superiority" "Trade Ukraine cheaply" │ ▼ │ Candidate Plan Generator │ │ deterministic algorithms │ │ 20-100 plans │ ▼ │ Simulation / Search │ │ MCTS / rollouts │ │ expected values │ ▼ │ Value Function │ │ NN + heuristics │ │ ▼ ACTIONThere is also a paper on something that might help.
https://repository.gatech.edu/server/api/core/bitstreams/a1b58887-d7df-48ee-a412-b13b3e236983/contentIn terms of models opus 5.5 should be enough.
-
@husky81
TripleA AI uses some of these values, its just a list to be used as a guide. It is more complicated than listed, but it gives you a start point for your AI.2x TT/SZ PU value +
10x if isFactory (personally should consider canProduceXUnits value) +
11x if TT is also MyCapital +
5x if it is isEnemyOrAlliedCapital) (personally Allied Capital should be 6x)Victory Centre does not appear to be taken into the AI calucation, but in the xml you can have
"victoryCity" value="1" so this should be a multiplier.Keep up the good work!
-
@xxXEddieXxx I spent most of the day building logic for combat moves. once I finish purchases and non combat moves i'll have an alpha version of my bot. To clarify, it had purchases and non combat moves before, but just something simple that opus cooked up, like everything to the nearest front, buy inf and arty plus transports if not mainland factories.
really simple stuff. I'm trying to build more thoughtful versions of those modules. Once that's done I'll pit it against triple a's fast ai and see how it does. I did this once before, beat easy, but lost badly to fast and hard. For sake of time i won't play hard again until I can beat fast. I have no idea how long that will take, but this vibe coding allows such quick and easy iteration that it's incredible. The most time consuming aspect is running simulations, as I only have 1 old ish pc with an 8 core cpu doing the work, i asked it about trying to use my GPU, But Claude said it wasnt't powerful enough to be useful, and it would require rewriting alot of the game software.
When i'm done with the modules i talked about above I'll post tournament results and a Claude writeup about how it works. Eddie, I'm not sure your schematic would be compatible with what i've built so far. But after the build I'm working on now, i'll feed it into claude code and see what it thinks. I may end up adding features from your outline, or building a competing bot around your model, recusing existing components where I can.
Currently my GitHub repo is set to private, but I am willing to open it up and share it once I'm somewhat happy with where the project is. I would want it to be a personal project, so wouldn't be looking for collaborators, at least at the moment, but of course people could fork it if they wanted, and no doubt someone would advance leaps and bounds beyond me.
-
@TheDog Thanks! I am trying to get claude to work things out from first principles to the extent possible. It's really amazing what can be done and the conversations that can be had. iteration is also extremely fast. It can write code in seconds that I couldn't in months, test and refine, all just by asking it questions and providing human information.
-
World War II v5 1942 SE TR
Sorry, I didn't see this. Yes, my bot plays 'World War II v5 1942 SE TR'
For development and testing I play in a Python engine that Claude built for that purpose. In those cases I played without bids. The only major rule difference is no landing on allied carriers, since that's not in the Beamdog game. Claude also built a bridge that allows my bot to play the bots that ship with Triple A. I've not changed the rules in Triple A, just my own engine and my bot's logic, so for the games played through the bridge, the Triple A AI could use allied carriers, but mine never will.
The results above are from a tournament I simulated between an early version of my bot and Triple A's Bots. For that tournament I used a bid of Russia +45!!, because based on simulations, that's what is required to equalize the Axis and the Allies on that map when played by the standard Triple A bots.
-
@husky81
Does Claude know the importance of a victoryCity and the loss of a Capital ? -
@husky81 Here is the thread of someone else who has tried, https://forums.triplea-game.org/topic/4240/game-engine-rules-ai-training
You have actually accomplished something by writing a bot and having it play against the Hard AI in TripleA. Also, your goal of creating a heuristic to evaluate positions is an interesting approach as opposed to those who think that LLM's have magical powers and they can solve anything with a prompt.
I suggest trying the other AI's in TripleA to get a greater variety of games. However, I suspect that Claude will not do a good job at doing positional evaluations. There is just too much going on. Note the Hard/Fast AI are effectively using TUV (total unit value). I suspect that relative TUV, relative income and objectives held would provide a fairly good statistical model.
You should check out the MiniMap scenario. It is the smallest non-trivial scenario and should allow you test your ideas more quickly.
I would be interested in how you got the Hard AI to play itself repetitively. It would be useful to determine what handicaps could be used to balance various scenarios.
-
f a victoryCity and the loss
Yes and no. Claude knows about the win conditions 9/10 VC's for Axis/Allies, and about the value of capitals.
My bot, (built by Claude, but not Claude) runs capital safety checks each turn. Its positional evaluator can see VC snipes, but it is not yet capable of making multi turn plans to either get a VC, or to block one. it only had limited capabilities for multi turn planning which is mainly done by leaving flags on the map that can be read on its next turn.
Running a full position search on multiple turns probably grows by something like n! in turn number, if not even worse. so that's probably impossible, I'll need a more clever way to get it to make future plans. I want to avoid hard coding as much as I can though.
-
Thanks,
I have tried playing Hard fast and easy ai's. My bot can beat easy, but lost to hard and fast.
I am building components and running self tests on my python game engine to try and get my bot to re-match Fast (I'll do hard later). It's quicker to iterate without bridging to triple a, so i think it's not worth doing again until my bot is actually decent. I am currently trying to build a proper logic for purchase and non combat move. Once I'm done those I'll try matching Triple A's fast again.
I asked Opus 5.5 to explain how it got Triple A to repeatedly self play: here is my prompt and its reply:
Me: a forum poster is asking me how I got you to setup the tournament where triple a's hard and fast ai played against each other, and our bot. could you write up a brief explanation of how you did that as well as the data that you collected. They are interested in doing it because they want to use self play to balance various scenarios. You could elaborate on what you might do with that objective in mind.
Opus 5.5 High:
"How the AI-vs-AI tournament was set up (TripleA 1942 2nd Edition)
Running TripleA headless. TripleA's own test suite already plays games between AIs without a GUI. I added one class next to those tests, Ww2v5SelfPlayRunner, about 470 lines in game-app/smoke-testing. It loads WW2v5_1942_2nd_TR.xml and assigns a player type per side: Hard (ProAi), Fast, Easy (WeakAi), Does Nothing, or our own bot. It then steps through the game delegate by delegate until someone wins or a round cap is hit (I used 25). Everything is configured through environment variables:
the map file;
the AI for each side;
the round cap;
the number of games;
an optional bid.A small Gradle task prints the test classpath. Each game then runs as its own JVM, launched with plain java -cp ... Ww2v5SelfPlayRunner, so several games run in parallel. The Hard AI already uses several cores, so running 4 games at once only gained 10–20%.
Bids. There were two ways to apply a bid:
set the map's "<power> bid" property and let the AI buy and place its own bid units;
load a fixed JSON of bid units, so every game starts from exactly the same position.The tournament used the fixed version: a 45 PU Russian bid, placed the way the Hard AI had chosen to place it.
Our bot. The bot is written in Python. A bridge AI on the Java side sends the game state to the Python process as JSON lines each time a decision is needed. Python answers with purchases, moves, casualty choices and retreats, and TripleA executes them. TripleA accepted all 15,732 of the bot's actions in the tournament, which doubles as a rules-parity check.
Recording. At the start of every power's turn the runner writes a snapshot as gzipped JSON lines, one file per game:
every territory's owner and units;
every power's bank and income.A header records the AIs and the bid, and a final line records the winner, the last round and the time taken. A Python driver (tournament.py) builds the schedule (every AI against every AI, 5 games per side), skips finished games so a run can be resumed, and runs 4 games at a time.
Data collected
Pilot, Hard vs Hard, no bid, 50 games: Axis 48, Allies 1, unfinished 1, with the median win in round 9. About 140 s per game.
Bid tuning, Hard vs Hard:
Russia bids of 12, 21 and 30 PUs: Axis still won 95–100%.
Russia 30 plus UK/US bids: Axis 80–85%.
Russia 40: roughly even (4 Axis, 3 Allies).
Evaluator dataset: 110 Hard vs Hard games at Russian bids of 40–50. A logistic-regression position evaluator trained on it predicts the eventual winner from any snapshot at 83% (log-loss 0.345, cross-validated by game).
Round robin, 120 games at the fixed 45 PU bid:
AI Points vs the others
Hard 89%
Fast 85%
Our bot (at the time) 51%
Easy 25%
Does Nothing 0%
Hard vs Fast head to head: 5 wins to 3, 2 unfinished.
At this bid, Axis 56 and Allies 57 across all games.
Speed: Hard vs Hard takes about 16 minutes per game; Fast vs Fast about 3.4 minutes, roughly 1/6 of the time at almost the same strength.Using self-play to balance scenarios: what I'd do
The biggest caveat: AI bias is not scenario balance. The Hard AI needed a 40–50 PU Russian bid to break even, while expert human play puts this game near even (an Allied bid of 3–6 PUs at most). The AI simply plays the Allies worse. A bid found by AI self-play measures how well the AI handles each side, not how balanced the scenario is. Use self-play to compare variants (scenario A vs scenario B under the same AIs), and calibrate the absolute number against human games where you can.
Check the AI understands the scenario. New units or rules the AI misjudges will skew results. Watch a few full games before trusting a batch.
Use Fast AI for volume and Hard for spot checks.
Plan the number of games. The standard error of a win rate near 50% is about 0.5/√n: roughly ±5 points at 100 games, ±2.5 at 400. Batches of 10 are much too noisy to tell balance apart.
Fit the even bid across all games rather than searching step by step. Play a spread of bids, fit P(Axis win) against the bid with logistic regression, and read off where it crosses 50%, with a confidence interval. I started with step-by-step adjustment and it wasted games.
Get more signal per game than win/loss:
Score unfinished games by TUV share or victory cities instead of discarding them.
Use a position evaluator's P(win) at a fixed round as a lower-variance measure. It needs care, because it inherits the AI's bias.
Record snapshots every turn, not just results. They show where and when a scenario tips: which territories swing, the round after which games are decided, and opening outcomes. That tells you what to change, not just which side is favoured.I didn't include the per-territory swing analysis or the evaluator's calibration tables; say if the poster would want either. The two purchase screens from earlier are still running."
Me: I only know what a small bit of that means, but it sounds pretty cool.
-
Can you combine scripted responses while adding some van dam uniqueness ?
It'd be a projrect. Start with the optimal 3-5 moves and there counters and then do the same from there. Can see that exploding quite quickly

But instead of just choose between these 3 and 5 choices add some claudeness to it as well.
Idk I have basically zero knowledge of how hard this is but it seems pretty hard

Keep it up though, we want you to succeed

-
@husky81 said
I am currently trying to build a proper logic for purchase and non combat move.
Have you/ClaudeBot seen TripleA AI logs?
In game
Debug> HardAI> Show Logs
Enable AI Logging =TICK
Log Depth=Finest
Log History To=99AFAIK this is only in memory and not a file.
It shows amongst other things how it categorizes its purchases and how it values the purchase of units and places them where needed.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login