TripleA Logo TripleA Forum
    • TripleA Website
    • Recent
    • Popular
    • Register
    • Login

    Creating a (hopefully) competitive bot using Claude Code.

    Scheduled Pinned Locked Moved AI
    25 Posts 7 Posters 115 Views 6 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • H
      husky81 @TheDog
      last edited by

      @TheDog Thanks! I am trying to get claude to work things out from first principles to the extent possible. It's really amazing what can be done and the conversations that can be had. iteration is also extremely fast. It can write code in seconds that I couldn't in months, test and refine, all just by asking it questions and providing human information.

      1 Reply Last reply
      Reply Quote 0
      • H
        husky81 @Cernel
        last edited by

        @Cernel said:

        World War II v5 1942 SE TR

        Sorry, I didn't see this. Yes, my bot plays 'World War II v5 1942 SE TR'

        For development and testing I play in a Python engine that Claude built for that purpose. In those cases I played without bids. The only major rule difference is no landing on allied carriers, since that's not in the Beamdog game. Claude also built a bridge that allows my bot to play the bots that ship with Triple A. I've not changed the rules in Triple A, just my own engine and my bot's logic, so for the games played through the bridge, the Triple A AI could use allied carriers, but mine never will.

        The results above are from a tournament I simulated between an early version of my bot and Triple A's Bots. For that tournament I used a bid of Russia +45!!, because based on simulations, that's what is required to equalize the Axis and the Allies on that map when played by the standard Triple A bots.

        TheDogT 1 Reply Last reply
        Reply Quote 0
        • TheDogT
          TheDog @husky81
          last edited by

          @husky81
          Does Claude know the importance of a victoryCity and the loss of a Capital ?

          https://forums.triplea-game.org/tags/thedog
          https://forums.triplea-game.org/topic/3741/curated-best-top-maps-triplea-guides

          H 1 Reply Last reply
          Reply Quote 0
          • RogerCooperR
            RogerCooper @husky81
            last edited by

            @husky81 Here is the thread of someone else who has tried, https://forums.triplea-game.org/topic/4240/game-engine-rules-ai-training

            You have actually accomplished something by writing a bot and having it play against the Hard AI in TripleA. Also, your goal of creating a heuristic to evaluate positions is an interesting approach as opposed to those who think that LLM's have magical powers and they can solve anything with a prompt.

            I suggest trying the other AI's in TripleA to get a greater variety of games. However, I suspect that Claude will not do a good job at doing positional evaluations. There is just too much going on. Note the Hard/Fast AI are effectively using TUV (total unit value). I suspect that relative TUV, relative income and objectives held would provide a fairly good statistical model.

            You should check out the MiniMap scenario. It is the smallest non-trivial scenario and should allow you test your ideas more quickly.

            I would be interested in how you got the Hard AI to play itself repetitively. It would be useful to determine what handicaps could be used to balance various scenarios.

            H 1 Reply Last reply
            Reply Quote 3
            • H
              husky81 @TheDog
              last edited by

              @TheDog said:

              f a victoryCity and the loss

              Yes and no. Claude knows about the win conditions 9/10 VC's for Axis/Allies, and about the value of capitals.

              My bot, (built by Claude, but not Claude) runs capital safety checks each turn. Its positional evaluator can see VC snipes, but it is not yet capable of making multi turn plans to either get a VC, or to block one. it only had limited capabilities for multi turn planning which is mainly done by leaving flags on the map that can be read on its next turn.

              Running a full position search on multiple turns probably grows by something like n! in turn number, if not even worse. so that's probably impossible, I'll need a more clever way to get it to make future plans. I want to avoid hard coding as much as I can though.

              B 1 Reply Last reply
              Reply Quote 0
              • H
                husky81 @RogerCooper
                last edited by

                @RogerCooper

                Thanks,

                I have tried playing Hard fast and easy ai's. My bot can beat easy, but lost to hard and fast.

                I am building components and running self tests on my python game engine to try and get my bot to re-match Fast (I'll do hard later). It's quicker to iterate without bridging to triple a, so i think it's not worth doing again until my bot is actually decent. I am currently trying to build a proper logic for purchase and non combat move. Once I'm done those I'll try matching Triple A's fast again.

                I asked Opus 5.5 to explain how it got Triple A to repeatedly self play: here is my prompt and its reply:

                Me: a forum poster is asking me how I got you to setup the tournament where triple a's hard and fast ai played against each other, and our bot. could you write up a brief explanation of how you did that as well as the data that you collected. They are interested in doing it because they want to use self play to balance various scenarios. You could elaborate on what you might do with that objective in mind.

                Opus 5.5 High:

                "How the AI-vs-AI tournament was set up (TripleA 1942 2nd Edition)

                Running TripleA headless. TripleA's own test suite already plays games between AIs without a GUI. I added one class next to those tests, Ww2v5SelfPlayRunner, about 470 lines in game-app/smoke-testing. It loads WW2v5_1942_2nd_TR.xml and assigns a player type per side: Hard (ProAi), Fast, Easy (WeakAi), Does Nothing, or our own bot. It then steps through the game delegate by delegate until someone wins or a round cap is hit (I used 25). Everything is configured through environment variables:

                the map file;
                the AI for each side;
                the round cap;
                the number of games;
                an optional bid.

                A small Gradle task prints the test classpath. Each game then runs as its own JVM, launched with plain java -cp ... Ww2v5SelfPlayRunner, so several games run in parallel. The Hard AI already uses several cores, so running 4 games at once only gained 10–20%.

                Bids. There were two ways to apply a bid:

                set the map's "<power> bid" property and let the AI buy and place its own bid units;
                load a fixed JSON of bid units, so every game starts from exactly the same position.

                The tournament used the fixed version: a 45 PU Russian bid, placed the way the Hard AI had chosen to place it.

                Our bot. The bot is written in Python. A bridge AI on the Java side sends the game state to the Python process as JSON lines each time a decision is needed. Python answers with purchases, moves, casualty choices and retreats, and TripleA executes them. TripleA accepted all 15,732 of the bot's actions in the tournament, which doubles as a rules-parity check.

                Recording. At the start of every power's turn the runner writes a snapshot as gzipped JSON lines, one file per game:

                every territory's owner and units;
                every power's bank and income.

                A header records the AIs and the bid, and a final line records the winner, the last round and the time taken. A Python driver (tournament.py) builds the schedule (every AI against every AI, 5 games per side), skips finished games so a run can be resumed, and runs 4 games at a time.

                Data collected

                Pilot, Hard vs Hard, no bid, 50 games: Axis 48, Allies 1, unfinished 1, with the median win in round 9. About 140 s per game.
                Bid tuning, Hard vs Hard:
                Russia bids of 12, 21 and 30 PUs: Axis still won 95–100%.
                Russia 30 plus UK/US bids: Axis 80–85%.
                Russia 40: roughly even (4 Axis, 3 Allies).
                Evaluator dataset: 110 Hard vs Hard games at Russian bids of 40–50. A logistic-regression position evaluator trained on it predicts the eventual winner from any snapshot at 83% (log-loss 0.345, cross-validated by game).
                Round robin, 120 games at the fixed 45 PU bid:
                AI Points vs the others
                Hard 89%
                Fast 85%
                Our bot (at the time) 51%
                Easy 25%
                Does Nothing 0%
                Hard vs Fast head to head: 5 wins to 3, 2 unfinished.
                At this bid, Axis 56 and Allies 57 across all games.
                Speed: Hard vs Hard takes about 16 minutes per game; Fast vs Fast about 3.4 minutes, roughly 1/6 of the time at almost the same strength.

                Using self-play to balance scenarios: what I'd do

                The biggest caveat: AI bias is not scenario balance. The Hard AI needed a 40–50 PU Russian bid to break even, while expert human play puts this game near even (an Allied bid of 3–6 PUs at most). The AI simply plays the Allies worse. A bid found by AI self-play measures how well the AI handles each side, not how balanced the scenario is. Use self-play to compare variants (scenario A vs scenario B under the same AIs), and calibrate the absolute number against human games where you can.
                Check the AI understands the scenario. New units or rules the AI misjudges will skew results. Watch a few full games before trusting a batch.
                Use Fast AI for volume and Hard for spot checks.
                Plan the number of games. The standard error of a win rate near 50% is about 0.5/√n: roughly ±5 points at 100 games, ±2.5 at 400. Batches of 10 are much too noisy to tell balance apart.
                Fit the even bid across all games rather than searching step by step. Play a spread of bids, fit P(Axis win) against the bid with logistic regression, and read off where it crosses 50%, with a confidence interval. I started with step-by-step adjustment and it wasted games.
                Get more signal per game than win/loss:
                Score unfinished games by TUV share or victory cities instead of discarding them.
                Use a position evaluator's P(win) at a fixed round as a lower-variance measure. It needs care, because it inherits the AI's bias.
                Record snapshots every turn, not just results. They show where and when a scenario tips: which territories swing, the round after which games are decided, and opening outcomes. That tells you what to change, not just which side is favoured.

                I didn't include the per-territory swing analysis or the evaluator's calibration tables; say if the poster would want either. The two purchase screens from earlier are still running."

                Me: I only know what a small bit of that means, but it sounds pretty cool.

                1 Reply Last reply
                Reply Quote 0
                • B
                  beelee @husky81
                  last edited by

                  @husky81

                  Can you combine scripted responses while adding some van dam uniqueness ?

                  It'd be a projrect. Start with the optimal 3-5 moves and there counters and then do the same from there. Can see that exploding quite quickly 🙂

                  But instead of just choose between these 3 and 5 choices add some claudeness to it as well.

                  Idk I have basically zero knowledge of how hard this is but it seems pretty hard 🙂

                  Keep it up though, we want you to succeed 🙂

                  TheDogT H 2 Replies Last reply
                  Reply Quote 0
                  • TheDogT
                    TheDog @beelee
                    last edited by

                    @husky81 said

                    I am currently trying to build a proper logic for purchase and non combat move.

                    Have you/ClaudeBot seen TripleA AI logs?
                    In game
                    Debug> HardAI> Show Logs
                    Enable AI Logging =TICK
                    Log Depth=Finest
                    Log History To=99

                    AFAIK this is only in memory and not a file.

                    It shows amongst other things how it categorizes its purchases and how it values the purchase of units and places them where needed.

                    https://forums.triplea-game.org/tags/thedog
                    https://forums.triplea-game.org/topic/3741/curated-best-top-maps-triplea-guides

                    H 1 Reply Last reply
                    Reply Quote 1
                    • H
                      husky81 @TheDog
                      last edited by

                      @TheDog said:

                      ow Logs
                      Enable AI Logging =TICK

                      No, but this is useful information. Am I right to think that this is more about seeing the current hard AI's reasoning? Currently i am attempting to solve the problems on my own (With Opus), I've told it explicitly not to look into the current ai's reasoning because I don't want to reverse engineer that ai at this point in the project.

                      Of course at some point in the future I could change my philosophy, especially if I either hit a brick wall and run out of ideas to improve my bot, or if i can finally beat triple A's bot, I would be more interested in mining the existing ai for ideas at that point. If there's another reason to look at the ai logs that I'm missing, let me know.

                      1 Reply Last reply
                      Reply Quote 0
                      • H
                        husky81 @beelee
                        last edited by

                        @beelee said:

                        s and there counters and then do the same from there. Can see that exploding quite quickly

                        But instead of just choose between these 3 and 5 ch

                        I can see the interest in making it do some weird plays, like going for fancy VC play, or doing different versions of KJF. I would definetly be interested in doing that. but currently I'm at the stage of trying to get it to stop doing ridiculous and stupid things like buying all infantry and turtling in moscow, or throwing it's Luftwaffe away trying to kill a naval stack that doesn't even had a single tt under it.

                        I'm very much at the walk before you can fly stage.

                        TheDogT 1 Reply Last reply
                        Reply Quote 1
                        • TheDogT
                          TheDog @husky81
                          last edited by

                          @husky81
                          Does the Claude bot have some simple stacking rules like the following?
                          Infantry > Artillery
                          Infantry > Armor
                          Destroyer > Cruisier
                          Destroyer > Battleship
                          Destroyer > Carrier
                          Destroyer = Transport
                          As this determines what to buy/place and then move.

                          If attacking it should have Armor in the stack
                          If attacking 66%+ chance of success, otherwise don't attack.

                          Also if Defending 50%+ chance of success, otherwise retreat

                          Just so you know , this is not a clone of what the TripleA AI does. It just me putting stuff on the page, for a starting point.

                          https://forums.triplea-game.org/tags/thedog
                          https://forums.triplea-game.org/topic/3741/curated-best-top-maps-triplea-guides

                          TheDogT H 2 Replies Last reply
                          Reply Quote 3
                          • TheDogT
                            TheDog @TheDog
                            last edited by

                            The above is only for that map.
                            It should be based on the units xml contents, not their unit names.

                            So like
                            Infantry 1-2-1, 3pu as cheapest land unit
                            Artillery 2-2-1 with artillery" value="true as support/mixed unit
                            Armour 3-3-1 with Blitz, so an Attack/mixed Unit
                            Destroyer 2-2-2 8pu as cheap surface sea unit

                            https://forums.triplea-game.org/tags/thedog
                            https://forums.triplea-game.org/topic/3741/curated-best-top-maps-triplea-guides

                            1 Reply Last reply
                            Reply Quote 2
                            • H
                              husky81 @TheDog
                              last edited by

                              @TheDog Currently it a values units by TUV (so cost), but also by "combat power" which is pips + 2.6* hit points, so an infantry is worth 2+2.6=4.6 in a defensive stack, 1 +2.6=3.6 in an attacking stack.

                              BB's are 4+2*2.6=9.2. This formula was arrived upon by simulating thousands of medium sized (6-12) unit per side land battles, and fitting the value of m in the equation pips + m * hitpoints that best predicted the favorite in the most circumstances. The simulation said that m was pretty flat between 2 and 4, so you could multiply hit points by anything from 2 to 4, then add pips to get a units approximate combat power when in a medium sized stack. (according to Claude)

                              the attack threshold is 70% to win I believe. Eventually, I want to have this become adjustable by the current board state. Every turn there is a function that the bot runs to test which side is winning, %to win estimate. (currently this function is not great since my training games only involve weak players). but my idea is that the bot will look for trends in the % to win, and if it notices that it's trending downward, or is really low, the bot will lower it's % to win threshold for attacks, choosing the most impactful ones.

                              So if I can ever get it to work right, it will 1-2 yolo on berlin or Karelia if it thinks it's getting close to the point of no return.

                              TheDogT 1 Reply Last reply
                              Reply Quote 0
                              • TheDogT
                                TheDog @husky81
                                last edited by

                                @husky81
                                What do think of having temporary hard ratios for Land, Sea, Air for each nation?
                                Losses are replaced in these ratios.
                                For example
                                Germany Land:80%, Sea:0% Air:20%
                                Russia Land:80%, Sea:0% Air:20%
                                Italy Land:50%, Sea:30% Air:20%
                                You get the idea.

                                This is to see if you can beat your current scores, as maybe your bots focus is too wide?

                                Ideally these ratios should be dynamic, perhaps based on the nearest VC objectives.

                                https://forums.triplea-game.org/tags/thedog
                                https://forums.triplea-game.org/topic/3741/curated-best-top-maps-triplea-guides

                                1 Reply Last reply
                                Reply Quote 1

                                Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                                Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                                With your input, this post could be even better 💗

                                Register Login
                                • 1
                                • 2
                                • 2 / 2
                                • First post
                                  Last post
                                Powered by NodeBB Forums