@Kindwind A neural net certainly makes more sense than a LLM. I didn't see anything obviously illegal in the play, but you didn't document the non-combat moves which caused a problem with your previous model
As you noted, your neural net is inferior to your algorithmic WeakAI, which is itself comparable to TripleA's Easy AI. However it does attack and would beat TripleA's Do Nothing AI, which is something. (I have seen a GGP (Ludii) that plays Chess so poorly that it has trouble against random play).
However, that it can't beat the training AI seems to point to a fundamental weakness with this sort of neural net. Note that AlphaZero played 44,000,000 games of Chess in its training process (600 years on your computer), and Chess is much simpler game computationally.