Having spent a good chunk of 2026 building AI agents for the AWS AI League Agentic Maze, including creating the open-sourced version to practice outside of competitions, the announcement of the Agentic Football challenge caught me by surprise. Where the dungeon challenge is a single agent solving a scored map run, football is five agents playing head-to-head against another team’s five agents, in real time, with a live crowd of other competitors watching the scores tick over. It is a different kind of problem, and a different kind of fun. The game has a number of issues, which is probably why AWS are calling it the alpha season, but I think it has the potential to be the most fun gamified learning since DeepRacer: it’s a spectacle to watch, it’s addictive, and there’s the unknown element of facing up against a team whose behaviour affects your behaviour, so the outcome is uncertain - just like real football!
This post is a reflection on the experience - the different formats I played, the architecture I deployed, and our run through the EMEA virtual knockouts to the final in the AI League game (at the time of writing Alpha season is approaching week 4/7). It deliberately stops short of the specifics that make up our competitive edge; this is still a live competition, after all. What I can share is the shape of the thing, the lessons, and the frustrations, in the hope it helps the next person decide how they want to approach it.
The Three Formats
I came at Agentic Football through three different doors, and they are worth distinguishing because the experience varies enormously.
The workshop is the hands-on, build-it-yourself format. You get an AWS account, a set of instructions, and you deploy your own agents into your own infrastructure. This is where the real learning happens - you see every moving part, you can read your own logs, and you own the whole stack. If you want to actually understand how the game works, this is the format to seek out.
The AI League is one of the competitive events: 10 matches in a league format, top 16 progress to the finals with head-to-head knockout matches against other teams, culminating in a live final. This is where the workshop preparation pays off - or doesn’t. The pressure is real, the matches are quick, and the bracket is unforgiving.
The Alpha League is an alternative competitive event - a longer-running league format where you play a series of fixtures over a seven-week period, refine between games, and climb a table rather than survive a single-elimination bracket. It rewards iteration and consistency over one-shot brilliance, and it is a far more forgiving place to experiment with ideas before you commit them to a knockout. That said, some people already seem to have unassailable leads in their leagues for the top re:Invent prize, but there are ‘wildcard’ slots for fastest goal, giant-killer, winning streak and comeback of the week, so there’s still all to play for. You can still join, refine your skills and go for the wildcards now, so if you’re learning about this for the first time, all is not lost!
Each format teaches you something. The workshop teaches you the machinery, the Alpha League teaches you to iterate, and the AI League teaches you to hold your nerve.
The Workshop Architecture
One of the things that makes the Agentic Football workshop genuinely interesting from an engineering standpoint is that your agents run in an AWS account you have access to. You are not uploading a prompt to someone else’s black box; you are deploying real, running infrastructure that the game engine reaches into.
The shape of it is a cross-account invocation pattern. The game engine lives in AWS accounts owned by the organisers. When a match runs, the engine invokes your agents through Amazon Bedrock AgentCore. Your agents run as AgentCore runtimes or harnesses, one per position on the pitch, and each one calls a Bedrock model to make its decision and return a command.
Game engine account
--> Bedrock AgentCore runtimes or harnesses invoked
--> each agent calls a Bedrock model, returns a command
--> engine advances the game tick
A few things about this setup are worth calling out:
- AgentCore runtimes, not Lambda. The agents deploy as managed AgentCore runtimes. You do not need to build a container or manage ECR images - the deploy step packages your code and stands up the runtime for you.
- One runtime per position. You deploy an agent per role - goalkeeper, defenders, midfielders, forwards - and register each runtime’s ARN against its position in the portal. That per-position separation is deliberate, and it opens up an interesting design space (more on that later).
Memory, Gateways and Observability
Three AWS capabilities did a lot of heavy lifting, and they are the parts I would point a newcomer at first.
AgentCore Memory gives your agents a place to persist context. In a fast-moving game where decisions come in a steady stream of ticks, having somewhere to hold state across those ticks changes what your agents can reason about. It is the difference between an agent that reacts to the current frame and one that has any sense of what just happened.
MCP gateways are how you give your agents tools in a clean, standardised way. The gateway pattern lets you expose tools through a consistent interface that the runtimes can call. It keeps the agent code focused on decision-making and pushes the plumbing out to the gateway, which is exactly where you want it.
AgentCore Observability is the one I would tell people not to skip. Every agent invocation lands in CloudWatch, with tracing through X-Ray and metrics emitted per runtime. When a match goes wrong - and matches go wrong in ways that are not obvious from the scoreline - the logs are the only honest account of what your agents actually did. I spent a lot of time in CloudWatch correlating what I saw on the pitch with what the agents were deciding tick by tick. You cannot tune what you cannot see, and observability is what makes the whole thing visible. This is what makes the AI League and Alpha Season variants so frustrating - you don’t get to see any of this!
What Comes Across Each Tick
The game runs on a tick cadence - every couple of seconds, the engine sends each of your agents a snapshot of the game state and asks for a command back. That snapshot describes the world: where the players are, where the ball is, who has possession. Your agent reads it, decides, and returns an action from the game’s command vocabulary - moving to a location, passing, shooting, tackling, pressing, marking, and so on.
There is genuine craft in interpreting that state well - that interpretation is a big part of where teams differentiate.
The command contract itself has one sharp edge worth flagging: the engine expects a clean, parseable response in a specific shape, and if your agent streams anything else - stray prose, extra lines, a malformed structure - the command fails to parse and your agent effectively does nothing that tick. Making your agents return exactly the right shape, every time, is unglamorous work that pays off immediately.
Models Matter - and So Does Position
Without getting into what we run where, one clear lesson from all our matches is that the choice of model matters, and not every model is equal for this task. Some models simply understand the game state better - they make more sensible decisions from the same information. Others are faster, and in a game that runs on a tick clock, latency is not a cosmetic concern; a model that thinks too long is a model whose decision arrives late or not at all.
That trade-off between understanding and latency is real, and it does not resolve to a single “best” model. Which leads to an idea worth sitting with: because you deploy an agent per position, you are not obliged to use the same model everywhere. A goalkeeper’s decision problem is not a forward’s. It is entirely reasonable to imagine matching different models to different roles based on what each position actually needs - reach for more reasoning where the decision is hard, and more speed where the decision must be quick.
The Frustrations of No-Code
The AI League and Alpha Season competitions offer no-code routes into building a team, and I understand why - they lower the barrier and let people get a team on the pitch quickly. But if your goal is to actually learn, or to compete seriously, the no-code options become frustrating fast.
The problem is visibility. When your team underperforms and all you have is a set of dropdowns and a prompt, you cannot see why. You cannot read the logs, you cannot correlate decisions with outcomes, and you cannot form and test a hypothesis about what to change. You are reduced to guessing, and guessing does not improve. The abstraction that makes the no-code route approachable is the same abstraction that hides everything you need to get better.
The workshop route is more work up front, but it is a far better learning experience precisely because nothing is hidden. You own the agents, you own the infrastructure, you own the logs. Every decision your team makes is inspectable, and every change you make is measurable. If you are the kind of person who wants to understand the system rather than operate it, do the workshop. The extra effort is the point. Most importantly - it builds transferable skills on agentic AI, troubleshooting and observability you can take straight into your day job.
My Run to the AI League Final
Which brings us to the EMEA virtual knockouts, where all of the above met the pressure of single-elimination football.
We went in as MarkRoss, and the bracket was a sixteen-team knockout - round of sixteen, quarter-finals, semi-finals, and final.

- Round of 16: we drew Blitz Cairns and had, frankly, our best game - a 14-1 win
- Quarter-final: a much tighter affair against Helm Claws, 6-4. Both teams scoring freely, and we edged a game that could have gone either way.
- Semi-final: the closest of the lot, 5-4 against Blitz Barbs. One goal in it, and we came back from 4-1 down!
- Final: and here we ran into Blitz Cougars. They had been imperious all tournament, including an 8-0 quarter-final and an 11-0 semi-final. They beat us 8-4.
Reaching a virtual EMEA final is something I’m proud of, and Blitz Cougars were worthy winners on the tournament as a whole.
A Note on the Final
There is one thing about the final I want to record honestly, without turning it into an excuse.
At kickoff, a previously unknown bug surfaced. Blitz Cougars lined up in a 1-1-1-1 formation, and the bug forced our team to mirror that same 1-1-1-1 formation despite us having selected a different one. That left several of our players out of their intended positions for the entire match. We flagged it to AWS immediately at kickoff; they were made aware and chose to proceed with the match as played, and the result stands. The formation I had chosen might have helped; it might not have mattered at all. Playing out of position certainly did not help, and the most frustrating thing for me is not knowing whether I need to go back to the drawing board with my team, or whether it’s already championship material on a day without formation bugs.

Final Thoughts
Agentic Football is the most engaging thing I have built in the AI League so far. It is a proper systems problem - real-time decision-making, model selection, observability - wrapped in a game that is genuinely enjoyable to watch. You come for the football and you stay for the engineering, or perhaps the other way around.
What makes it more than a game is that every part of the stack you touch is a production AWS service doing a production job. The football is the fun wrapper; underneath, you are learning the same building blocks that power real agentic systems. Here is each service the workshop puts in your hands, what it does in the game, and where it earns its keep in the real world:
- Amazon Bedrock AgentCore Runtime - hosts each of your five agents as a managed, serverless runtime, with session isolation and scaling handled for you. In the game it stands up one agent per position and runs it every tick. In the real world this is how you deploy and scale production AI agents without managing containers or servers - customer-support agents, coding assistants, research agents - so the team can focus on agent behaviour rather than infrastructure. Docs: AgentCore Runtime.
- Amazon Bedrock (foundation models) - the actual reasoning engine each agent calls to turn a game-state snapshot into a decision. In the game, model choice is a genuine competitive lever. In the real world, Bedrock is the single API through which you access and compare frontier models (Anthropic, Meta, Mistral, Amazon Nova and more) for chatbots, summarisation, extraction and RAG, letting you pick the right model per task on cost, latency and quality. Docs: Amazon Bedrock.
- Amazon Bedrock AgentCore Memory - persists context for an agent across ticks so it can reason about what just happened, not only the current frame. In the real world this is the memory layer for stateful assistants: remembering a user’s preferences across a conversation, carrying context through a multi-step workflow, or maintaining session history in a support bot. Docs: AgentCore Memory.
- Amazon Bedrock AgentCore Gateway - exposes tools to your agents through a clean, standardised (MCP) interface, keeping the plumbing out of the agent code. In the real world this is how you give agents governed access to your APIs, databases and internal services - turning existing systems into agent-callable tools without rewriting them. Docs: AgentCore Gateway.
- Amazon Bedrock AgentCore Observability - traces and monitors every agent invocation so you can see what your agents actually did. In the game it is the honest account of each decision, tick by tick. In the real world it is how you debug, monitor and optimise agents in production - latency, token usage, error rates and full traces - which is essential once agents are making decisions that matter. Docs: AgentCore Observability.
- Amazon CloudWatch - the destination for AgentCore’s logs and metrics, where I spent hours correlating on-pitch behaviour with per-tick decisions. In the real world CloudWatch is the standard observability backbone for AWS workloads: dashboards, alarms and log analysis across virtually every service. Docs: Amazon CloudWatch.
- AWS X-Ray - provides distributed tracing through the agent invocations, so a single decision can be followed end to end. In the real world X-Ray is how you find the bottleneck or failure in a distributed or microservice application, tracing a request across every hop. Docs: AWS X-Ray.
If you are thinking about getting into it: do the workshop, deploy your own agents, live in your logs, and treat model choice as a first-class decision. Iterate in the Alpha League, then hold your nerve in the AI League. And if you find yourself in a final against a team that has been winning by double figures, enjoy it - you earned your place there.
See you on the pitch.
Want to join the AWS AI League? Visit the AI League Page for details on upcoming events. Follow the league in the AWS AI League Builder Space for updates, or join the AWS AI Community to connect with fellow competitors.
Mark Ross, Atos
