AEO, GEO, LLMEO; it goes by many names and it's the modern version of SEO for AI Chatbots. In essence, AI-eats-everything. Keeping in that spirit, this summer Gruntwork decided to use AI to track its AI "search" performance. We vibe coded an addition to our internal portal, our "Gruntwork AEO Monitor", and in this post I'd like to share how that's going.
This is the first post in a series. Nobody has AEO figured out yet, including us, so, we're going to write it up as we go. This first post is about what the tool is and how it works, and in subsequent posts (after we've had more time to gauge results) we'll discuss whether it worked.
AEO Monitoring Fundamentals - Ask the same questions every week
Once a week we have an automated job that sends the same set of questions to seven models (ChatGPT, Claude, Gemini, DeepSeek, Llama, Mistral and Qwen) and saves every request and response. If a model supports web search, we ask it twice, once with search turned on and once with only what it learned in training. The two answers often disagree. Without search you're seeing what made it into the training data. With search it's closer to what the web says about us this week.
Some of the questions name us, like "What is Terragrunt and what problems does it solve?" or "Is Terragrunt still worth using?" For those, we want to know whether the answer is accurate. The rest are questions someone would ask before they've heard of us, like "What are the best Terraform orchestration tools?" or "How do I set up an AWS landing zone with Terraform/OpenTofu?" There we just want to know if we got mentioned at all.
We track three products this way: Gruntwork (21 questions), Terragrunt (24) and Terragrunt Scale, our commercial product for teams running Terragrunt (34).

Grading the answers
Nobody is going to read several hundred AI answers a week, so a second model grades each one. It gets the question, the answer and a short description of the product, and it grades against a fixed rubric. When a question names us, it rates the answer good, OK or bad. Otherwise, it records whether we were named and, if the answer was a list, where we landed in it.
The tool saves every answer in full, so when a grade looks off you can click through and read what the model actually said. In the example below, DeepSeek answered "What are the best tools for managing Terraform at scale?" and put Terragrunt third.

Each product also gets a grid of questions against models. It's the quickest way to see that one model has us wrong, or that one question is a blind spot for all of them.

Tracking scores over time
The grid rolls up into two numbers per product that we follow week to week. Mention rate is the share of non-branded answers that name us, and accuracy rate is the share of branded answers that get us right.

AI is non-deterministic and so the models don't give the same answer every time. The scores move around a little every week, noise basically, so we're looking for larger trends over longer time-periods.
When a number does move, we can compare two runs cell by cell and see which models and which questions changed.

Turning scores into content ideas
A score doesn't tell you what to write. So when the grader marks an answer that left us out or got us wrong, it also suggests content that might have changed the answer. The tool collects those suggestions from the whole run and ranks them by priority and by how often they came up.

The list is long, and I don't expect anyone to read it start to finish. Mostly it gets read through Claude instead.
Using it through Claude
Our internal portal has an MCP server, and the AEO results and content priorities are available through it. The marketing contractor we work with has it hooked up to Claude Code, so he doesn't spend much time in the dashboard. He asks Claude to pull the current priorities and the answers behind them, and they work out content ideas together. Those turn into blog posts and copy changes on gruntwork.io and terragrunt.com, and the cycle continues.
How it's built
Our internal tool runs on Google Cloud and the AEO monitoring engine reuses infrastructure our other internal tools already had. The app sits behind Identity-Aware Proxy and is backed by Firestore. Every week a Cloud Scheduler job kicks off a Python Cloud Function, and that function queues one Cloud Tasks job for each combination of question, model and search mode. Each job asks the model its question (Gemini through Vertex AI, the others through OpenRouter), has Gemini Flash grade the answer, and saves the result in Firestore. The dashboard reads from Firestore as results come in, and the MCP server reads the same data.
The cost of running ~1000 AI queries/week
When we started we used nine models and about 1,000 calls, and a run cost about $30. We've since cut it to seven models and about 790 calls, so a run now costs about $15, or around $60 a month.
Where we're starting
We started experimenting in late July, the tool settled down in early August so we've only been paying close attention for four to six weeks. As you can see in the charts, nothing has moved much yet, which is perhaps not surprising. New content has to get written and then crawled and cited before it shows up in search answers, and it takes longer still to show up in what a model knows on its own.
We do have a baseline now. We know where we're missing (most non-branded questions about Terragrunt Scale come back without mentioning us), and we've gotten into the habit of turning that into content every week. I think if we keep at it, the numbers will start to move.
The SaaS tools in this space are more sophisticated than ours, and they generate questions dynamically where we use a fixed list. Ours is a quick and dirty in-house tool, but we can customize it to our needs and it sits next to the rest of our go-to-market data and plugs into the Claude setup we already use every day.
Next in the series
In the next post we'll cover what we changed, in the tool and in our content, and whether any of it moved the needle.


-
No-nonsense DevOps insights
- Expert guidance
-
Latest trends on IaC, automation, and DevOps
-
Real-world best practices