Back to Blog
Tools

How we vibe coded our own AI answer monitor (Part 1)

Zach Goldberg
Zach
Goldberg
,
Chief Executive Officer
October 8, 2026
Grunty the yeti points at a rising trend line on a score dashboard, fed by seven graded AI chat answers

AEO, GEO, LLMEO; it goes by many names and it's the modern version of SEO for AI Chatbots. In essence, AI-eats-everything. Keeping in that spirit, this summer Gruntwork decided to use AI to track its AI "search" performance. We vibe coded an addition to our internal portal, our "Gruntwork AEO Monitor", and in this post I'd like to share how that's going.

This is the first post in a series. Nobody has AEO figured out yet, including us, so, we're going to write it up as we go. This first post is about what the tool is and how it works, and in subsequent posts (after we've had more time to gauge results) we'll discuss whether it worked.

AEO Monitoring Fundamentals - Ask the same questions every week

Once a week we have an automated job that sends the same set of questions to seven models (ChatGPT, Claude, Gemini, DeepSeek, Llama, Mistral and Qwen) and saves every request and response. If a model supports web search, we ask it twice, once with search turned on and once with only what it learned in training. The two answers often disagree. Without search you're seeing what made it into the training data. With search it's closer to what the web says about us this week.

Some of the questions name us, like "What is Terragrunt and what problems does it solve?" or "Is Terragrunt still worth using?" For those, we want to know whether the answer is accurate. The rest are questions someone would ask before they've heard of us, like "What are the best Terraform orchestration tools?" or "How do I set up an AWS landing zone with Terraform/OpenTofu?" There we just want to know if we got mentioned at all.

We track three products this way: Gruntwork (21 questions), Terragrunt (24) and Terragrunt Scale, our commercial product for teams running Terragrunt (34).

The Gruntwork AEO Monitor: a list of weekly runs with their progress and cost, next to sections for trends, run comparison, per-product question grids and content priorities
The Gruntwork AEO Monitor, with every weekly run and what it cost.

Grading the answers

Nobody is going to read several hundred AI answers a week, so a second model grades each one. It gets the question, the answer and a short description of the product, and it grades against a fixed rubric. When a question names us, it rates the answer good, OK or bad. Otherwise, it records whether we were named and, if the answer was a list, where we landed in it.

The tool saves every answer in full, so when a grade looks off you can click through and read what the model actually said. In the example below, DeepSeek answered "What are the best tools for managing Terraform at scale?" and put Terragrunt third.

A single graded response: DeepSeek, knowledge only, listing Terragrunt as an option at rank 3
Each grade links to the full response, so you can check the grader's work.

Each product also gets a grid of questions against models. It's the quickest way to see that one model has us wrong, or that one question is a blind spot for all of them.

The Terragrunt question-by-model grid: branded questions graded Good, OK or Bad, non-branded questions marked Listed with a rank, Miss, or N/A
Terragrunt's grid for one run. The branded questions mostly come back accurate. The non-branded rows are where we have work to do.

Tracking scores over time

The grid rolls up into two numbers per product that we follow week to week. Mention rate is the share of non-branded answers that name us, and accuracy rate is the share of branded answers that get us right.

Two line charts of mention rate and accurate rate per product, with web search and knowledge-only lines, from July to September 2026
Mention rate and accuracy rate for each product since late July. Solid lines are with web search, dashed lines are model knowledge only.

AI is non-deterministic and so the models don't give the same answer every time. The scores move around a little every week, noise basically, so we're looking for larger trends over longer time-periods.

When a number does move, we can compare two runs cell by cell and see which models and which questions changed.

A run comparison showing which models improved or regressed and which questions flipped between Listed and Miss
Comparing two runs. Gemini named us a bit more often this week and DeepSeek a bit less, and the list below shows which questions changed.

Turning scores into content ideas

A score doesn't tell you what to write. So when the grader marks an answer that left us out or got us wrong, it also suggests content that might have changed the answer. The tool collects those suggestions from the whole run and ranks them by priority and by how often they came up.

A ranked list of content priorities, each marked HIGH, such as comparison content against AWS Control Tower and content on multi-cloud landing zones
Content priorities from one run.

The list is long, and I don't expect anyone to read it start to finish. Mostly it gets read through Claude instead.

Using it through Claude

Our internal portal has an MCP server, and the AEO results and content priorities are available through it. The marketing contractor we work with has it hooked up to Claude Code, so he doesn't spend much time in the dashboard. He asks Claude to pull the current priorities and the answers behind them, and they work out content ideas together. Those turn into blog posts and copy changes on gruntwork.io and terragrunt.com, and the cycle continues.

How it's built

Our internal tool runs on Google Cloud and the AEO monitoring engine reuses infrastructure our other internal tools already had. The app sits behind Identity-Aware Proxy and is backed by Firestore. Every week a Cloud Scheduler job kicks off a Python Cloud Function, and that function queues one Cloud Tasks job for each combination of question, model and search mode. Each job asks the model its question (Gemini through Vertex AI, the others through OpenRouter), has Gemini Flash grade the answer, and saves the result in Firestore. The dashboard reads from Firestore as results come in, and the MCP server reads the same data.

The cost of running ~1000 AI queries/week

When we started we used nine models and about 1,000 calls, and a run cost about $30. We've since cut it to seven models and about 790 calls, so a run now costs about $15, or around $60 a month.

Where we're starting

We started experimenting in late July, the tool settled down in early August so we've only been paying close attention for four to six weeks. As you can see in the charts, nothing has moved much yet, which is perhaps not surprising. New content has to get written and then crawled and cited before it shows up in search answers, and it takes longer still to show up in what a model knows on its own.

We do have a baseline now. We know where we're missing (most non-branded questions about Terragrunt Scale come back without mentioning us), and we've gotten into the habit of turning that into content every week. I think if we keep at it, the numbers will start to move.

The SaaS tools in this space are more sophisticated than ours, and they generate questions dynamically where we use a fixed list. Ours is a quick and dirty in-house tool, but we can customize it to our needs and it sits next to the rest of our go-to-market data and plugs into the Claude setup we already use every day.

Next in the series

In the next post we'll cover what we changed, in the tool and in our content, and whether any of it moved the needle.