← Compendium

An operations handbook that writes along

An undocumented platform that keeps breaking, and next to it the new world is supposed to take shape. Anyone who only fixes each outage meets it again a few weeks later in a slightly different form. What helps is an operations handbook in Markdown in your own Git, where every insight ends up, including those from outages. An AI with read-only access collects data, takes notes and links earlier outages, and people review every change through a pull request. The result is a knowledge base that grows with every outage and makes the next one faster to fix. And the same source yields a report for the tech team and a summary for the CEO, with the real business impact in numbers.

The situation in one of our client projects was what it looks like in many companies that have grown over time. A platform built up over years, barely documented and constantly breaking. The business depended on it, so it had to keep running. At the same time the new world was to be built in AWS, so that the old one could be replaced one day. Our job was both: stabilise operations to keep the business alive, and make room for the migration.

It quickly became clear that both at once would not work, at least not the way we started out. We fixed every outage, removed the obvious cause as well as we could and moved on. After a few weeks similar problems came back, but each time slightly different, because we had of course removed the obvious cause. Stabilising ate the time we needed for building, and we were burning out doing it.

Understand the old world to replace it

There was plenty to do for the new world. There was no monitoring and there were no backups. The application was not containerised, and there was no safe way to deploy: new versions reached the server via git pull and a PHP compile script, which should never have gone into production like that, but that is how it was. The application also could not run on more than one instance, and taking it apart for that was a challenge in its own right.

To do that, we had to understand the existing system, and better than it was written down anywhere. Every outage revealed something about how the platform really worked, which parts depended on each other and which assumptions were wrong. That knowledge was just as valuable for the migration as for operations, but it got lost as soon as the outage was fixed. So we needed a proper operations handbook where we record every new insight about the platform.

A handbook in your own Git

To make writing easy, we chose Markdown as the format and a dedicated repository in the client’s Git. The structure was quickly clear: an architecture overview showing which parts exist and how they fit together, and a glossary for the domain terms that everyone uses but not everyone understands the same way. Both quickly helped build a shared picture, between us and the client’s team.

Instead of keeping the reviews after outages in some email thread, we put them in the same repository. Every outage got its own file, with symptoms, cause, fix and open points. That put everything in one place: how the platform is built, what the terms mean and what has gone wrong in the past.

text
operations-handbook/
├── README.md
├── architecture/
│   ├── overview.md
│   └── data-flow.md
├── glossary.md
├── operations/
│   ├── deployment.md
│   └── backup-and-restore.md
└── incidents/
    ├── 2025-03-04-checkout-timeout.md
    └── 2025-03-18-db-connections.md
A typical basic structure: architecture and glossary for the shared picture, operational procedures for everyday work, and one file per outage.

The AI writes along

This is where AI comes in. We wrote the documentation with Kiro, a development environment based on VS Code with an AI assistant, from our input. Any other VS Code with a language model, such as Claude or GitHub Copilot, works in a similar way. We described what we had found out, and the AI wrote it up, filed it in the right place and extended the architecture overview.

We soon used this for handling outages too. Instead of taking notes on the side, we had the AI record the findings and collect data in the background: from Cloudflare, from Grafana, from the database and from other metrics. With every step of the migration to AWS this became even more effective, because more and more metrics and details could be fetched through the AWS CLI. The IaC repository and the code repository could also be linked up quickly this way, so it became visible which change takes effect where in the infrastructure.

Just as important was that the AI never wrote straight into the handbook. Every change came as a pull request, and a person reviewed it before it was merged. That takes little time, because you only review rather than write, and it prevents a plausible-sounding but wrong explanation from becoming the supposed knowledge of the platform.

No, the AI did not do our work. We found and fixed the causes ourselves. But in stressful situations especially, writing good sentences is not the priority, and afterwards the details are no longer top of mind. The AI closes exactly that gap: it records while you work, and at the end there is a clean note instead of a half-filled document that nobody ever finishes.

Worth its weight in gold: linking outages

The greatest value came from something we had not planned at first. Because all outages were in the same repository, the AI could draw on earlier ones when a new one occurred. It recognised that the symptoms resembled an earlier outage, which cause had been found back then and which fix had helped. That was exactly what we had been missing at the start, when we kept investigating the same problems from scratch in slightly different forms.

1 Outage Symptoms appear, the team fixes them. 2 AI collects Read-only: metrics, logs, database, cloud. Takes notes. 3 Pull request A person reviews the note and links. 4 Handbook grows Outages, architecture, glossary. Next outage: earlier cases are at hand, the pattern is recognised faster
The cycle: every outage fills the handbook, and the handbook makes the next outage faster to solve. The AI only reads and proposes; nothing is merged until a person has reviewed it.

That made every fix faster, and the next steps of the migration could be planned more precisely. When three outages pointed to the same weak spot, we knew which part of the old platform should be replaced first. Along the way the architecture overview became more detailed piece by piece, because every outage brought another detail to light that would otherwise have stayed in the heads of individual people.

What emerges is a knowledge base that maintains and extends itself almost on its own. It does not go stale, because it is updated exactly when something new shows up, and it does not depend on a single person who has everything in their head. For us, this is one of the most effective ways to put today’s AI capabilities to use in operations.

One report for the team, one for the CEO

An outage has more than one audience. The tech team needs the details: timeline, cause, affected components, links to earlier outages and the next steps. The CEO needs something else. With every outage the CEO saw the business at risk, and at the same time the migration was not moving fast enough, because the team spent night after night fighting fires to keep the business alive. What matters at that level is not which cron job held connections open, but what the outage cost and what happens next.

Because all knowledge about an outage was already in the handbook, both versions came from the same source. The AI turned it into a technical report for the team and a short summary for the CEO. This kind of translation is exactly where it shines: it turns “connection pool exhausted” into a statement about what customers experienced and what that means for the business, without anyone having to write two texts after a long night.

Real numbers made the biggest difference. Through read access to the data, the AI could automatically analyse revenue around every outage and show the actual business impact instead of estimating it. A pattern emerged: during an outage revenue dropped, as expected, but after it ended it rose disproportionately, because customers caught up on what they had not been able to complete before.

Outage usual revenue Drop Catch-up Time
Schematic: during the outage revenue drops, afterwards customers catch up on part of it. The real damage is the drop minus the catch-up, and only with both numbers does it become tangible.

With these numbers the situation could be presented honestly: what an outage really costs the business, how often it happened and how much of the team’s time went into firefighting instead of the new world each time. That made it clear how important it was to put all resources into the fastest possible move instead of patching the old platform further. A gut feeling became a decision that could be backed by numbers, and the CEO could stand behind it because it was understandable.

text
What happened?
  One sentence, no jargon.

What did it cost?
  Revenue during the outage,
  catch-up afterwards, net.

What are we doing about it?
  Immediate action and replacement status.

What does it mean for the migration?
  Which part gets priority now.
A possible outline for the summary to the CEO: four questions, each answered in a few sentences.

How to start

Getting started does not take a big project. Create a repository in your own Git and give it a simple structure: an architecture overview, a glossary and a folder for outages. The first overview may be rough, a few boxes and arrows are enough, and it will get more precise over time.

From the next outage on, you write things down there, with the AI in your development environment. Tell it what you see and what you do, and let it turn that into a note. Step by step, set up read-only access to the data sources you look at during outages anyway, and require a pull request for every change to the handbook. After a few outages you will notice the AI drawing connections you would not have thought of yourself.

markdown
# 2025-03-18 Database connections exhausted

## Symptoms
Checkout aborts, error 502 from approx. 14:10.

## Cause
Connection pool full, a cron job keeps
connections open.

## Fix
Cron job stopped, pool restarted.

## Related
Similar to 2025-03-04 (checkout timeout).

## Open
Run the cron job as its own service with
limited connections in the new world.
What an outage note can look like. The “Related” section is the most important one: this is where the AI links the new outage to earlier ones.

What this means for you

An operations handbook rarely fails for lack of good intentions, but because nobody has time to write it, least of all in the middle of an outage. An AI that writes along, collects data and draws on earlier outages takes exactly that load off, as long as it can only read and people review every change. That closes the gap between what a system actually does and what you know about it, a gap we describe from the security angle in the article It was secure yesterday, wasn’t it?. And because the same knowledge can also speak the language of the executive team, every outage becomes an argument backed by real numbers. If you are facing a similar platform or finally want your operational knowledge in one place, we are glad to tackle it together.

Frequently asked questions

What belongs in an operations handbook?

At least an architecture overview, a glossary of domain terms, the key operational procedures such as deployment, backup and restore, and a note on every outage with symptoms, cause, fix and open points.

Why Markdown in Git rather than a wiki?

Markdown is easy to read and write for people and for an AI alike, and Git brings versioning and pull requests. That makes every change traceable and reviewed before it applies.

Should an AI have access to production systems?

Only read-only, and only with separate credentials per system. That lets it fetch metrics, logs and configuration, but not change anything. Access to production should always be granted with care, and for an AI all the more so.

How do you stop the AI from writing wrong things into the handbook?

By not letting it merge anything directly. Every change comes as a pull request, and a person reviews it. That takes little time, because you only read rather than write.

Which tools do you need?

A Git repository and a development environment with an AI assistant. We used Kiro; VS Code with Claude or GitHub Copilot works in a similar way. On top of that, read-only access to the data sources you look at during outages anyway.

How do you make the business impact of an outage visible?

By analysing revenue around the outage, not just during it. Customers often catch up on part of it afterwards, and the real damage is the drop minus that catch-up. An AI with read access to the data can analyse this automatically and summarise it in a way the executive team understands.

Is it worth it without a migration?

Yes. Any system that has outages and whose knowledge sits in a few heads benefits from it. The migration just makes the benefit especially visible, because knowledge about the old world directly improves the planning of the new one.

Related topics