Monitoring an AI feature: what to record so you can see why it failed
10 October 2026
When ordinary code fails, you get an error and a line number. When an AI feature fails, you often get nothing: the model returned something, a check rejected it, the user saw "couldn't make that change", and the evidence is gone.
MooTrip is an AI trip planner where a language model writes and edits itineraries. In its first two weeks I kept finding that I could not answer the question "why did that fail?". This article lists the seven things I now record, each with the incident that made me add it.
1. Keep the whole rejected answer
When a model's answer fails a check, the answer itself is the evidence. For the first few days I logged the error and threw the answer away, which told me that a plan had overlapping times but not what the model had written.
Now every rejected answer is logged in full, next to the errors that rejected it, so I can read exactly what the model wrote.
2. Send failures somewhere that lasts
A plan that failed after its retry was written to the server log, which my host keeps for about a day. If I did not look that day, it was gone.
Failed plans and edits are now reported to an error tracker with their reasons, where they stay and can be counted. A failure you cannot find a week later might as well not have been recorded.
3. Count which model answered
I use a first-choice model with others as fallbacks. For a while I had no idea how often the first choice was answering.
I added a count of answers per model and of how often the first choice fell back. Around the same time I learned that the first choice had been turning requests away, because one plan was larger than its rate limit allowed. The plans still arrived, because the fallback caught them.
A fallback that works hides the failure of the thing in front of it. Count it.
4. Time every step and every model call
Building a plan has several steps: the model call, the checks, the repairs, the look-ups. When a run was slow or timed out, I could not tell which step was responsible.
Each run now logs one line per step with how long it took and the running total. When a request times out, the log shows which step used the time.
5. Watch what the model spends its budget on
One model I tried spent 14,592 of its 16,000 available tokens on reasoning for a five-day trip, and had almost nothing left to write the plan with. A setting to reduce reasoning was being ignored.
The timing line now records how much went on thinking and when the first text arrived, so the same thing is visible at a glance.
6. Check that monitoring is receiving anything
I had two tracing tools installed. One of them replaced the other's global setup when it started, so the server traces I thought I was collecting were never arriving.
Nothing looked broken, because an empty dashboard and a healthy system look the same. After adding any monitoring, cause a failure on purpose and confirm you can see it.
7. Keep the error log quiet enough to read
A harmless database warning was filling the error log, which made real errors harder to spot.
I fixed the cause of the warning. If the log is mostly noise, you stop reading it, and then it is not monitoring anything.
What I would set up on day one
- Log every rejected answer in full, with its errors.
- Report every failure that reaches the user to somewhere permanent.
- Count answers per model, and fallbacks.
- Time each step.
Those four would have answered most of the "why did that fail?" questions I had in the first two weeks. The other three I would add as soon as something looked odd.
One caution: a model's answers and a user's requests can contain personal detail. Decide what you are allowed to keep before you log everything, and label records with an id you can look up, not with the person's name or email.