Testing an AI feature: what my evals changed my mind about
10 October 2026
When an AI feature fails, the obvious fix is to add another instruction to the prompt. I did that, measured it, and found it made things worse.
MooTrip is an AI trip planner where a language model writes the itinerary and edits it on request. This article covers how I tested those edits, and six things the tests taught me that I would not have guessed. It is for anyone building on a model who has not yet set up tests for it, often called evals.
The setup, briefly
Edits kept going wrong in the same few ways. A new stop would overlap the bus ride next to it. The model would change the wrong row. "Add two days" would come back a day short.
I built a small test set: eight edit requests on saved plans from two trips. Each test replays the request through the same code the app runs, then through the app's own checks. I ran each case several times and counted passes.
1. Build the test set from real plans
The cases use saved plans, including three versions of a real trip, and the kinds of request that had been going wrong: adding a lunch and a snack, adding attractions to a full day, removing a stop and moving dinner.
Invented cases test what you imagine the model finds hard. Real plans and real requests test what it does find hard.
2. Count first-try passes and after-retry passes separately
The app gives the model one retry when an answer fails a check. So I recorded two numbers for every run: passed first time, and passed after the retry.
They tell you different things. First-try rate is how good the instructions are. After-retry rate is what the user experiences. An edit format can look weak on the first and fine on the second, and you need both to choose.
3. When the AI keeps failing, make the task easier, not the instructions longer
The main failure was overlapping times: 16 of 23 first-attempt failures in the first comparison.
My first fix was more rule text. I reworded the timing rule twice and told the model to look at the whole day before answering. First-try passes went from 17 of 24 down to 14, and the retry rescued fewer failures than before.
What worked was changing the task. I had been asking the model to patch a day: edit this row, add one after that row. I tested four formats, two that patch and two that have the model write the day out in order, from morning to night.
| Format | Passed first try | Passed after retry |
|---|---|---|
| Patch rows by number | 15 of 24 | 18 of 24 |
| Patch rows by id | 17 of 24 | 22 of 24 |
| Rewrite the whole day in full | 19 of 24 | 24 of 24 |
| Rewrite the day, unchanged rows as id only | 17 of 24 | 21 of 24 |
The full rewrite did best. In both rewrite formats, overlaps fell to between one and three per 24 runs. A model spaces a day well when it writes it top to bottom. It does badly when asked to insert one row and leave the rest alone, because nothing prompts it to move the bus ride that now overlaps.
I adopted the last format. Its answers were under half the size of the full rewrite and took about half as long. In this first run it only matched patching by id, but once I corrected the test samples (see point 5) it pulled ahead: 18 of 24 first try against 16, and 22 against 18 after the retry. On a longer run it passed 33 of 40 first time and 39 of 40 after the retry.
4. Test the changes that obviously help
I was sure that showing the model each row's end time would reduce overlaps. If it can see that lunch ends at 13:30, it will not start the next stop at 13:15.
It made no measurable difference: 17 passes against 16.
I kept it because it costs nothing. But I would have claimed it as a fix if I had not measured, and I would have been wrong.
5. Keep the test samples in step with the format
Partway through I found that some of the saved plans were out of date. They had been stored in an older format, with stops that had no durations, so the model was being tested on input the app no longer produces.
A stale sample measures the sample. Whenever the format changes, the saved plans and the expected results have to change in the same piece of work.
6. Test the model the way the app calls it
My tests call the same function the app calls, with the same model, provider and settings. That way a pass in the tests means the same thing as a pass in the app.
The same goes for choosing between models. A single fast run tells you little. Running your own cases tells you whether a model can do your task.
How much to trust this
Eight cases from two trips is a small test set, and 39 of 40 means a true pass rate somewhere between about 87% and nearly 100%. A difference of two or three passes between formats is noise at this size.
That is still far better than no tests. It was enough to stop me shipping a prompt change that made things worse, and enough to choose a format on evidence.
If you take one thing from this: before adding a rule to your prompt, write down five real failures and run them. Then add the rule and run them again.