mootripTechnical blog

Refactoring an AI pipeline: what I cleaned up, and what it fixed

10 October 2026

Code that wraps a language model gets messy quickly. Every failure adds a check, a retry or a special case, and after a week the function that runs the model is a long tangle that nobody wants to touch.

MooTrip is an AI trip planner where a model writes and edits itineraries. Nine days in, I spent a day cleaning up the code that runs the model. This article covers the seven changes that mattered, as rules you can apply to your own pipeline.

1. Name the steps

The function that ran a plan had grown into one long sequence: call the model, check the answer, fix a few things, check again, build the reply.

I turned it into a list of named steps. Doing that exposed something: a function called "evaluate" was also quietly changing the plan, dropping flight prices in the wrong currency. A step that only reads and a step that changes things had been hiding under one name.

If you cannot name a step, it is probably doing two things.

2. Separate checks from repairs

Once the steps had names, they fell into two kinds, and I gave each its own rules.

  • A check judges the plan. If it fails, the plan goes back to the model.
  • A repair improves a plan that already passed. It can never reject the plan. If a repair fails, or leaves the plan worse, the repair is dropped and the plan is kept as it was.

This matters because repairs are where things quietly break. A repair that can fail the whole request turns a small improvement into a new way to lose a plan.

3. Keep one source of truth for the format

The plan's format was described in more than one place: the definition used to validate saved plans, and a hand-written guide to the fields in the prompt. A change to the format meant editing each by hand.

Now there is one definition. The model's version and the prompt's field guide are generated from it. The generated guide came out almost exactly the length of the hand-written one (8,877 characters against 8,880), so nothing was lost, and they can no longer disagree.

4. Do not describe to the model what it should not write

The prompt had a line telling the model never to write certain fields, because code fills those in.

I removed the line and removed those fields from what the model is shown. Naming a field in order to forbid it puts the field in front of the model, which invites it. A field the model has never seen is much less likely to be written, and code strips it if it is.

5. One name for one thing

A row's length was stored under two names: one for time spent at a stop and one for the length of a journey. A taxi ride given a zero under the stop's name caused a whole edit to be refused.

I merged them into one field. Two names for nearly the same thing give the model a choice to get wrong.

6. Delete what you replaced, the same day

On the day the model took over planning, I removed the old planner: 1,046 lines of code and its tests, and a further 264 lines of tests that were no longer needed.

Some of the deleted tests checked the exact wording of the prompt. They broke on every copy edit and said nothing about whether the product worked. A test that fails when you improve a sentence is testing the wrong thing.

Leaving the old code "in case" would have meant reading around it for weeks.

7. Refactor without changing behaviour, in its own commit

Each restructuring change was made on its own and marked "no behaviour change". The plan that came out before and after was the same.

That separation is what makes a refactor safe to review. If a change both reorganises the code and alters what it does, nobody can tell which part caused a new bug.

What it bought

The cleanup did not add a feature. The aim was to make later changes smaller: adding a check is now adding one entry to a list, and changing the format is editing one definition. One real bug also surfaced on the way: checks that were still running in the background after a failed call, and could write to the page while the retry was starting.

If your model-running code has become the file you avoid, start with the first two: name the steps, then separate the ones that judge from the ones that change. Most of the rest follows from that.

← All posts

PrivacyTerms