Making an AI feature faster and cheaper: eight things that worked
10 October 2026
MooTrip is an AI trip planner: a language model writes the itinerary, and you change it by chatting. In the first two weeks, one edit to a 12-day plan went from taking two to four minutes to taking 14 seconds.
No single trick did that. This article lists the eight changes that made the product faster or cheaper, each as a rule you can apply to your own AI feature, with what happened when I applied it.
1. Ask the AI for only what changed
At first, every edit made the model write the whole plan again. Changing one flight price on a 12-day trip meant regenerating twelve days, which took two to four minutes.
I changed the edit so the model returns only the days it changed, and code merges them into the plan. The same edit took 14 seconds.
If your model rewrites a document to change one line, this is the first thing to fix.
2. Shrink what the AI has to write
Later I tested two ways to describe a changed day. In one, the model wrote every row in full. In the other, an unchanged row was just its id and time, and only new or changed rows were written out.
Both passed my tests at a similar rate. The short version produced answers under half the size (about 4,000 characters against 9,200) and took about half as long (3.9 seconds against 7.3).
Output is the slow, expensive part of a model call. Anything the model copies without changing is time you are paying for.
3. Check the output as it arrives, and redo only what failed
A plan arrives day by day as the model writes it. I check each day the moment it is complete, and stop the model at the first day that fails.
Before this, a failed plan was only discovered at the end, and the retry started from nothing. One trip spent about 285 seconds writing two full plans because of one bad day. Now the retry keeps the days that passed and writes only the rest.
4. Show progress as it happens
A five-day plan took about 35 seconds to write when I measured it. That is a long time to look at a spinner.
So the page shows the plan as it is written: the summary after about 6 seconds, the packing list at 10, then each day as it passes its checks. The total time is the same, but the wait feels much shorter because there is something to read.
The same applies outside AI. The photos on my landing page used to finish appearing after 4.2 seconds; starting them earlier brought that to 2.4.
5. Check a free tier's limits against one real request
I moved to a model that wrote a four-day plan in 25 seconds, against 50 to 70 on the one before. It looked like a clear win.
Its free tier allowed 8,000 tokens a minute. A single plan needs more than that, so every request was turned away, and the setup ended up costing money I had not planned for.
Before depending on a tier, compare its limits with the size of one real request, not with a small test.
6. Run your code next to your data
My database was in Singapore. My pages were being built in the eastern United States, the hosting default. Every page read the database several times, so every click waited about half a second for data to cross the Pacific and back.
One setting moved everything to Singapore.
7. Measure before you speed anything up
I added three timers: from a click to the page appearing, for each step the server takes to build a page, and for each model call. Each run now logs one line per step with how long it took.
Without that, a slow request is a guess. With it, the log says which step used the time.
8. Do not fetch what you already have
The list of trips was loaded once on the server and then again by the page when the sidebar opened. Photo searches that found nothing were repeated on the next visit.
The page now uses the list it was given, and an empty photo search is remembered for five minutes. Neither change is clever. Both removed work that produced nothing.
Which to do first
If I were starting again, I would do them in this order:
- Measure (7), so you know where the time goes.
- Ask for less (1 and 2), because model output is usually the largest cost.
- Fail early and retry narrowly (3).
- Show progress (4), for whatever wait is left.
The rest are ordinary web performance work, and worth doing once the AI part is under control.
A note on the numbers: the 14 seconds, the 285 seconds and the half second each come from one observed case, not an average over many runs. They show the size of the effect, not a benchmark.