Launch day gets treated like a finish line more often than it should. In reality, the first 90 days after an agent goes live are when we learn the most about how well it actually fits a business — and where most of the real tuning happens, based on real customer conversations instead of our best pre-launch guesses.

Week one: tighter monitoring than usual

Every agent launches with elevated monitoring for the first two weeks — more frequent transcript review, faster response time if something looks off, and daily check-ins with the client's team rather than the weekly cadence that follows later. This is deliberate: the gap between how an agent performs in staging against sandbox data and how it performs against real, unpredictable customer traffic is where the most valuable and sometimes surprising information shows up.

In week one, we're specifically looking for patterns that didn't appear in our evaluation set — a phrasing style unique to this client's customer base, a product category question we hadn't anticipated, a tone mismatch that reads fine in testing but lands oddly with real customers. None of this means the agent launched broken. It means real traffic surfaces things testing can't fully replicate.

Weeks two through six: tuning against real data

Once the initial monitoring window closes, we shift into a tuning phase built entirely around real transcript review. Every escalation gets reviewed for whether it was necessary, every low-confidence response gets reviewed for whether the agent should have been more decisive, and every case where a customer had to rephrase becomes a candidate for a scope or prompt adjustment.

This is also when we usually have our first substantive conversation with the client about expanding or narrowing the agent's scope, informed by actual usage instead of the initial plan. A client might discover their agent is handling far more of a certain question type than expected, making it worth building out deeper handling there, while another category barely comes up at all and can be deprioritized.

Days 60 to 90: settling into steady state

By this point, the tuning cadence usually slows, not because we've stopped paying attention, but because the agent's behavior has genuinely stabilized against the patterns in that store's real traffic. This is when we run the first formal performance review with the client — resolution rate, escalation rate, and wherever possible, a direct comparison against the support metrics from before the agent launched.

It's also the point where we have an honest conversation about what's next: whether the current scope is the right long-term footprint, whether a second workflow is worth mapping, or whether the model choice underneath the agent still makes sense given how traffic actually played out versus what was projected at launch.

Why we frame launch as the start

Treating the first 90 days as a distinct, structured phase — rather than a quiet tail-off after a launch announcement — is what lets an agent actually improve past its launch-day state instead of staying frozen at whatever it happened to be on day one. Most of the meaningful improvement we've seen across every build we've shipped happened in this window, not before it.

What we measure, specifically

The 90-day review isn't a subjective check-in. We track resolution rate without escalation, average time to resolution, and how those numbers trend week over week rather than looking only at a single snapshot. Where the client had baseline support metrics before launch — average handle time, tickets per agent, first-response time — we compare directly against them, because a percentage improvement means more to a client's leadership than a description of how things "feel" better.

We share this review with the client in full, including the parts that didn't go as well as hoped. An honest 90-day review, warts included, is what makes the next conversation about expanding scope a real decision based on evidence, rather than an assumption that automation is automatically working just because it launched.

Want to know what launch really looks like?

Ask us about the first 90 days of any build we've shipped — we're happy to walk through it.