Technology leaders often think of engineering systems like household plumbing: an overlooked system hidden inside, which only gets attention when a pipe bursts. But this thinking can be very costly. Whether a software as a service (SaaS) product collapses under its own weight or scales rapidly is not decided by some brilliant architecture drawn on a whiteboard. The real difference comes down to three things: how fast your team learns, how quickly it catches failures, and whether your users are willing to keep up with constantly changing software. Mike L. Swafford, who has led engineering systems teams at Microsoft, spent years untangling these very knots. He has seen the era when software shipped in boxes after three years, and now he is also witnessing the fast-paced era of cloud services. His straightforward opinion is: the teams that win are the ones that make feedback loops their top priority and put everything else after it.

Feedback Loops Are The Product Behind The Product

Ask Swafford what the secret to a product operating at scale is, and he doesn’t start a long, winding debate on architecture. He goes straight to loops: an automated system that catches errors on the spot and fully validates the user experience; telemetry that tells you what is happening inside the live system; and the ability to correlate all of this with release and testing, so that the resulting data does not become mere decoration, but helps in making concrete decisions. He says: “That is the telemetry that you have coming from the system as it runs so that you have informed understanding of what’s going on the system, and that you can correlate that telemetry with things like rolling out new builds or experiments, so you have real actionable data on understanding the changes that you’re making.”

The real point is what most monitoring dashboards miss. All lights being green on a dashboard is no proof that the user is happy. Swafford looks for signals in customer behavior. Are people clicking on the help button more often? Are they visiting the system-down page? This can be a signal that something is slipping through your telemetry. The goal is to catch the flaw before the user notices it. Safety nets should be laid beneath the system. But mistakes will happen anyway. The loop is only completed when learning from the failure leads to permanent changes in the tools. Swafford genuinely believes in this principle: never waste a good crisis. An investigation after which the system does not change remains merely paperwork.

What Got You Here Will Not Get You There

The transition from the old software model to continuous delivery is often considered just an engineering task. But according to Swafford, this is an inversion of fundamental business principles, which is quite difficult to manage. Delivering a new update after three years is one thing, and providing a continuous service is an entirely different game. He explains: “Box software is about a set of amazing new capabilities delivered on a slow cadence, whereas services is about a continuous stream of innovation.” The biggest illusion is assuming that the customer is always ready for change. The control over the schedule slips out of the customer’s hands. New things arrive automatically, and for organizations whose critical operations depend on it, feeling anxious and resisting is only natural.

Then this burden falls right back onto engineering, and that too in the form of requirements that teams often overlook – for instance, pausing changes when necessary, or giving the customer the right to adopt a new feature early or late. Swafford clarifies that every new feature should be controllable at the user and company level, so that users can turn it on at their discretion. Inside the team itself, the biggest bottleneck is testing speed. He calls it the habit of eating vegetables every day – that is, work that everyone agrees is good, but no one is willing to do. In the three-year era, slow tests could be run every two months. But now: “You can’t afford to do that if you’re continually shipping software. You need to be able to validate everything before you give it to your customers.” If delivery is to be sped up, the speed of testing must also increase; otherwise, your own code will become a wall in your path.

AI Raises The Price Of Weak Test Coverage

Agentic AI has radically changed the way code is written. Swafford’s experience over the past year shows where AI works magic and where it causes harm. When there are straightforward instructions and clear standards, AI works wonders, especially on brand-new code. Trouble arises when a simple sentence written in English is not understood by the model. Then teams have to split hairs and write even stricter tests to maintain the entire structure. The real headache is the legacy codebase, and the whole business rests on it. Swafford has blunt advice for teams about to try AI on legacy systems: “If you don’t have good test coverage for a section of the codebase, you probably should not point AI at it, because AI, if it doesn’t have tests to know whether it’s breaking something or not, it will break it.”

There is another subtle risk as well. When a test fails, an AI agent might alter the test itself instead of fixing the code. That means the system’s typical behavior will change, and a user feature will quietly break. The team must know which test validates real user workflows. Where Swafford has seen teams successfully rewrite code, they did it very methodically. They explained the existing app’s design and behavior to AI, provided reference code, and built rapid prototypes to throw away. Then they rebuilt each prototype in roughly 20 separate pieces so humans could keep an eye on it. In just three months, the app was made ready for desktop as well. But he is not in favor of erasing old code merely out of preference: “I am not a big fan of rewriting codebases for the sake of rewriting codebases.” This step should only be taken when the existing code completely gives up. Speaking of future threats, he warns: newer models are extremely fast at spotting security vulnerabilities. If your legacy code cannot change rapidly, it will prove an easy target for hackers who now possess these weapons.

Follow Mike L. Swafford on LinkedIn for more insights on engineering systems, cloud delivery, and building software at scale.