Release day. The community launch was an hour old, traffic was light, and the vendor’s API was already refusing us.
Rate limit errors. Not even the standard kind: the API answered with 409s, a conflict status doing a rate limit’s job. Which was fitting, in a way nobody was in the mood to appreciate that morning, because everything about that integration was one thing doing another thing’s job.
The traffic that broke it was a small fraction of our user base. Not a spike. Not a viral moment. A quiet fraction of our users, doing ordinary things, was more workload than the platform we had spent months integrating could survive.
And the strangest thing about the failure was how unsurprising it was. Engineers had asked, months earlier, to check for exactly this kind of thing. The only person surprised was the person who had made the decision.
A sound decision, evaluated on look and feel
Rewind a couple of months. The capability was a community layer: discussions, member profiles, the social surface around the core product. It was not core to what the company sold, and the reasoning went the way that reasoning should go. A specialist vendor would do this better than we would. Buy it, integrate it, keep the engineers on the product that makes the money.
I want to be fair to that logic, because it was correct. If this story has a villain, it is not the decision to buy.
The vendor search was a solo job. Our VP of Product ran it alone. He was not a technical person, and there is no crime in that either. But it meant the evaluation ran on the evidence he could personally read, and the evidence he could personally read was the vendor’s interface. It looked polished. It felt right. He liked it, and he wanted it now.
What never happened was the other half of the evaluation. Nobody in engineering was asked to read the technical documentation. Nobody reviewed the API we were about to couple our platform to. We asked to run a short technical investigation of the vendor before the commitment was made, the kind of check that takes days, not months. The request went nowhere. Not refused, exactly. Just never taken up, the way a suggestion dies when the decision it questions has already been made.
The demo answered every question a demo can answer. Look. Feel. Features. Nobody asked the questions release day was going to ask.
What the delivery evidence showed
For a couple of months, the integration went fine, which is the treacherous part. Endpoints responded. The test environment behaved. Integration traffic is polite: one developer, a handful of test accounts, requests arriving one at a time. Nothing about it resembles what your full user base does at once, and everybody technical knows that, in the way you know things you have not been asked to write down.
So the work advanced, the status was green, and the assumption underneath the whole initiative had still never been tested or even stated: the vendor has to survive our traffic.
Release day tested it. Here is what the evidence showed, once we finally looked.
The vendor’s API could not sustain traffic from a small fraction of our users. The rate limits we were hitting had never appeared in any evaluation, any contract conversation, or any integration plan. And the vendor, it turned out, was a very small company. An engineering team you could count on one hand, and infrastructure sized to match. They had not deceived anyone about that. Nobody had asked.
We say the missing requirements were discovered after delivery. Discovered is a generous word. Throughput, error behavior, rollback, support capacity: those were requirements from the day the contract was signed. They did not appear on release day. They just stopped being optional.
A warning with no record is trivia
The gap here was not knowledge. Any engineer on the team could have told you that APIs have rate limits and that small vendors have small ones. Several of us had said as much when we asked for the investigation.
The gap was that none of it existed as a record. The spec, as far as the decision was concerned, was the demo. The operational requirements lived in the heads of people who had not been consulted, and the one attempt to move them into the decision, the investigation request, was made in conversation and died in conversation.
That is why the VP’s surprise on release day was genuine, and the genuineness is the instructive part. He had not suppressed a warning. From where he stood, there had never been one. A concern raised verbally in a meeting months earlier had never entered anything he would encounter again: no risk item, no decision record, no written objection attached to the commitment. It could be forgotten without malice, so it was.
A warning that lives only in a meeting is not a warning. It is trivia.
I am not excusing the decision. Choosing a platform vendor on look and feel, without letting anyone technical near the evaluation, was a bad process, and it stayed a bad process even before it produced a bad outcome. But the engineers, and I include myself here, treated saying it once as having handled it. We raised the risk in the wrong medium, watched it go nowhere, and then integrated for months without ever making it awkward again. Being right in a meeting is not a control. It is a consolation prize.
Who paid for it
The engineers paid first and most visibly. Release day became firefighting, and the weeks after it became a coordination grind.
The vendor paid too. A very small company suddenly had a client an order of magnitude beyond its infrastructure, hammering it with traffic it could not absorb, on the most public possible day. The pattern was consistent with a product sold one client size above its infrastructure. That is not negligence. It is a mismatch that a single technical conversation before signature would have surfaced, to everyone’s benefit, including theirs.
The VP paid in credibility, in the slow way executives do when a confident decision fails in public and the people who warned about it are the ones cleaning it up.
And the company paid twice for one capability. Once for the vendor, and once for the engineering the vendor was supposed to make unnecessary.
A very dumb load balancer
Because here is what fixing it actually meant. The vendor could not scale their side, not on any timeline that mattered to our launch. So we fixed it from our side of the API.
Application caching, so repeated reads never reached the vendor at all. A queue, and I use the word generously, that held requests and released them gradually instead of letting our users’ clicks hit the vendor directly. Counters tracking how many requests per second we were sending, so we could stay under a ceiling we had discovered empirically, in production, on release day. Piece by piece, we built a very dumb load balancer, whose entire purpose was to protect our vendor from our users.
Sit with that inversion for a second. The reason you buy instead of build is that the specialist absorbs the hard operational problems for you. We ended up paying the specialist and building the operational layer, which is both purchases at once.
By the time it was stable, the buy had cost more than building the capability ourselves would have. That is an estimate. Nobody ever measured it formally, because measuring it would have required records of what the recovery consumed, and record keeping was not that initiative’s strong suit at any stage. But the shape of the cost was unmistakable, and most of it was not code. It was meetings. Meetings with the vendor about capacity. Meetings with the VP about the meetings with the vendor. Meetings where engineers explained rate limits to people who had signed a contract governed by them.
We spent more time in rooms explaining the problem than in editors fixing it.
How to see it coming
A buy decision heading for this failure gives off signals well before release day:
- A vendor commitment is being made and no engineer has read the API documentation.
- The evaluation record contains what the product looks like and nothing about how it behaves under pressure.
- Nobody has multiplied your user count against the vendor’s documented limits, or noticed that no limits are documented.
- A technical objection was raised in conversation, went nowhere, and was never converted into a written risk.
- Integration traffic looks nothing like release traffic, everyone technical knows it, and the launch plan does not care.
Any one of these is survivable. Two or more mean the operational half of your spec is being written by luck, and luck publishes its findings on release day.
The release day requirements review
Here is the durable version: the checklist I wish had been stapled to that contract before anyone signed it. Before any buy instead of build commitment, make someone answer the second half of the spec, in writing.
1. Throughput
What sustained and peak request rates will the vendor commit to, in writing, against your actual user numbers?
2. Error contract
What does the API do under pressure? Which status codes, what retry guidance, what backoff?
3. Observability
How will you know the integration is degrading before your users do?
4. Rollback
If it fails on release day, what is the path back, and how long does it take?
5. Support capacity
Who fixes an outage on the vendor’s side, and how many of them are there?
6. Evidence
Has anyone technical read the documentation and pushed realistic load at a sandbox?
None of this is exotic, and none of it insults the vendor. A good vendor answers these in a day, and the answers become part of the commitment. The point of writing them down is not bureaucracy. Written requirements have standing. They are what a warning becomes when it stops being trivia, and they are what the release day conversation points back to when someone asks how this was allowed to happen.
The Scopeworth lesson
The expensive part of this story was never the rate limit. Rate limits are ordinary. The expensive part was that the decision could not answer basic questions at any point in its life. What was promised? Nothing, in writing, about anything operational. What was checked? A demo. Who raised what, and when? Memory says the engineers did. Nothing else says anything. What did the failure cost? Nobody knows, because the recovery was never measured either.
That is an evidence gap wearing a vendor problem as a costume. The requirements were missing from the record, so the production incident became the first written spec, and the invoice for it was paid in engineering time and meetings that nobody ever added up.
Scopeworth is being built to turn fragmented delivery signals into client-ready evidence: what was promised, what changed, what shipped, and what it cost, connected closely enough to survive a hard question. It is early, and honest about that. But the principle behind it is the one this story left with us:
A requirement you have not written down is still a requirement. You are just choosing to discover it in production.
Next time buy instead of build is on the table, ask for the second half of the spec before the signature: throughput, errors, rollback, support. And ask the diagnostic question about your last vendor commitment: could your team reconstruct, today, what was actually checked before it was signed, without relying on memory?
