All Insights

The Impossible Quarterly Review

A client asked one simple question — what did we ship this quarter, and what did it cost? Every tool gave a different answer. Here is why, and how to make the next quarter answerable.

Flat evidence map titled The Impossible Quarterly Review, showing one executive question and four delivery records — Committed, Changed, Shipped, Costed — each returning a different number across a misaligned quarter boundary before converging into a single reconciled quarterly account.
Four honest systems, four different answers, and one question none of them was built to answer.

This story combines patterns from multiple delivery situations. Details have been changed, merged, or omitted to protect confidentiality.

The question that broke the composite quarter was not a hard one.

We were an hour into a quarterly review with one of our larger clients. The slides were going fine. Then their operations lead put down her pen, leaned back, and asked the thing that everyone assumes is easy.

Before we get into next quarter — just walk me through what we actually shipped this quarter, and what it cost us.

Reasonable. Completely reasonable. The kind of question a client is entitled to ask, and the kind a delivery team should be able to answer in its sleep.

I said I would pull it together and send it over.

I thought it would take an hour.

It took days. And even then, the answer I handed over was one I could not fully defend.

We had every tool. We did not have an answer

Here is the part that still bothers me. We were not a chaotic team. We had the tools. All of them.

We had a work tracker with every ticket, every sprint, every board. We had the full code history. We had a deployment pipeline that logged every release. We had time tracking. We had the invoices. We had months of chat and a folder of decision documents. If you had asked me the morning before that meeting whether we could account for our delivery, I would have said yes without blinking.

We had confused having tools with having a record.

Those are not the same thing, and I did not understand the difference until I sat down to answer one simple question and watched every system I trusted give me a different number.

Every tool was right. None of them agreed

I started with the work tracker, because that is where you start.

It told me a story. A confident-looking story, right up until I read it properly. A chunk of what was “done” this quarter had actually started last quarter. Some items had been carried over the other way and were sitting half-finished. Tickets had been renamed as we understood the work better. A few epics had been split into follow-up items. One had been reopened twice. If I counted “completed this quarter,” I got one number. If I counted “meaningfully worked on this quarter,” I got a very different one.

So I went to the code history to get something more solid. It showed sustained, obvious work — plenty of it. It also showed refactors, a couple of reverts, and a run of follow-up fixes that did not map cleanly onto any ticket I could point a client at. The code could tell me a lot changed. It could not tell me which of it was the thing we had promised.

Then the deployment log. That at least should be objective: this reached production, on this day. It was objective. It was also on a cadence that had nothing to do with the client’s quarter. Work built late in the period did not deploy until the next one. Some of what deployed this quarter had been built in the previous one.

Then finance, for the “what did it cost” half. The invoices ran on a billing cycle that did not match the sprint boundary, which did not match the release boundary, which did not match the quarter the client had in her head.

Four systems. Four honest answers. Not one of them answered the question she actually asked.

None of them was lying. None of them was enough.

The quarter was not unaccountable because we did little

This is the part I want to be careful about, because it is the part that is easy to get wrong.

The quarter was not unaccountable because the team had slacked off. We had worked hard and shipped real things. The problem was that nothing in our operating setup agreed on where the quarter began and ended, or on what “shipped” even meant.

“Done” in the tracker meant a developer had finished. “Released” in the pipeline meant it had reached production. “Delivered” in the report meant the client could use it. “Billed” in finance meant time had been invoiced. Those four words pointed at four different moments, sometimes weeks apart, and we had quietly assumed they were the same event.

On top of that, the boundary itself leaked. Work in flight across the quarter line belonged partly to two quarters and cleanly to neither. There is no honest way to count a half-finished migration as either fully this quarter’s or fully last quarter’s, and we had dozens of those.

And underneath everything, the thing I least wanted to admit: I could no longer separate the work we had planned from the work that had simply landed on us. A production incident and a roadmap feature leave almost identical footprints in the tools three months later. By quarter close, I genuinely could not tell you how much of our capacity the firefighting had eaten, because nothing had marked it as firefighting when it happened.

Who paid for it

Days later I had an answer. It was, I am fairly sure, roughly correct.

But look at what it had already cost.

I had sent nothing for days on a question the client thought would take a quick lookup. That silence says something on its own. When the account finally arrived, it arrived after the doubt had already formed — and evidence that shows up after the doubt always looks like a defence, even when it is just the truth.

The client did not accuse us of anything. She did not have to. The scope had grown over the quarter through a dozen reasonable mid-quarter “could you also…” conversations, none of which we had re-baselined against the original commitment. So expanded scope looked exactly like original scope, which meant the extra work looked like slowness, which meant the bill looked high for what was “planned.” I had no clean way to show her the difference, because we had never recorded it.

I spent that review defending a legitimate quarter with numbers I had reconstructed under pressure. We kept the client. But I had turned a routine review into a small trial, and I was the one who had built the courtroom by never keeping the receipts.

What I actually did

The reconstruction itself was not clever. It was just slow, and it was the sort of thing that should never be done in a hurry before it is due.

Before the composite client reconstruction, there was the temporary reconciliation application: a small tool that used the available tracker context and change timestamps to assemble an estimate. The confidence meter was not decorative. It was a visible admission that ambiguous categories and missing effort records could not produce a pristine account. That tool then needed storage, deployment instructions, and maintenance. In other words, the reporting gap had produced more reporting surface.

I took each material initiative and pinned it to one boundary — the client’s quarter — and one definition of “shipped,” which we agreed meant “the client could use it in production.” Then, per initiative, I lined up what we had committed to, what actually changed, what genuinely shipped inside the boundary, and what it cost against the same boundary. I flagged everything that straddled the quarter line and split it honestly. And I went back through the chat history to find every scope change we had accepted and never re-baselined, so I could finally separate “the plan” from “the plan plus everything you asked for in March.”

It worked. It produced an account I could stand behind.

It also should have been a standing record I could export in an afternoon, not an archaeology project. Everything I reconstructed by hand had existed all along, scattered across five tools. What was missing was the thin layer that tied it together while the work was still fresh.

How to see it coming

You do not need my bad reporting stretch to know whether you are exposed. The warning signs are quiet and specific:

  • Someone asks “what did we ship last quarter?” and your honest first move is to open four tools.
  • “Done,” “released,” “delivered,” and “billed” are used interchangeably in your reporting, and nobody can say which moment each one means.
  • Your invoice period, your sprint boundary, and your release cadence are three different calendars.
  • You cannot point to where a mid-quarter scope change was re-baselined, because it never was.
  • You cannot separate the quarter’s planned work from its unplanned recovery work without relying on memory.

If more than one of those is true, your next quarterly review is already going to be harder than it should be. The time to fix it is now, while the work is fresh — not immediately before the review.

The four-record quarterly reconciliation

Here is the practical version. It is boring, which is exactly why it works.

Answer the one question — what did we ship this quarter, and what did it cost? — with four records reconciled to one boundary and one definition of “done.”

1. Committed

What did we agree to deliver this period, and how did that commitment change? Record the original commitment and every re-baselined scope change.

2. Changed

What implementation and recovery work actually happened, and how much of it was unplanned?

3. Shipped

What reached a genuinely usable state inside the period boundary — not “code complete,” not “deployed to staging,” but usable?

4. Costed

What did the period cost, on the same boundary as the delivery?

Then, per material initiative, write one short account: the original and re-baselined commitment; planned delivery versus unplanned recovery; what actually shipped and reached a usable state; the cost on the same boundary; anything in flight across the line and how you apportioned it; the defects and cleanup still open; and how confident you are.

One more rule, learned the hard way: do not reach for story points, commit counts, hours logged, or ticket totals to answer this. They can flag a contradiction worth investigating, but they cannot tell a client what they got for their money. A big number is not an account. It is just a bigger thing to explain.

The Scopeworth lesson

The uncomfortable truth of that quarter was not that we lacked data. We were drowning in it. We lacked a chain that connected what we promised, what we changed, what actually shipped, and what it cost — closely enough, and on the same calendar, to answer one fair question without an improvised dig.

Scopeworth is being built to turn fragmented delivery signals into client-ready evidence. Not to watch developers, not to score activity, and not — to be clear — to have magically rescued that review for me. It is early, and it is honest about that. The point of the product is the same as the point of this story: a delivery record you keep while the work is fresh, so the quarterly account is something you run, not something you survive.

Because if answering “what did we ship last quarter, and what did it cost” needs every tool open and an improvised reconstruction, you do not have a delivery record.

You have delivery exhaust.

Could you answer that question today — what shipped this quarter, and what it cost — without opening four tools and clearing your calendar?

Sample report

See what client-ready proof looks like.

Download an anonymised example SDLC ROI report — what shipped, what changed, what it cost, and how confident the data is.

Book a discovery call