Monte Carlo vs velocity: how to forecast a delivery date you can defend
A single-date forecast from a velocity average is a promise you will almost certainly break. A probabilistic forecast gives you a range and the odds. Here is the difference — and why leaders trust the range.
Somebody asks the question every quarter: "So when will it ship?" And the reflexive answer — take the team’s average velocity, divide it into the remaining backlog, count the sprints, read off a date — feels rigorous because it involves arithmetic. It is not. It is a single point pulled from a cloud of uncertainty and handed over as if it were a measurement. The date has, realistically, a near-zero chance of being exactly right, and when it slips, the credibility cost lands on you. There is a better way to answer the question, and it starts by admitting the question has a distribution for an answer, not a number.
Why the single-date velocity forecast breaks
The velocity-divided-by-backlog method makes three assumptions that are all false in practice. It assumes your velocity is stable — but it wobbles sprint to sprint. It assumes the backlog is fixed — but scope grows and shifts. And it assumes a hard end date is a useful thing to promise about work months away — when the honest truth is a range. As a rough sanity check on a two-week horizon, velocity is fine. As a commitment about a quarter-distant release, it is false precision dressed as data, and it fails at exactly the moment it matters most.
A single date says "trust me." A range with a probability says "here is what I actually know, and how sure I am." Only one of those survives contact with a steering committee.
What Monte Carlo actually does
A Monte Carlo forecast does not try to predict the future. It simulates thousands of plausible futures. It reaches into your team’s real, historical throughput — the actual number of items finished in each of the last several dozen periods — and replays the remaining work against those samples, over and over, with the natural variability baked in. Some runs get lucky and finish early; some hit a bad streak and run long. Count them up and you get not a date but a curve: an 60% chance of shipping by this date, an 85% chance by that one, a 95% by a later one.
The magic is not the math — it is the honesty. Instead of hiding uncertainty behind a confident single number, the forecast puts the uncertainty on the table where a stakeholder can make a real decision with it. "We can commit to the July date at 85% confidence, or the June date at 55% — which risk do you want to take?" is a fundamentally more mature conversation than "we said June," and it is one you can only have with a range.
The sprint-level version: probability of missing
The same philosophy scales down to a single active sprint. Rather than a straight-line "ideal" burndown that assumes the last point is as easy as the first, a good sprint forecast models the S-curve real teams actually finish on, and outputs a calibrated probability of missing the commitment — with a confidence band, not a false-precision point. Early in a sprint, or on a new team with thin history, an honest forecast says so and widens the band, instead of inventing certainty it has not earned.
- A probability, not a color — "78% chance of missing committed scope," updated as the sprint moves.
- A confidence band — a range that visibly narrows as the evidence grows, so you can see how much the model actually knows.
- Honest about thin data — new team or early sprint? The forecast leans on a prior and tells you it is doing so.
- Denominated in delivery, not points — what will land, by when, and the cheapest scope to cut to change it.
A forecast is only worth as much as its track record
The catch with any probabilistic forecast is that a probability is unfalsifiable on a single event — you cannot check "70%" against one sprint. You check it across many. So the forecast has to keep score of itself: of every sprint it called at 70%+ risk, how many actually missed? KalMatrix back-tests its forecasts against your own closed sprints and publishes that calibration next to the number, so the range you are asked to trust comes with its own report card. That willingness to be graded is the whole difference between a forecast and a guess — and it is the reason a leader will act on it two days early instead of confirming it two days late.
Velocity is not the enemy; false precision is. Keep velocity for the quick sanity check, and when the stakes are real — a release date, a customer commitment, a board update — trade the single confident number for a range with the odds attached. It is the difference between a forecast that impresses in the meeting and one that is still right when the date arrives. If you are still turning a points average into a date, start with why "story points can’t forecast a date."
The demo is a fully-populated workspace — real sprints, real forecasts, the same daily diagnosis this post describes. No signup, nothing to configure.