Velocity is one of the few numbers in delivery that looks like a productivity measure. It has a unit, it goes up and down, and it can be charted, which is enough for it to end up in a board pack about four months after a team starts using it. Almost all of the damage velocity does comes from that resemblance, and almost all of its value survives as long as the number stays where it was produced.
The short answer to what it should be used for: forecasting, by the team that generated it, for its own near-term planning, expressed as a range. Everything beyond that runs into trouble quickly.
Velocity is an observation of how much a particular team, working on a particular kind of work, in a particular organisation, has completed per cycle in the recent past. The unit is whatever that team uses for relative size, and its meaning is entirely internal. A team's eight-point item is comparable with its own three-point item and with nothing else anywhere, because the scale was calibrated by those people against their own earlier work.
That is not a weakness in the measure. It is the source of its usefulness. A relative scale can be applied quickly and consistently by a group who share context, which is why estimation in those units takes minutes where an absolute estimate takes hours and is no more accurate. The scale works precisely because it never has to mean anything outside the room.
Forecasting a window. Take the last several cycles, use the range and not the average, and you can say something honest about how much of a backlog is likely to be finished by a given date. The Schedule Performance Domain of the PMBOK® Guide Eighth Edition, at Section 2.3, is concerned with producing schedule information people can rely on, and an empirical range from the actual team, stated as a range, is usually more reliable than a precise date derived from a plan nobody has tested.
Noticing that something has changed. A sustained shift in the number is a signal worth investigating, and the investigation is the point. A drop usually means the work changed character, people were pulled away, the environment broke, or something is being carried that nobody has named. The number does not tell you which. It tells you to go and ask, and asking within a cycle is considerably cheaper than finding out at the end of a quarter.
Estimates inflate. This is the first and most reliable effect, and it requires no dishonesty at all. When the number matters, the same piece of work quietly becomes a five where it used to be a three, and the chart rises beautifully while precisely the same amount of work leaves the team each fortnight.
Work gets split to be counted. Items are broken into pieces because pieces score, and the pieces stop corresponding to anything a user would recognise as finished. A backlog of forty fragments, each individually complete and none of them usable, is a familiar sight in teams whose throughput is being watched.
Quality moves to where it is not counted. Refactoring, test coverage, documentation and the small maintenance work that keeps a system habitable produce no points, so under pressure they stop happening, and the cost arrives three quarters later as a system that is slow to change. The velocity chart records none of this and will look strong throughout.
Unknown work gets avoided. A team measured on points will, quite rationally, prefer items it understands to items it does not, which means the genuinely uncertain work, usually the work that matters most, drifts to the bottom of the list and stays there.
The other misuse worth naming is comparison. Two teams' numbers cannot be compared, because the units were never calibrated against each other, and a manager comparing them is comparing two different currencies without an exchange rate. Where that comparison enters performance conversations, the scales adjust within about two cycles and the measure is gone for good.
An NHS trust was rolling out electronic observations ward by ward. The delivery team of seven had eight cycles of history and a velocity that ranged between twenty-four and forty-one, averaging around thirty-two. Asked by the programme board for a completion date, the team divided the remaining backlog by the average and gave one.
That date left the team and became an operational plan. Ward-by-ward go-live dates were set from it, training was booked, backfill was rostered, and a four-bed bay was scheduled to close for a fortnight of cabling and device installation. None of those commitments were unreasonable. Each was somebody doing their job with the best number available.
Two cycles later the team delivered nineteen points, then twenty-two. Two of the seven had been pulled onto a major incident in another system, which was the correct call and nobody disputed it. The bay had already closed. It stood empty for three weeks beyond its fortnight, at a cost measured in cancelled admissions, and the ward manager's view of the programme did not recover for a year.
What changed afterwards was the shape of the answer, not the arithmetic. Forecasts went to the board as a range with the assumptions attached: at the lower end of recent cycles this date, at the upper end this one, on the basis that the team stays whole. Ward closures were committed only two cycles ahead, where the forecast was firm enough to bear an operational decision. The board stopped asking whether velocity could be improved, which had been the most damaging question in the programme, because the honest answer to it was always yes and the honest method was always to write bigger numbers on the cards.
For a PMP® candidate, what matters is that an empirical measure describes the team that produced it, so a situation where velocity is being compared or targeted is describing a measurement problem before it is describing a delivery problem. A response that accepts the number as a performance statement has accepted a unit that has no meaning outside one team. A structured PMP exam preparation course keeps forecasting situations in play where the number is sound and the use of it is not.
The practical check on a live project takes ten seconds. Ask where your team's velocity figure is currently reported, and who reads it there. If it appears anywhere that funding, comparison or appraisal happens, its useful life is already short, and the honest response is to replace it in that forum with a forecast range and let the unit go back to being an internal planning tool.
Forecasting well is a judgement about how much certainty a number can carry before somebody builds an operational commitment on it. Omega's PMP® Exam Preparation works through schedule forecasting where the estimate is sound and the use made of it is not.
Empirical forecasting and the wider schedule work around it are covered in the PMBOK® Guide Eighth Edition.
Ad · Amazon affiliate link.
A110: The PMBOK 8 Schedule Performance Domain: What It Really Covers
A121: When a Schedule Baseline Needs to Change
A120: Schedule Compression: Fast Tracking vs Crashing
A113: Float Explained: How Much Delay Can You Really Absorb?
A115: Agile Estimating vs Predictive Estimating
PMP and PMBOK are registered marks of the Project Management Institute, Inc.