Schedule Risk

Choosing Monte Carlo uncertainty ranges when you have no data

Published 6 August 2026 · 8 minute read

Start with the honest position: almost nobody running schedule risk analysis has calibrated historical duration data. The ranges get invented in a workshop, by people with opinions and a flight to catch. That is not a scandal, it is the normal condition of the trade, and pretending otherwise is how risk analysis loses its audience. The engineers in the room know the inputs were made up on a Tuesday, so when the output is presented as measurement rather than structured judgement, they stop listening. The useful question is not "where is the data" but "how do we make invented ranges defensible", and there is a reasonable amount of craft in the answer.

The uniform spread that says nothing

The most common shortcut is a blanket percentage: minus 10, plus 25, applied to every remaining task. It produces an S-curve that looks scientific and carries no information, because it encodes exactly one assumption, that all work is equally uncertain, and that assumption is false on every project ever run. A concrete pour with a crew that has done forty of them does not share a risk profile with a development approval sitting with a council.

There is a second problem, and it is mathematical rather than philosophical. On a network of any size the result is dominated by the handful of tasks that genuinely deserve wide ranges, the long-lead procurement, the approvals, the weather-exposed fieldwork. Smearing a uniform spread across three hundred well-understood tasks adds noise that partially cancels itself out, while the ten tasks that will actually move the completion date get the same timid range as everything else. You have spent effort making the model bigger and made it less informative. The fix is to spend the workshop's limited attention where the variance lives, and accept deterministic durations, or narrow ranges, on the work everyone genuinely understands.

Reference classes by work type

You do not have project-specific data, but you are not starting from zero either. Different classes of work have recognisably different tail shapes, and any room of experienced practitioners can rank them. Procurement of engineered equipment overruns through vendor slippage and shipping, and its tail is long but bounded by escalation: someone eventually gets on a plane. Statutory approvals have the nastiest tails in most programmes because the duration is not under anyone in the room's control, and a rejection restarts the clock rather than extending it. Repetitive fieldwork with an established crew is the tightest class you have, its variance driven by weather and access rather than by the work itself. Design and engineering sits in between, prone to rework loops that look like 20 or 30 percent overruns rather than doublings.

So instead of asking the workshop to range three hundred tasks, agree half a dozen work-type classes, argue properly about the shape of each one, and apply the class ranges as the default. Then interview only the exceptions. This is faster, more consistent, and much easier to defend in review, because the reviewer can attack six documented assumptions instead of three hundred undocumented ones. The counterargument is real: classes flatten genuine task-level knowledge, and a scheduler who knows a specific vendor is in trouble should override the class. That is the point. Classes are the default, not the ceiling.

Asymmetry is the default, not the exception

Durations overrun more than they underrun. A ten-day task can finish three days early at best, and there is rarely a reason for anyone to make it do so, while the list of things that can make it finish twenty days late is effectively unbounded. Parkinson takes care of the left side of the distribution and physics takes care of the right. A symmetric range around the planned duration is therefore a claim that the plan is unbiased, which contradicts everything the industry knows about its own estimates. Right-skewed distributions, a triangular with a long right tail being the workhorse, should be the starting shape, and a symmetric range should require justification rather than the other way around.

Watch the interaction with task length. The old 44-day rule of thumb exists because long tasks hide their own variance: a 120-day "detailed design" activity is really thirty smaller pieces of work whose individual overruns are invisible inside the bar. Ranging that bar at plus or minus 15 percent understates it, because the estimator mentally averaged the pieces. Where long tasks survive into the risk model, either break them down or widen the range beyond what feels comfortable, since the comfort is exactly the problem.

Ask for the story, never the number

The interview technique matters more than the distribution choice. Ask a construction manager "what's the P90 on this?" and you get the planned duration plus a socially acceptable margin, anchored on the number already in the schedule. Ask instead "tell me what would have to go wrong for this to take twice as long" and you get a story: the crane booking, the survey resubmission, the wet season window. Then price the story. If the events in it are plausible, the range must accommodate them, and the interviewee has argued for the wide range themselves, which is worth a great deal when the P80 lands on a director's desk.

Sequence matters too: get the extremes before anyone says the most-likely number out loud, because once the planned duration is spoken it anchors every estimate that follows. This is standard elicitation practice and it is routinely skipped because workshops run out of time, which is another argument for ranging classes rather than tasks.

Correlation is what actually moves the curve

Here is the part that gets lost in arguments about triangles versus betas: on a large network, independent task variation mostly cancels. One task runs long, another runs short, and the central limit theorem quietly narrows your S-curve into a confident sliver. If your P10 to P90 spread on a two-year programme is three weeks, this is almost certainly what happened, and the model is wrong in the dangerous direction.

Real projects do not fail through independent noise. They fail through common causes: the wet season that hits every earthworks task at once, the approvals authority that slows every consent, the labour market that stretches every crew-constrained duration together. When those factors bite, dozens of tasks slip in the same direction in the same iteration, and the tails of the distribution fatten dramatically. Modelling this does not require precision, correlating each work-type class against a shared driver factor at even a coarse 0.3 to 0.6 is a far better model than independence, which is a strong assumption disguised as a default. If you take one thing from this piece: a crude correlation structure with rough ranges beats exquisite ranges with no correlation, every time.

What to report, and what not to

Given invented inputs, the absolute dates in the output deserve modest confidence. Three things in the output deserve rather more, because they are robust to the exact numbers chosen:

And the failure mode to refuse: precision theatre. A P80 of 14 March quoted from inputs that were argued over a sandwich implies a calibration nobody in the room possesses, and sophisticated audiences discount the whole exercise the moment they see it. Report to the week. State the assumptions on the same page as the curve. The credibility you give up in false precision comes back doubled in trust.

Where PathProof fits

PathProof's Monte Carlo engine runs on Microsoft Project and Primavera P6 files locally, with per-phase uncertainty presets applied by work type, so the reference-class approach above is the default workflow rather than a spreadsheet exercise. The output is an S-curve with P50, P80 and P90 markers and a tornado of drivers, and because a run takes seconds, re-running on every monthly update makes the drift visible as a trend. The beta is free while we refine it with working schedulers.