Back to Resources

Advisor Summary: A Monte Carlo simulation tests a retirement plan against randomized market and inflation futures and reports how often the money lasted. An 85% probability of success means 850 of 1,000 simulated futures ended with assets remaining, given the plan’s assumptions. The simulation method is mathematically sound but its use in financial planning is conceptually bankrupt. Its practical limits show up in what the score cannot say: is that score good or bad? how large would a shortfall be? when should a change be made? or what to do next? Those questions need a different tool.

Ask advisors or clients what they most want to know about Monte Carlo simulation, and the question is rarely about the math. It is some version of: “A client’s plan read 91% at retirement. Eighteen months of rough markets later, the re-run says 65%. What exactly am I supposed to do with this number?”

Monte Carlo simulation is the dominant analytical engine in retirement planning software. This article walks through what this common analytical method does, what a score like 85% means mathematically, where the framework runs out of road, and the course-correction the industry needs to deliver real client-centered planning.

What a Monte Carlo simulation actually does

A Monte Carlo simulation is a mathematical technique that answers a question by running the same model many times through different scenarios, seeing what happens in each. In retirement planning, the model is a client’s plan and the randomized input is varying market returns (though other things, like inflation and mortality can be varied as well).

The name of this technique comes from the casino in Monaco, on the Mediterranean coast. This name makes sense because the scenarios in the Monte Carlo simulation are typically produced through randomization. Nicholas Metropolis and Stanislaw Ulam, mathematicians at Los Alamos, formalized the technique for problems too complex to solve directly and published it in the Journal of the American Statistical Association (September 1949); Metropolis suggested the name, recalling in a 1987 Los Alamos Science retrospective that it was inspired by an uncle of Ulam’s who gambled at the Monte Carlo casino. It spread from physics to finance because retirement planning has the same shape: too many interacting uncertainties to solve with one equation.

Applied to a retirement plan, one simulation trial works like this:

  1. The software builds the plan: starting portfolio, asset allocation, planned spending, Social Security and pension income, and a planning horizon, commonly to age 90 or 95.
  2. For each year of the horizon, it draws a random portfolio return (and in the best software also a random inflation rate) from a statistical distribution calibrated to a set of assumptions.
  3. It applies that year’s withdrawals, income, and growth, then moves to the next year.

The software repeats the whole exercise, typically about 1,000 times. Each trial is one simulated financial future.

So far so good, but what does someone do with all of these simulated futures? Typically, they are interpreted using a single result: in how many scenarios was there still money left at the end? The famous output, the “probability of success,” is simply the fraction of those simulated futures in which the portfolio never hit zero; a different name for it would be the plan’s simulated survival rate. If 850 of 1,000 trials end with money remaining, the result reads 85%.

The mathematics of this machine are fine. But the implementation and interpretation of the results in financial planning have been anything but. The complications live in two places: the assumptions that feed it, and the way its output gets used.

Where the assumptions come from

Every Monte Carlo output is conditional: the “probability of success” score is a property of the plan plus a specific set of assumptions about how markets behave. Change the assumptions and the score changes, with no change in the client’s actual situation.

Input What it controls Where it typically comes from
Expected return The average growth rate of simulated portfolios Long-run historical data, asset manager capital market assumptions, or software defaults
Volatility (standard deviation) How widely each simulated year swings around that average Historical data, sometimes adjusted by the vendor
Correlations How asset classes move together within a simulated year Historical co-movement
Inflation How much purchasing power is eroded over time Historical averages or forward-looking estimates
Planning horizon How long the money must last Advisor-selected, commonly age 90 to 95
Spending path The withdrawals the plan must support The client’s plan itself

Expected returns are estimates, not facts. A simulation calibrated to 1926-to-present U.S. equity returns will produce meaningfully higher scores than one calibrated to a major asset manager’s forward-looking assumptions (since these assumptions are rarely higher than historical averages). Neither is wrong; they are answers to different questions, and the client sees only the final percentage, with no indication of which question was asked.

The distribution shape matters. Most engines draw each year’s return independently from a normal or lognormal distribution, like pulling a numbered ball from a jar and putting it back before the next draw. Real markets are not quite like that: extreme years happen more often than a normal curve implies, and bad periods can cluster, because recessions, bear markets, and inflation regimes persist. But reversion to the mean is also present in real life: if bad times have just happened, it’s typically then less likely that things will experience another downturn of the same size. Recovery is more likely.

The typical Monte Carlo method also pays no attention to where things are today. It’s true that no one knows if markets are at a high now or are going to continue to go up, but we do know it isn’t at a low (since markets have been going up). That information can be useful, but it isn’t included in most Monte Carlo simulations. Regime-based Monte Carlo methods address this by modeling markets as moving between distinct environments. The refinement matters most where retirees are most vulnerable, in extended bad stretches early in retirement; how Monte Carlo can understate retirement income risk compared to historical simulation examines this limit in depth.

Advisor takeaway: Monte Carlo results quoted as a probability of success make the whole thing seem simple. But the devil is in the details, and the simulations used in most Monte Carlo software can differ meaningfully from how markets really work. That doesn’t make Monte Carlo simulation useless. It just means we should not take results as accurate predictions of the future. They must be interpreted and taken with a grain (or a giant scoop) of salt.

What Monte Carlo does well

Monte Carlo fixed a real problem: straight-line projections. A projection that grows a portfolio at a constant 7% every year answers “what happens on average?” and silently assumes the average shows up on schedule. No one should confidently walk across a river after hearing it is, on average, three feet deep. Retirees do not live in averages. They live in one specific sequence of returns, and the order of those returns can matter nearly as much as their size.

This is sequence-of-returns risk, and a simple example shows it. Consider two retirees, each age 65, each starting with $1,000,000 and withdrawing $50,000 at the end of each year. Over three years, both experience the same three returns: one year at -20% and two years at +10%. Only the order differs.

Retiree A takes the loss first. The portfolio falls to $800,000, and the withdrawal leaves $750,000. Two +10% years later, after withdrawals, the balance is $802,500.

Retiree B takes the loss last. Two good years carry the balance to $1,105,000 after withdrawals; the -20% year then drops it to $884,000, and the withdrawal leaves $834,000.

Same returns, same average, same withdrawals: Retiree B ends $31,500 ahead, and the gap compounds from there, because Retiree A must fund every future withdrawal from a smaller base. A straight-line projection shows both retirees the same line.

Monte Carlo sees the difference directly. Because every trial shuffles the order of returns, the simulation naturally includes futures where the bad years land early, and those trials are disproportionately the ones that deplete. That was a genuine advance. The simulation also communicates one idea well: the future is a range, not a number, and a fanned-out set of 1,000 possible paths is more truthful than any single line on a chart.

The big failure: What an 85% score actually means

An 85% score means that 850 of the plan’s 1,000 simulated futures ended with portfolio assets remaining and 150 did not, under the plan’s fixed assumptions.

Take a concrete case. Paul and Diane, both 64, plan to retire next year with $1,100,000, combined Social Security of $4,600 a month beginning at 67, and planned spending of $8,600 a month. Their advisor’s software models the plan to age 95 and runs 1,000 trials. In 850 of them, the portfolio still holds assets at 95. In the other 150, it runs out somewhere before then. The plan reads 85%.

Mathematically, the claim is much narrower than it sounds. Strip “85% probability of success” into its parts and it makes four separate assertions at once: the futures were generated under one specific set of return, inflation, and longevity assumptions; spending was held rigid through every simulated crash and boom (no adjustments allowed); the only thing tested was whether the portfolio avoided full depletion by age 95; and 850 of the 1,000 futures cleared that bar. Every qualifier is load-bearing, and every one is invisible in the number itself, which is why the score and its plain-English unpacking feel so different.

Here are some of the questions the probability of success score doesn’t answer:

  1. How bad were the failures? Were they $1 short or $1,000,000?

  2. When did the failures happen? Early in the plan when I’m likely to feel them, or late when I might not even be around?

  3. What would have happened if I had adjusted instead of stubbornly staying the course?

  4. How big of an adjustment would I have needed in order to avoid running out of money?

  5. How long might I have needed to stick with an adjustment before returning to the original plan?

  6. In the “successful” scenarios, how much did I overshoot my end-of-plan goals? Did I have $1 left, or $100,000,000?

  7. How often could I have actually spent more than I have planned? How much more? For how long?

  8. How dependent are my results on my assumptions? What if I received better returns or inflation? What about worse?

  9. How wide is the dispersion of results? Do the simulations range from me being $10,000,000 in the hole at the end of the plan to $100,000,000 above? (If so, that doesn’t seem like a very useful simulation!)

To be precise about what the number is not:

The score tells you The score does not tell you
850 of 1,000 simulated futures ended with assets remaining How much was left in the 850, or how deep the hole was in the 150
The plan, held perfectly rigid, weathered most simulated markets When trouble would arrive, or what early warning looks like
Risk exists somewhere in the plan Which change, of what size, would address it
A frequency inside a model, conditional on assumptions A guarantee, or even a forecast, about the one future the client will actually live

Advisor takeaway: The term “probability of success” is deceivingly simple. In fact, it doesn’t tell you much, and it doesn’t answer basic client questions. It’s best to avoid using this misleading score entirely.

Where Monte Carlo breaks down in practice

The core failures of Monte Carlo aren’t in the math. They are in the interpretation and the use of “probability of success”. The success/failure framing usually used for delivering Monte Carlo results implies a world that doesn’t exist and leads to bad decision making and heightened client (and advisor) anxiety.

The endpoints are binary, but real outcomes are not

A simulated future that reaches age 95 with $1 remaining counts as a success. A future that ends with $4 million also counts as a success, weighted identically. On the other side, a future that depletes at 94 counts the same as one that depletes at 78. The score is blind to magnitude in both directions: two plans can carry the same 85% while one risks a small, late shortfall easily absorbed by a spending trim and the other risks running dry with fifteen years left. Those plans deserve very different conversations, and the metric cannot tell them apart.

The binary framing also buries a quieter problem: many of the 850 “successes” are futures in which the client dies with multiples of their starting wealth, having spent three decades declining trips and deferring generosity they could easily have afforded. The score books those futures as wins. Research by Derek Tharp and Justin Fitzpatrick on Kitces.com, The Retirement Distribution ‘Hatchet’: Using Risk-Based Guardrails To Project Sustainable Cash Flows, demonstrates the alternative: project the actual cash flows a plan with built-in adjustments would deliver over retirement, including the spending path those adjustments trace, and the binary survived-or-depleted tally stops being the headline at all.

The model assumes no one ever adjusts

A standard Monte Carlo run holds spending fixed, adjusted only for inflation, for thirty or more years, whatever the simulated markets do. That assumption is false for essentially every household. Real retirees adjust, and spending drifts downward on its own: according to David Blanchett, then head of retirement research at Morningstar Investment Management, in “Exploring the Retirement Consumption Puzzle” (Journal of Financial Planning, May 2014), real retiree spending declines by roughly 1% per year through much of retirement rather than tracking inflation upward. The simulation is stress-testing a robot that spends identically at 67 and 90 and never reacts to a bear market. The 150 depleted futures in Paul and Diane’s run are largely futures in which nobody ever touched the dial. In real life, someone would have.

This is the deepest limit of the framework. If the plan’s response to trouble is “we will adjust,” then the client needs answers to two different questions: how much can we spend now, and what specifically will we change, and when, if the future misbehaves?

The score only measures one thing

Mathematically, “probability of success” is the same as “chance of underspending.” A 100% probability of success aligns with a planned spending level that would have survived in every single imagined scenario. In other words, in 100% of those scenarios, the client could have spent more and still not run out of money! In other words, a 100% probability of success score should not be seen as a great plan. It should be seen as a suboptimal plan that leaves money on the table. It should be viewed as a plan that doesn’t help people live the best life they can. But, because of the word “success”, no one realized this.

A better approach would help people understand that spending decisions are about trade-offs. Spend more and get to enjoy more life now, but take on a higher risk of a future downward adjustment. Spend less now and gain more long-term stability, but maybe give up on life experiences you actually can afford (and regret that later in life).

The score is a communication hazard

The percentage format invites clients to hear things the math never said. “Success” and “failure” are loaded words for a number that only counts portfolio depletion; a client who hears “15% chance of failure” hears a threat to their retirement, not a conditional frequency inside a model that never let them adjust. And when the number moves for reasons a client did not cause, the natural reading is that someone made a mistake. These are different problems from the mathematical ones above; why probability of success can mislead advisors and clients takes them up in full.

“What do I do when it drops to 85%?”

Nothing in the score itself tells you: a probability is a reading, not a plan of action.

It may be the single most common question advisors raise when Monte Carlo comes up in webinars and study groups, and it is the question that exposes the gap.

Return to Paul and Diane. At retirement the plan read 91%. Eighteen months of poor markets later, it reads 85%. The advisor’s realistic options look like this:

  • Reassure and hold. “85% is still strong; we stay the course.” Often analytically defensible, but expensive over time: a framework that answers every decline with “no action” teaches the client that the plan never actually calls for anything, and it leads to the client questioning whether the advisor really knows what they are doing.
  • Cut spending. The score says nothing about how much, to what number, or for how long, and a cut deep enough to move the score meaningfully is usually far deeper than conditions warrant.
  • Revisit the assumptions. Sometimes legitimate, but the metric quietly invites tuning inputs until the output looks better. The client’s situation has not improved because an expected return assumption moved up half a point.
  • Re-anchor the conversation on a framework built for adjustment. The only option that actually answers the question, and the subject of the next section.

Advisor takeaway: Decide before the decline what reading triggers a change and what the change will be; anything defined only after the number drops is improvisation under stress, not a course-correction mechanism.

Risk-based guardrails: the course-correction layer

No one has ever walked into an advisor’s office and said, “the main thing I want to know is my probability of success.” “Risk-based” guardrails are the framework built around the real questions clients have: How much can I spend? What would change that? What could a change look like?

Instead of reporting odds, a guardrails plan answers these questions directly, and in dollars: a monthly spending amount, the portfolio balance at which the plan would call for spending more, and the balance at which it would call for slowing down.

The thresholds are set in advance, monitored continuously, and recalculated from the full plan, so the “what do we do?” conversation happens before it is ever needed. (The complete methodology is covered in the complete guide to retirement income guardrails.)

Three things distinguish this from the Monte Carlo status quo.

The analysis stays; the score leaves the client conversation. Sophisticated guardrails still run simulation analysis underneath the surface. And, unlike typical Monte Carlo systems in the most widely-used legacy software systems, modern guardrail software like Income Lab varies not just across the client’s investment returns, but also handles variability of inflation and mortality. Risk is unique to client circumstances: income sources, taxes, longevity, and spending. What changes is the translation layer: the engine’s risk readings become portfolio-balance thresholds and a monthly “retirement paycheck” in dollars, so the client hears “your plan supports $8,600 a month, and here are the two balances that would change it” instead of a percentage. The client’s question was never statistical; guardrails answer the question that was actually asked.

Adjustment is built in, not improvised. A guardrails plan assumes from day one that spending will flex, which is exactly the behavior the standard simulation assumes away. When the lower guardrail is reached, the plan calls for a measured step, typically a 5% to 10% pullback in spending even when the portfolio is down 20% to 30%, rather than a panic cut improvised after bad news. When the upper guardrail is reached, the plan calls for a raise. Underspending is treated as a real risk with a real trigger, not an unaddressed and unrealistic 9-figure estate in the success column.

Risk-based is not Guyton-Klinger “Withdrawal Rate” guardrails. The withdrawal-rate guardrails introduced by Jonathan Guyton and William Klinger deserve credit for popularizing the adjust-as-you-go instinct, but their fixed-percentage trigger rules ignore everything about a plan except the ratio of withdrawals to portfolio balance. They apply the same trigger at 65 and 85. Derek Tharp and Justin Fitzpatrick’s Kitces.com analysis, Why Guyton-Klinger Guardrails Are Too Risky For Most Retirees (And How Risk-Based Guardrails Can Help), shows how withdrawal-rate triggers misfire in practice. Wade Pfau’s 2015 Journal of Financial Planning review of variable spending strategies and Karsten Jeske’s 2017 safe withdrawal rate series reached similar conclusions independently: withdrawal-rate guardrails simply do not perform well in the real world or in randomized simulations. The instinct toward adjustment is the right one, but withdrawal rates cannot deliver an approach that works. Risk-based guardrails are different. They base balance guardrails on the entire plan context: investments, inflation, plan length, mortality risk, etc. So, as the client ages and circumstances move, the guardrails adjust as well. The full comparison is in risk-based guardrails vs. Guyton-Klinger.

Income Lab runs this methodology natively, inside a platform built to be second to none in the full lifecycle of financial planning, and the simulation engine remains a first-class choice within it: the default analysis runs plans through actual historical market sequences, and both traditional and regime-based Monte Carlo are available as selectable analysis methods. This approach insists on elite-level analytics, but also delivers easy to understand answers for clients. The complex math belongs in the kitchen, with results translated into dollars before it reaches the dining room table. For what that dollar-first answer looks like across real households, see how much a retiree can actually spend.

When the market drops and an 85% would have appeared on the screen, Paul and Diane instead hear something like: “Your portfolio is still above your lower guardrail, so nothing changes. If it ever falls to $850,000, the plan calls for a measured spending trim, and we agreed on the shape of that step two years ago.” That $850,000 is not a line drawn once at retirement, either; the plan recalculates it every month as the household ages and circumstances move. The mathematics underneath is the same; the conversation on top is entirely different.

FAQ

What is a Monte Carlo simulation in retirement planning?

A Monte Carlo simulation is a technique that tests a retirement plan against roughly 1,000 randomized market and inflation futures. Each trial draws a different sequence of annual returns, applies the plan’s spending and income year by year, and records whether the portfolio lasted the full horizon. The fraction of trials that end with assets remaining becomes the plan’s reported score.

What does an 85% Monte Carlo score mean?

It means 850 of 1,000 simulated futures ended with portfolio assets remaining, under fixed assumptions about returns, inflation, longevity, and spending that never adjusts. The “probability of success” label makes it sound like a forecast about the client, but it is a conditional frequency inside a model, which is a different thing. It says nothing about how large a surplus or shortfall would be, or what to do if conditions worsen.

Is a 100% Monte Carlo score the goal?

No. The score only measures how careful a plan is being, so it can always be pushed higher by spending less, and a different problem appears at the top: a 100% probability of success is, mathematically, a 100% risk of underspending and regret. A plan that never risks depletion in any simulated future is nearly always a plan that leaves intended living unlived in most of them.

Is a Monte Carlo simulation the same as a retirement stress test?

No. A Monte Carlo simulation generates around 1,000 randomized hypothetical futures. A retirement stress test runs a plan through specific named historical sequences, such as retiring into 2008 or the 1970s inflation era, and shows how the plan and its adjustments would have played out. Testing plans against actual history has a long pedigree; the Journal of Financial Planning study that produced the 4% rule was built on exactly that approach (Bengen 1994). Stress tests trade statistical breadth for a story a client can actually follow, and the two methods are complementary rather than competing.

What should an advisor do when a plan’s score drops?

Recognize what the drop is: a reading on a rigid model that moves with its assumptions, where a point or two (or ten) is often statistical noise. The need to deal with changing probability of success scores is another good reason to abandon them in practice. The durable alternative is a guardrails framework in which spending levels, adjustment triggers, and adjustment sizes are agreed in dollars before any decline, so the plan itself calls for the change. See the complete guide to retirement income guardrails.

Sources

Monte Carlo simulation is sound mathematics. It describes risk; it does not manage risk, and it was never built to. The households sitting across from you need both: an engine that understands uncertainty, and a plan that already knows what to do about it. To watch a real plan answer “what do we do now?” in dollars, with the course corrections agreed before they are needed, Book a Walkthrough.

Justin Fitzpatrick, PhD, CFA, CFP - President and Co-Founder of Income Lab

Justin Fitzpatrick is President and Co-Founder of Income Lab, retirement income planning software used by thousands of financial advisors. He developed the guardrails-based approach to retirement income distribution after a decade in financial services at Jackson and seven years in academia at MIT, Harvard, and UCLA. His research on adjustment-based planning has been published on Kitces.com, ThinkAdvisor, AdvisorPerspectives, and FinancialPlanning Magazine.

Ready to see this in action?

Watch how Income Lab helps advisors answer clients' toughest retirement income questions with guardrails-based planning.

Book a Walkthrough Start Free Trial