-
1
Pull submitted commit snapshots for the last six to twelve quarters
A back-test starts with raw history. Pull every forecast submission for the window you want to calibrate across. Six quarters is the minimum signal, twelve is better because seasonality and ramp effects get smoothed out. Grab the Commit, Best Case, and plan numbers at the moment of submission for each week, not the final restated version after the quarter closed. Attach the rep, manager, segment, product line, and new-logo-versus-expansion split that was in effect at submission time. Store the pull as an immutable dataset with its own version tag so later cycles can rerun the same analysis against the same inputs.
- Export each weekly forecast snapshot with its lock timestamp from the CRM
- Preserve the organizational hierarchy as of the submission date, not the current one
- Save the pull as a dated artifact so follow-up cycles can diff cleanly against prior runs
Tip: If you only have five quarters of clean data, start the loop anyway. Waiting another two quarters to begin calibration is a worse decision than running an imperfect first pass and tightening it next cycle.
-
2
Join snapshots to actual closed revenue by period
Match every submitted number to what actually closed in the same period. Build the join on three keys: the forecast period, the rep or team the deal was credited to, and the deal identifier when available. Count the actual as realized bookings or recognized revenue, whichever the forecast was meant to predict. Separate deals that closed inside the forecasted period from deals that slipped in or slipped out, because slippage is its own bias signal that gets washed away in a simple total. Keep expansion and new logo on separate rows. The output is a tidy table with submitted Commit, submitted Best Case, actual close, and the delta per rep per period.
- Join on period plus rep plus deal id to catch mid-period rep changes cleanly
- Flag every deal that landed in a different period than it was forecasted into
- Preserve new-logo versus expansion as distinct rows because they bias in opposite directions
-
3
Measure accuracy percent per rep, manager, and segment
Compute accuracy at three altitudes so the story is specific. Per rep: actual close divided by submitted Commit, expressed as a percent. Per manager: the aggregate across all reps in their team, which exposes whether a manager is coaching the number up, down, or holding steady. Per segment: enterprise, mid market, and velocity each roll separately because their variance profiles are different animals. Report four columns: average accuracy, median accuracy, standard deviation, and the trend over the last four cycles. A rep at 102 percent average with a 25 point standard deviation is less reliable than a rep at 97 percent with a 4 point standard deviation, and the model should treat them differently.
- Compute both mean and median because a single blowout quarter will skew the mean
- Add standard deviation as a volatility signal so steady underpredictors are not confused with wild swingers
- Chart the four cycle trend so coaching conversations can anchor on direction, not just level
Tip: Do not grade a rep on fewer than three cycles of history. One quarter of noise is not a pattern, and premature labeling destroys the trust you need to run the loop sustainably.
-
4
Identify systematic bias: sandbagging, inflating, late slips
Accuracy tells you whether the number was right. Bias tells you how it was wrong. Sort each rep and manager into one of four patterns based on the direction and timing of their miss. Sandbaggers submit Commit numbers that consistently close above 105 percent, which hides upside from the board and starves the next period of coverage. Inflaters submit above actuals, usually driven by quota pressure or optimism bias on verbal commitments. Late slippers are accurate on total count but keep moving close dates into later periods, which signals weak qualification. Early callers lock Commit too soon and then churn the number as deals shift. Tag every rep and every manager with their dominant pattern so coaching matches the actual failure mode.
- Classify sandbagging when average accuracy is above 105 percent across three or more cycles
- Classify inflating when average accuracy is below 90 percent and standard deviation is low
- Classify late slippers by tracking the median close-date change per deal from first commit to actual close
-
5
Decompose misses by reason code and stage
A missed forecast has a cause. Pull every deal that was in Commit or Best Case but did not land, and tag it with a loss reason: competitor win, loss to no decision, budget frozen, security review stalled, champion left, deal shrink, or forecasted but never qualified. Cross tabulate the reasons against the stage the deal was in when it was forecasted. If forty percent of missed Commit deals were sitting in a stage without a procurement contact, the exit criteria for that stage are too loose. If most misses tag to no decision, the qualification bar is the problem, not the forecast model. The output is a short list of concrete process defects that caused the variance, not a vague accuracy complaint.
- Require a reason code on every missed deal before the cycle is signed off
- Pivot reason codes against the stage at commit time to expose which exit criteria are weakest
- Separate shrinks from full losses because they signal different sales motions and need different fixes
Tip: A high no decision rate is almost never a forecasting problem. It is a qualification problem that is masquerading as forecast noise. Fix the stage definition before you tune the rollup.
-
6
Adjust rollup multipliers using the calibration data
The point of the back-test is to feed the math. For each rep and segment, compute a calibration multiplier equal to the trailing four cycle average accuracy. If a rep averages 92 percent of their Commit, their submitted Commit gets multiplied by 0.92 in the leader rollup until the trend changes. Apply the same logic at the segment level. Use median rather than mean when standard deviation is high, so one blowout cycle does not warp future cycles. Publish the multipliers so reps and managers understand how the system is reading their history, and refresh them every cycle so improvements show up in the number quickly, not six months late.
- Compute per rep and per segment multipliers from trailing four cycle accuracy
- Use median when standard deviation exceeds 15 points so outliers do not dominate
- Make the multiplier visible to the rep inside the CRM so improvement feels earned, not hidden
-
7
Re-publish categories, coverage targets, and the risk score
Calibration cascades beyond the rollup. Rewrite forecast category exit criteria if the back-test shows deals move into Commit without the required evidence. Reset coverage targets per segment using the newly measured close rate rather than last year's assumption. Retune the Strkr AI risk score weights based on which signals correlated with misses in the last window and which did not. Publish all three changes in one package with a changelog so the whole org reads the same version of the system. The teams that run this step well treat forecast calibration as a configuration release, not an unspoken drift, and they can point at which cycle introduced which rule.
- Rewrite Commit exit criteria where back-test shows deals missed from inside Commit at a high rate
- Reset coverage targets to the inverse of newly measured segment close rate plus a slippage buffer
- Retune Strkr AI risk score weights toward signals that correlated with the last window's misses
Tip: Version every change. If you cannot tell a rep which cycle changed which rule, you cannot defend the rollup to the board when the number moves.
-
8
Communicate the calibration output to the field
A calibration that is not communicated is a secret tax on the field. Send a one page summary at the end of every cycle covering four things: the headline accuracy numbers at company, segment, and manager level, the top three bias patterns identified, the specific changes being published (multipliers, categories, coverage targets, risk score weights), and the two or three coaching focus areas for the next cycle. Call out improved reps by name as well as reps who need coaching, so the loop feels like a fair accounting rather than a one way audit. Give managers forty-eight hours to question the data before the new configuration goes live.
- Write a one page summary with headline accuracy, bias patterns, configuration changes, and coaching focus
- Name improved reps alongside reps who need support so the loop feels balanced
- Open a 48 hour window for managers to challenge the numbers before the new rules lock
-
9
Lock the cadence and automate the pull
A back-test loop that only runs when someone remembers is not a loop. Set a fixed cadence and defend it. Monthly works for high velocity teams where a lot of deals close inside a thirty day window. Quarterly is right for enterprise motions where the signal needs a full cycle to form. Put the pull, the join, the accuracy math, and the bias classification into a saved report inside Strkr so the next cycle starts with the dataset already assembled. Assign a named owner, usually in RevOps, who runs the loop whether the quarter was good or bad. Hitting plan is the most dangerous moment to skip calibration, because success without accounting hides the failures that are quietly building.
- Pick monthly for velocity motions and quarterly for enterprise motions, then do not switch
- Save the pull, join, accuracy math, and bias tags as a repeatable report in Strkr
- Assign a named owner in RevOps who runs the loop every cycle regardless of attainment
Tip: Run the loop in the quarters you hit plan as hard as you run it in the ones you miss. The hidden luck is where next quarter's miss is being incubated.