-
1
Audit historical data quality before you train anything
An AI forecast is only as good as the pipeline history feeding it. Spend the first week auditing data quality across the opportunities, activities, and stage transitions you plan to score. Measure field completeness on every closed deal in the last four to six quarters, not just the open set. Look for stage skips, missing close dates, backdated amount changes, and duplicated opportunities that inflate win rate. Compare rep-reported close dates to actual close timestamps, because systematic optimism bakes into the model as a confident wrong answer. Clean the obvious errors, document the known gaps, and park the fields that are too noisy to use rather than letting them silently distort the signal.
- Run a field completeness report by stage and segment, and set a 90 percent floor on the fields the model will consume
- Pull a sample of 50 closed deals and compare the CRM story to the actual email, call, and meeting record for ground truth
- Flag opportunities with stage skips, reopened cycles, or amount jumps above 25 percent and decide whether to include or exclude them
Tip: Do not try to fix every historical record. The goal of this audit is to pick the subset of history that is clean enough to train on, not to rewrite the past.
-
2
Define the risk signals you want the model to watch
AI value in forecasting comes from signals rep intuition cannot track at scale. Before you pick an approach, write down the six to ten signals that most consistently predict slippage in your motion. Deal age inside the current stage is almost always the strongest single signal. Activity decline on the buyer side, measured by inbound email replies and meeting acceptance rate, is a close second. Stage skips, discount creep above a documented floor, missing decision maker contact roles, and a stale next step beyond a defined threshold all deserve a slot. Agree on definitions with sales leadership before you touch a model, because the signal list is a policy document as much as a feature list.
- Draft a signal catalog with name, definition, source field, and the hypothesis for why it predicts risk
- Review the draft with first line managers and veteran reps so the final list reflects buyer behavior and not just seller activity
- Rank signals by expected weight, and reserve the top three or four for a transparent rules-based layer anyone can audit
Tip: Treat the signal catalog as living. Add new signals when a loss review surfaces one, and retire signals that stop correlating after a quarter of data.
-
3
Choose an AI approach that matches your data and team
There is no single right AI approach for forecasting. Match the method to your data volume, your explainability expectations, and the team that will maintain it. A rules-based scoring layer is the fastest to ship and the easiest to defend, because every number traces to a weighted combination of named signals. A machine learning model trained on your history delivers tighter accuracy once you have thousands of closed deals, at the cost of harder explainability. An LLM-assisted narrative layer is powerful for summarizing why a deal looks risky and surfacing patterns in rep notes, but it is weakest at point predictions. Most mature teams run a hybrid: rules-based score as the backbone, ML overlay once data is sufficient, and Strkr AI narrative on top for coaching conversations.
- Score your motion on deal volume, data quality, and in-house analytics capacity before picking the primary approach
- Start with rules-based scoring if closed deal count is under 2,000 per year or if explainability is a hard requirement
- Reserve the ML overlay for the signals where human-written rules keep missing, and keep the rules layer live as the audit trail
-
4
Run the model in shadow mode next to the human forecast
Do not replace the human forecast on day one. Run the AI call in shadow mode for a full 60 days, which usually spans the back half of one quarter and the opening of the next. During shadow mode, the model scores every deal and produces its own Commit, Best Case, and risk-adjusted total, but the number submitted to the board is still the human roll-up. Store both numbers as immutable snapshots at every weekly lock so you can compare them cleanly after the quarter closes. Resist the urge to tune the model mid-cycle. The point of shadow mode is to measure the gap, not to erase it.
- Snapshot the AI forecast and the human forecast at the same weekly cadence, with the same categories and definitions
- Instrument an explainability panel so every AI call shows the top three signals that drove the score
- Communicate clearly to the sales team that shadow mode is not a performance review of their forecasting
Tip: Run shadow mode for at least one full close cycle before anyone sees the comparison. Partial-cycle numbers mislead on both sides.
-
5
Present the AI call and the human call side-by-side
Once shadow mode has produced two or three weekly snapshots, start publishing the comparison inside the pipeline review. Show the AI Commit next to the rep Commit on the same screen, with the top risk signals for every deal where they disagree. Keep the format consistent week over week so patterns are easy to spot. Do not make the AI number the headline. The headline is the delta and the signals behind it, because that is what drives a coaching conversation. Managers should walk out of the review with a short list of deals to inspect, not a sense that the model is grading them.
- Build a single scorecard that shows rep Commit, AI Commit, delta, and the two or three driving signals per flagged deal
- Use color only to flag disagreement above a defined threshold, so the eye lands on the deals that need a conversation
- Rotate the deal review order so the same reps are not always on the hot seat in week one
-
6
Resolve disagreements with a why-did-AI-disagree post-mortem
Every meaningful disagreement between the rep call and the AI call becomes a short post-mortem after the deal closes. If the rep had the deal in Commit and the AI called it at risk, and the deal slipped, log the signals that fired and the rep context that was missing. If the AI called it at risk and the deal closed anyway, log what the model did not see. The point is not to pick a winner. The point is to build a shared library of patterns, because the model learns from the loss and the rep learns from the pattern. Over three to four quarters, these post-mortems are the single biggest driver of accuracy, far more than any model tuning.
- Create a lightweight post-mortem template with deal id, final outcome, rep call, AI call, driving signals, and lesson learned
- Run post-mortems only on the top 10 to 15 disagreements per quarter so the loop stays focused on signal, not noise
- Feed every post-mortem lesson back into either the signal catalog or the rep enablement library, with a named owner
Tip: Record the post-mortem conclusions in writing. A verbal reconciliation inside a pipeline review disappears by the next quarter and the same pattern repeats.
-
7
Calibrate weights, signals, and thresholds quarterly
An AI forecast that is not calibrated goes stale fast, because your motion shifts under it. Hold a formal calibration review on the first business day after each quarter closes. Measure accuracy of the AI call versus the human call against actuals, broken down by segment, product line, and rep tenure. Tune signal weights based on the variance analysis, retire signals that stopped predicting, and promote a signal from the watch list if it earned its slot. Document the version and date-stamp the change inside Strkr, so a forecast run on March 10 can always be traced to the exact signal set that produced it. Calibration is the discipline that separates an AI program from a one-time pilot.
- Compute AI versus human accuracy against Commit, Best Case, and plan, segmented at least three ways
- Adjust signal weights with the smallest change that explains the error, and never more than two weights at once
- Version the signal catalog with a date and a changelog entry so audits and board questions have a clean trail
-
8
Treat AI as a copilot, not an oracle
The last step is cultural, and it is the one most programs miss. Set an operating principle that the AI forecast is a copilot for the rep and the manager, not the authority that overrides them. The submitted number is still a human decision, informed by the AI call and the signal evidence behind it. When the AI is right and the rep is wrong, coach the rep on the pattern. When the rep is right and the AI is wrong, feed the context back into the signal catalog. This posture protects two things at once: it keeps rep accountability intact, which matters for quota and comp, and it keeps the model honest, because no one is told the output is final. Teams that frame AI this way adopt it faster and keep it longer.
- Write a one page operating principle on AI use and share it in the sales kickoff, not just in a Slack channel
- Keep the human sign-off on the submitted forecast, with AI inputs visible but never binding on the number
- Measure adoption by how often reps read the AI call before their roll-up, not by whether they agree with it