-
1
Inventory the in-product signals that pattern-match to buyer intent
Start by writing down every in-product event that paid customers tend to do and that non-converting workspaces tend to skip. The usual list for a B2B SaaS product is short and durable: feature adoption on the paid-tier features, workspace creation, seat invites sent and accepted, data imported from a source system, integration connected, core-workflow runs, and any admin-settings action that only a buyer would perform. Write each one as a named event in the telemetry tool, with the workspace identifier and the user identifier both attached, so the signal can be rolled up two ways. OpenView PLG and Reforge research both land on the same discipline here: the signal inventory is the artifact the model depends on, and teams that write it down once and reuse it ship a durable program, while teams that improvise the list every quarter watch the model drift because the input layer keeps changing.
- List every in-product event that paid customers reliably do in their first two weeks, and every event that non-converting workspaces reliably skip
- Attach both the workspace identifier and the user identifier to every event, so signals can be rolled up at the account level and at the user level
- Separate the signals into four families: feature adoption, workspace growth, data depth, and buyer behavior, because the model will weight the families differently
- Write a one-line description for each signal and keep the list under twenty events for the first version, because a short list is one the whole team can trust
Tip: If a signal cannot be written as a one-line SQL filter against the event table, drop it from the first version. A fuzzy signal produces a fuzzy score, and the model will learn against the noise rather than against the intent.
-
2
Roll signals up to the account, not just the user
The unit of a PLG signal model is the account, not the user. A single power user who invites nobody is a different buying signal from a workspace with four invited seats and two of them active daily, and a model that scores at the user layer will miss the second case entirely. Build an account-rollup view in the warehouse that joins every event to its workspace, counts active users per week, sums seats invited and seats accepted, and flags whether an admin-settings action has fired in the last fourteen days. The rollup is where PLG scoring stops being a usage report and starts being a sales signal, because buyers are companies and the signal has to live where companies live. Appcues and ProductLed research are both consistent on this point: account-level rollups outperform user-level scores on sales-assisted conversion by a wide margin, because the sales motion operates on accounts and the score has to speak the same language.
- Build an account-rollup view in the warehouse with one row per workspace and columns for every signal, summed over the last seven, fourteen, and thirty days
- Add a weekly active users count and a weekly active seats count, because the ratio of active to invited seats is one of the cleanest buyer-intent indicators
- Flag any workspace with an admin-settings action in the last fourteen days, because that behavior usually means a buyer is configuring the product for a team
- Version the rollup view with a dated name so monthly reruns do not silently overwrite the training set that produced the deployed model
Tip: Treat the account rollup as the source of truth for the sales team. If product, marketing, and sales each pull their own workspace activity numbers from three different places, the score will not be trusted no matter how predictive it is.
-
3
Weight each signal against paid outcomes using the ninety-day cohort
Before you combine signals into a composite, measure how predictive each one is on its own against the paid outcome. Pull the ninety-day cohort of workspaces, join each to its final paid outcome, and compute the conversion rate for workspaces that fired each signal versus workspaces that did not. The signals that lift conversion by a factor of three or more are the ones that deserve the most weight, the signals that lift by less than fifty percent are candidates to drop, and the signals that correlate with conversion only because they correlate with time in the product get a smaller weight so they do not dominate. Reforge research is clear that the first version of a PLG signal weighting should be interpretable, because the sales team has to trust it to act on it, and a weighting that cannot be explained to an AE will be routed around no matter how precise it looks on paper.
- Pull the ninety-day cohort of workspaces with one row per account, every signal as a boolean or a count, and the paid-conversion outcome as a boolean
- Compute the lift of each signal: conversion rate for workspaces that fired the signal divided by conversion rate for workspaces that did not
- Rank the signals by lift, keep the top ten to fifteen, and drop any signal with a lift below one-point-five because it is noise more than signal
- Share the ranking with marketing and sales before you set the weights, so the composite score does not land in the CRM as a black box
Tip: Lift is more trustworthy than raw correlation. A signal that fires on every workspace will correlate with everything, including conversion, but it will not lift the conversion rate on the workspaces that fire it, and that is the number that tells you whether the signal predicts intent.
-
4
Build the composite score from the four signal families
Combine the weighted signals into a single composite score, scaled zero to one hundred, with the four signal families contributing explicit shares. A durable split is roughly feature adoption at thirty, workspace growth at thirty, data depth at twenty, and buyer behavior at twenty, tuned against the validation set rather than against intuition. Keep the composite interpretable by exposing the four family subscores alongside the headline, so an AE opening the CRM record sees the composite, the four family scores, and the top three contributing signals in one view. ProductLed and OpenView benchmarks both show that interpretable composite scores outperform black-box scores on sales adoption, because the sales team actually looks at them, and a score that an AE will not open is a score the model never gets to deploy. The point is not model accuracy in a notebook, the point is operational trust in the CRM.
- Scale the composite to a zero-to-one-hundred range so it reads cleanly alongside other CRM scores and so thresholds are easy to set and communicate
- Expose the four family subscores on the account record, not just the headline, so the AE can see which family is driving the composite
- Attach the top three contributing signals to every scored account as a short list, so the first touch carries context rather than a bare number
- Hold out the most recent two weeks of the cohort as a validation set, and tune the family weights on holdout calibration rather than on in-sample fit
Tip: Resist adding a fifth family on version one. Four families with ten clean signals will outperform six families with twenty-five noisy signals, because trust compounds and complexity does not.
-
5
Score every account daily and write the composite to the CRM
A model that scores once a week is a report. A model that scores every day and writes to the CRM is an operating layer. Run the scoring job every twenty-four hours against every active account, write the composite to a stand-alone numeric field on the account object, write the four family scores to their own fields, and attach the top signals as a short readable string. The frequency matters because PLG buying windows are short: a workspace that crosses the hot threshold on a Tuesday and does not get a touch until the following Monday has usually either already decided or already moved on. Appcues and Reforge research both call out the same discipline here: the lift from a daily scoring cadence over a weekly one is larger than the lift from most model-quality improvements, because latency compounds inside a short buying window and the sales touch has to arrive while the signal is still fresh.
- Schedule the scoring job to run daily against every active account, with a stable completion time so downstream automations can trigger on fresh scores
- Write the composite, the four family scores, and the top signals to dedicated fields on the CRM account object, not to a free-text note or a dashboard column
- Expose the composite as a column and a filter on the account list view, so AEs can sort on it rather than scanning for context in a note field
- Monitor the scoring job with a simple health check, because a silent failure that leaves yesterday's score in place is the kind of bug that compounds for weeks before anyone notices
Tip: Never let the composite land in a derived dashboard column. The score has to live on a first-class CRM field so it can be filtered, routed, and reported against, and so the next automation can trust it as a stable input.
-
6
Route hot accounts to a named AE with a short SLA
Once the composite lives in the CRM, wire the routing. Define the hot threshold against the validation holdout, usually the top ten to fifteen percent of active accounts by composite, and route those accounts to a named AE based on account ownership or territory rather than to a round-robin queue. Agree on an SLA for first touch, typically under four business hours during business days for a hot account, and track SLA adherence weekly for the first ninety days. The point of the AE touch is not to pitch at the workspace, it is to show up with context, because the model has already told you which signals fired, which family is driving the composite, and which seats are active. OpenView and ProductLed are both clear that a product qualified account program dies when sales treats a PQA like a cold inbound, and it compounds when sales treats a PQA as a warm conversation the product has already opened.
- Define the hot threshold from the validation holdout and agree on the number of hot accounts per week sales can realistically touch at the agreed SLA
- Route hot accounts to a named AE based on account ownership or territory, not to a round-robin queue, so the first touch carries context and continuity
- Equip the AE with a short talk track that references the top signals, the family driving the composite, and the active seats in the workspace
- Measure SLA adherence, accepted rate, and converted rate on the hot cohort weekly for the first ninety days, and tune the threshold up or down on accepted rate, not on volume
Tip: Do not loosen the hot threshold to raise sales activity. A loose threshold drops accepted rate, trains the AE team to ignore the score, and quietly destroys the trust that makes the program compound in the first place.
-
7
Nurture the warm accounts with product-led content tied to signals
The majority of active accounts will sit below the hot threshold, and that is correct. Build a product-led nurture track for the warm band that targets the specific signals the model is watching, so an account that is adopting features but has not invited seats gets a seat-invite nudge, and an account that has invited seats but not imported data gets a data-import nudge. Sequence the nurture against in-product state rather than against calendar days, so a workspace that imports data on day two does not receive the day-three import nudge. Keep the content short, action-oriented, and tied to the exact signal the account is missing. ProductLed and Appcues both show the same pattern here: nurture sequenced against user and workspace state converts materially better than nurture sequenced against the signup date, because state-based sequencing meets the account where it actually is rather than where the lifecycle tool thinks it should be.
- Write a message map that pairs each hot-signal with the specific email and in-app message that drives at it, and suppress messages for workspaces that have already fired the signal
- Sequence the nurture against in-product state, not against calendar days, so a workspace that invites seats on day two does not receive the day-three invite nudge
- Add upgrade triggers at the moments the workspace hits a free-tier ceiling, with a clear path to the paid plan and a backup path to request a demo
- Measure signal lift and composite lift from the nurture program weekly, segmented by score band, so you can see which messages are moving which part of the funnel
Tip: Do not route warm accounts to the sales team as a safety net. Nurture is a different motion from sales, and loading the AE team with warm accounts to backstop the model is how both motions quietly fail.
-
8
Review model performance monthly and tune one lever at a time
Once the model has thirty days of production data, run a formal review every month for the first two quarters and quarterly thereafter. Lead with four numbers: conversion on the hot cohort versus the base rate, accepted rate on the hot cohort inside the AE team, signal lift from the nurture program, and composite calibration on the last thirty days of accounts. The calibration number is the one most teams miss and the one that most tells you when to retune: if accounts scored at seventy are converting at forty, the model is drifting and the weights are stale. Retune on a rolling ninety-day window every month, keep the signal set stable unless a review decision changes it, and version the deployed weights with a dated name so you can roll back cleanly. Reforge research is consistent that durable PLG signal programs are retuned on a cadence and that teams that leave the model static for a quarter watch its lift decay without seeing it until the sales team quietly stops acting on the score.
- Retune the signal weights monthly on a rolling ninety-day window, hold out the most recent two weeks for validation, and compare calibration against the previous deployed weights
- Report the four headline numbers to the weekly revenue review, so the PLG signal program stays funded on performance rather than on vibes
- Pick one layer per month to tune: the signal inventory, the family weights, the hot threshold, or the AE SLA, and change one thing at a time so the delta is readable
- Document every retune with the training window, the weight list, the holdout calibration, and the production decision, so the next person to touch the model has a clean trail
Tip: Seats-invited-per-active-workspace is the leading indicator most worth obsessing over, because it moves weeks before paid conversion does. A program where that number is climbing is a program where the hot cohort is about to convert at a higher rate, and the model should be retuned before that happens, not after.