-
1
Define the activation and conversion events the model will learn against
Before you fit any model, write down the two outcomes it is trying to predict and the signals it is allowed to use. The conversion event is the clean one: a trial converts to paid when the billing system records a successful first charge or an executed order form against the trial account. Write it down as a boolean on the trial record, not as a derived field in a dashboard, because the model will join against it dozens of times. The activation event is the leading indicator and it is the richest input the model will have. Pick two or three in-product events that pattern-match to what paid customers do in their first session, and instrument them cleanly in the analytics tool. OpenView PLG and ProductLed research both land on the same point: products that explicitly define activation and conversion as named events outperform products that leave them implicit, because the model is only as good as the labels it learns from.
- Write the conversion event as a boolean on the trial record, sourced from the billing system, not from a sales dashboard filter or a CRM stage
- Pick two or three activation events that are observable in-product and that pattern-match to what paid customers did in their first week
- Instrument the events in the analytics or product telemetry tool with the trial account identifier attached, so they join cleanly to the signup record
- Decide the time windows you care about: activation at day one, day three, and day seven are the standard set, and the model will score across all three
Tip: If you cannot write the conversion event as a one-line SQL join from the billing table to the trial table, stop and fix the data model first. A trial-to-paid model that learns against a fuzzy label will produce fuzzy scores, and no amount of fitting will rescue it.
-
2
Pull the ninety-day cohort history and label every trial
The model learns against history, so the next step is to build a clean cohort table that joins every trial from the last ninety days to its final paid outcome. One row per trial, with the signup timestamp, the signup source, firmographic fields from your enrichment provider, every activation event with its timestamp relative to signup, and the paid-conversion boolean. Ninety days is the minimum because shorter windows do not give the slow-converting trials time to resolve, and because the signal you want to find is the difference between a seven-day converter and a thirty-day converter. Appcues trial research and Reforge PLG research both call out the same discipline here: the cohort table is the artifact the whole program depends on, and teams that build it as a view in the warehouse rather than as a one-off query ship a durable model, while teams that pull it ad hoc find themselves rebuilding it every month.
- Build a cohort view in the warehouse with one row per trial and columns for signup, firmographics, every activation event, and the paid-conversion boolean
- Confirm the join from trial to billing is clean by spot-checking twenty converted trials and twenty non-converted trials against the raw billing records
- Add a column for days-to-conversion for the paid cohort and days-to-churn or days-to-expiration for the non-paid cohort, because the model will segment on both
- Version the view with a dated name so monthly reruns do not silently overwrite the training set that produced the deployed model
Tip: Treat the cohort view as the source of truth. If marketing, sales, and product are each pulling their own trial numbers from three different places, the model will not be trusted no matter how well it predicts, because stakeholders will cherry-pick the dashboard that matches their intuition.
-
3
Segment conversion by signup source before you fit anything
Before you fit a model, segment the cohort by signup source and look at the base-rate conversion for each. Organic signups, paid signups, referral signups, and existing-list signups usually convert at wildly different rates, and a single global model will quietly bias toward whichever source dominates volume. The segmentation step is the one most teams skip and the one that pays off the most, because it tells you whether you need one model or several, whether paid acquisition is converting below its channel baseline, and whether a specific referral source is doing the work the sales team thinks the product is doing. OpenView PLG benchmarks are consistent that trial-to-paid rates vary by source by multiples, not by percentage points, and that any model fit without the segmentation is modeling the channel mix as much as the user behavior.
- Compute trial-to-paid conversion for the last ninety days segmented by signup source, with a minimum cohort size of fifty signups per source before you report a rate
- Pair the conversion rate with days-to-conversion and activation rate at day three for each source, so you can see which sources convert fast and which drag
- Flag any source whose conversion is more than fifty percent below the organic baseline, because that source is a candidate for its own model or for a volume cut
- Write the segmentation down and share it with marketing and sales before you fit the model, so the downstream scores are not a surprise when they land in the CRM
Tip: If a signup source has fewer than fifty trials in the ninety-day window, do not report a conversion rate for it. Small-sample rates look precise and are not, and the model will overfit to the noise if you treat them as signal.
-
4
Build the predictive signal from fit, activation, and engagement
The first version of the model should fit on an index card and should combine three families of inputs: fit, activation, and engagement. Fit is the firmographic layer: company size, industry, geography, and ICP match from your enrichment provider. Activation is the event layer: whether each activation event fired, how fast, and in what order. Engagement is the depth layer: sessions per week, seats invited, data imported, and the volume of core-workflow runs. Fit a logistic regression or a gradient-boosted tree against the paid-conversion label, and keep the feature set small enough that every input is explainable in a sentence. Reforge PLG research is clear that the first version of a PQL or trial score should be interpretable, because the sales team has to trust it to act on it, and a model that cannot be explained to a rep will be routed around no matter how accurate it is.
- Define the three input families as a short list: fit features from enrichment, activation features from the event instrumentation, and engagement features from product telemetry
- Fit a logistic regression as the baseline, note its lift over the base rate, and only move to a gradient-boosted tree if the lift justifies the loss of interpretability
- Hold out the most recent two weeks of the cohort as a validation set, and judge the model on its calibration on that holdout, not on in-sample accuracy
- Document every feature in a one-line description, because the sales team, the growth team, and the next person to retrain the model all need the same reference
Tip: Resist stacking twenty features on the first version. A clean five-feature model that a rep can explain in a sentence will outperform a twenty-feature model that nobody trusts, because trust is what gets the score acted on inside the CRM.
-
5
Score every active trial daily and write the score to the CRM
A model that runs once a week is a report. A model that runs every day and writes a score to the CRM is an operating layer. Set up the scoring job to run every twenty-four hours against the full set of active trials, write the score to a stand-alone field on the trial or lead object in the CRM, and attach the top contributing features so the rep can see why. The frequency matters because trials are a time-boxed asset: a trial that becomes hot on day five of a fourteen-day window needs a sales touch within hours, and a weekly batch loses half the window to lag. Appcues trial research and ProductLed benchmarks both call this out: the lift from a daily scoring cadence over a weekly one is larger than the lift from most model-quality improvements, because latency compounds inside a short trial window.
- Schedule the scoring job to run daily against the full set of active trials, with a stable completion time so downstream automations can trigger on fresh scores
- Write the score to a dedicated field on the CRM record, along with the top two or three contributing features, so the rep sees the signal and the reason in one view
- Expose the score as a column and a filter on the trial and lead list views, so sales can sort on it rather than scanning for context in a note field
- Monitor the scoring job itself with a simple health check, because a silent failure that leaves yesterday's score in place is the kind of bug that compounds for weeks before anyone notices
Tip: Never let the model write to a free-text note or a derived dashboard column. The score has to live on a first-class CRM field so it can be filtered, routed, and reported against, and so the next automation can trust it as a stable input.
-
6
Trigger a sales touch on hot trials and route to a named rep
Once the score lives in the CRM, the next job is to wire the triggers. Define the hot threshold against the validation holdout, usually the top ten to twenty percent of active trials by score, and route those trials to a named rep rather than a round-robin queue. Agree on an SLA for first touch, typically under two business hours during business days for a hot trial, and track SLA adherence weekly for the first ninety days. The point of the sales touch is not to sell at the user; it is to show up with context, because the model has already told you which activation events fired, which friction points are showing up in the session data, and which firmographic signals matter. OpenView PLG and Reforge both call out the same discipline: a product qualified lead program dies when sales treats a PQL like a cold inbound, and it compounds when sales treats a PQL as a warm conversation that the product has already opened.
- Define the hot threshold from the validation holdout and agree on the number of hot trials per week sales can realistically touch at the agreed SLA
- Route hot trials to a named rep based on account ownership or a territory map, not to a round-robin queue, so the first touch carries context
- Equip the rep with a short talk track that references the activation events the model flagged, the firmographic match, and the specific friction points in the trial session
- Measure SLA adherence, accepted rate, and converted rate on the hot cohort weekly for the first ninety days, and tune the threshold up or down based on accepted rate, not on volume
Tip: Do not open the hot threshold to raise sales activity. If accepted rate drops below the baseline, the threshold is too loose, and loosening it further trains the sales team to ignore the score. Tighten first, then widen only when accepted rate holds.
-
7
Nurture the low-scoring trials with product-led emails and in-app content
The majority of trials will never be hot, and that is correct. The job for the low-scoring segment is to run a product-led nurture program that keeps the activation events in reach and that surfaces upgrade triggers at the moments the user is near the ceiling. Build the nurture against the same activation events the model learned on, so a user who has not reached activation at day three gets a message that drives at that specific event, and a user who has activated but not expanded gets a message that drives at the next workflow. Keep the content short, action-oriented, and tied to in-product state rather than to calendar days. ProductLed and Appcues both show the same pattern: nurture that is sequenced against user state converts materially better than nurture sequenced against the signup date, because state-based sequencing meets the user where they actually are rather than where the lifecycle tool thinks they should be.
- Write a message map that pairs each activation event with the specific email and in-app message that drives at it, and suppress messages for users who have already completed the event
- Sequence the nurture against in-product state, not against calendar days, so a user who activates on day two does not receive the day-three activation nudge
- Add upgrade triggers at the moments the user hits the free-tier ceiling or the time-boxed expiration, with a clear path to the paid plan and a backup path to extend
- Measure activation lift and upgrade lift from the nurture program weekly, segmented by score band, so you can see which messages are moving which part of the funnel
Tip: Do not route low-scoring trials to the sales team as a safety net. Nurture is a different motion from sales, and loading the sales team with low-score trials to backstop the model is how both motions quietly fail.
-
8
Review model performance monthly and retrain on fresh data
Once the model has thirty days of production data, run a formal review every month for the first two quarters and quarterly thereafter. Lead with four numbers: trial-to-paid conversion on the hot cohort versus the base rate, accepted rate on the hot cohort inside the sales team, activation lift on the nurture program, and model calibration on the last thirty days of trials. The calibration number is the one most teams miss and the one that most tells you when to retrain: if trials scored at a seventy percent probability are converting at fifty, the model is drifting and the features are stale. Retrain on a rolling ninety-day window every month, keep the feature set stable unless a review decision changes it, and version the deployed model with a dated name so you can roll back cleanly. Reforge research is consistent that durable PQL and trial-score programs are retrained on a cadence and tuned one lever at a time, and that teams that leave the model static for a quarter watch its lift decay without seeing it until the sales team stops acting on the score.
- Retrain the model monthly on a rolling ninety-day window, hold out the most recent two weeks for validation, and compare calibration against the previous deployed version
- Report the four headline numbers to the weekly revenue review, so the trial-to-paid program stays funded on performance rather than on vibes
- Pick one layer per month to tune: the feature set, the hot threshold, the nurture message map, or the sales SLA, and change one thing at a time
- Document every retrain with the training window, the feature list, the holdout calibration, and the production decision, so the next person to touch the model has a clean trail
Tip: Days-to-activation is the leading indicator you should obsess over, because it moves weeks before trial-to-paid does. A program where days-to-activation is dropping is a program where the hot cohort is about to convert at a higher rate, and the model should be retrained before that happens, not after.