-
1
Pick exactly one variable and write the hypothesis down
Pricing experiments fail when they try to test strategy. Strategy is the framework: packaging, segmentation, value metric, and positioning. An experiment is a controlled test on one lever inside that framework. Pick exactly one: seat price, tier threshold, or usage-overage rate. Not two. Not seat price with a trial-length change. One. Then write a single-sentence hypothesis in the form 'if we change X from A to B for a randomly selected cohort of new accounts, we expect metric M to move by at least D over 60 days of exposure, without degrading guardrail metric G.' Save the written hypothesis in the shared RevOps folder before randomization starts. It is the only defense you will have against hindsight storytelling in the read-out.
- Choose one of three variables: seat price, tier threshold, or usage-overage rate.
- Write the hypothesis sentence with named primary metric, minimum detectable effect, and guardrail metric.
- Have the CFO and CMO both sign the hypothesis document before cohort assignment begins.
- Pre-register the stop criteria and the read-out date in the same document.
Tip: If the sponsors want to test three variables at once because 'we only get one shot,' walk away and run three sequential 90-day tests instead. A multi-variable test tells you nothing you can act on and burns the same quarter.
-
2
Define cohort randomization and the control arm cleanly
The quality of a pricing experiment is set the moment cohorts are assigned. The default pattern for B2B is a prospect-level random assignment at the point of first qualified touch, so a given account sees either the control price or the variant price for the full 60-day exposure window and never sees both. Randomize on a stable hash of the account identifier so the assignment is reproducible and auditable. Do not randomize on session, because the same buying committee will see both arms. Do not post-hoc segment by industry or size after the fact, because that is how teams invent findings that are not there. Document the control arm with the same precision as the variant: current list price, current discount policy, current packaging. If the control is ambiguous, the read-out will be too.
- Randomize at the account level using a stable hash, not session or user.
- Lock the control arm definition: list price, discount guardrails, and packaging as of day one.
- Exclude existing customers, expansions, and partner-sourced deals unless the test is scoped to them.
- Confirm the Strkr CRM is tagging every new opportunity with its cohort assignment at creation.
Tip: A 50 or 50 split is almost always correct. Teams that split 90 or 10 to 'limit exposure' end up underpowered and reading noise. If the test is too risky at 50 or 50, the real problem is the hypothesis is too aggressive.
-
3
Instrument the funnel for conversion, ARPU, and trial-to-paid
Three metrics carry a pricing experiment: funnel conversion from qualified lead to closed-won, average revenue per account, and trial-to-paid rate for self-serve motions. Instrument each one by cohort before the test starts, not after. In Strkr, that means a cohort field on the opportunity, a billing export that joins to the cohort tag, and a dashboard that updates daily so the operator can watch guardrails. Pull a dry run against the last 60 days of baseline data and confirm the pipeline math reconciles to the finance system within one percent. If the dry run drifts, the live read-out will drift too, and no one will trust the result when it matters.
- Add a cohort field on the opportunity and inherit it onto any resulting subscription or invoice.
- Build a single daily dashboard that shows funnel conversion, ARPU, and trial-to-paid by cohort.
- Reconcile the dashboard numbers against the billing system before day one of the live test.
- Lock the dashboard view so no one can filter by segment mid-flight and reshape the story.
Tip: If finance and RevOps cannot agree on how ARPU is calculated during the dry run, stop the experiment and fix the definition. You cannot resolve a definition dispute after you have a result.
-
4
Set guardrails on margin, win rate, and logo quality
The primary metric tells you whether the variant moved the needle. Guardrails tell you whether the move came at an unacceptable cost. For a B2B pricing experiment, three guardrails are standard: gross margin per deal in the variant arm cannot drop more than a pre-agreed percentage below control, win rate on competitive deals cannot drop more than a pre-agreed percentage below control, and the variant arm cannot overweight small or low-fit logos that will churn in year two. Each guardrail gets a numeric threshold in the hypothesis document, and each one has a named owner on the Finance side who signs off on the threshold before randomization. If any guardrail breaches during the run, the test stops and the sponsors meet before day sixty.
- Set a numeric margin floor for the variant arm, with the CFO as the named owner.
- Set a numeric win-rate floor on competitive deals, with the CMO as the named owner.
- Set a logo-quality guardrail: cap the share of small or low-fit accounts the variant can absorb.
- Agree the stop meeting is automatic the first time any guardrail breaches, no debate on the day.
Tip: Guardrails without named owners are decorations. If a breach happens and no one has signing authority on the stop call, the test keeps running and the quarter is already lost.
-
5
Launch, hold the design still, and run for the full window
The hardest part of a pricing experiment is leaving it alone once it is live. Sales will ask to flex the variant price on a strategic deal. Marketing will want to re-skin the pricing page mid-flight. The CEO will see a bad week and ask to pause. Each of those is a design change that invalidates the read-out. Agree in advance that the only mid-flight actions are the pre-registered guardrail stops and nothing else. Hold the variant exactly as launched for the full 60 days of exposure. Communicate the lockdown to Sales leadership and Marketing in writing before day one so no one is surprised when a request is declined. The point of running an experiment is to earn a defensible answer, and defensibility only survives if the design survives.
- Send a one-page briefing to Sales and Marketing leadership on day zero describing the lockdown rules.
- Route any mid-flight discount, packaging, or page change request through RevOps for impact review.
- Snapshot the pricing page, quote templates, and billing config on day zero for audit.
- Hold a short weekly check-in to review guardrails only, not the primary metric.
Tip: If the sponsors cannot resist peeking at the primary metric weekly, hide it. Give the operator a private dashboard and share only the guardrail view until day sixty. Peeking at primary metrics is how teams stop good experiments early and keep bad ones alive.
-
6
Watch the pre-registered stop criteria, not your gut
A disciplined experiment has two stop paths and only two. First, a guardrail breach: margin, win rate, or logo quality drops below the pre-registered threshold, and the test ends immediately pending a sponsor meeting. Second, a futility stop: by the midpoint, the primary metric is tracking so far below the minimum detectable effect that completing the window will not change the decision. Everything else is noise. A big win in week two is noise. A bad week after a holiday is noise. A loud deal that went sideways is noise. The pre-registration exists so the sponsors do not have to argue in the moment about which noise is real. The operator reports against the stop criteria at the weekly check-in and the sponsors ratify or override in writing.
- At each weekly check-in, the operator reports yes or no against each pre-registered stop condition.
- Any override of a stop condition requires both sponsors to sign the override in the shared document.
- Record every override with a timestamp so the final read-out can account for design drift.
- If a stop fires, pause new cohort assignment immediately and let in-flight deals complete under their arm.
Tip: Teams that override their own stop criteria more than once per experiment do not have an experiment, they have a rolling opinion. Rewrite the pre-registration process before the next test.
-
7
Read out the result and ship a decision in writing
At day ninety, the operator produces a single read-out document with four sections: hypothesis as pre-registered, cohort sizes and exposure dates, primary metric and guardrails with confidence intervals, and the recommended decision. The decision is one of three: adopt the variant, keep the control, or run a follow-up test on a tightened hypothesis. The CFO and CMO sign the decision in the same document that holds the original hypothesis so the paper trail is complete. Then communicate the outcome to Sales, Marketing, and Finance in a short internal memo that names the lever tested, the result, and the effective date of any change. A quiet, documented decision is the point of all this. A loud roll-out without a signed read-out is how pricing changes get rolled back three months later.
- Produce the read-out document with the four fixed sections, no appendices beyond the raw dashboard snapshot.
- Decision is one of adopt, keep, or re-test with a tightened hypothesis. No 'adopt with modifications.'
- Both sponsors sign the decision in the same file as the original hypothesis and the override log.
- Send a one-page internal memo to Sales, Marketing, and Finance with the lever, result, and effective date.
Tip: If the read-out is longer than one page plus a dashboard screenshot, the team is burying the finding. Make it short enough that the next pricing debate can start by quoting it directly.
-
8
File the artifacts and schedule the next test
The last step of a pricing experiment is the one most teams skip, and it is the reason the next experiment starts from scratch. Save every artifact in a durable location: the signed hypothesis, the cohort assignment log, the daily dashboard snapshots, the override log, the read-out document, and the final decision memo. Tag each artifact with the lever tested and the quarter so a future RevOps hire can find it. Then book the next pricing experiment on the finance calendar now, while the lesson is fresh, with a candidate hypothesis in draft form. B2B pricing power compounds across tests. A single experiment answers one question. A rolling cadence of one test per quarter, each one tightening the last, is how a company earns durable pricing leverage over three to five years.
- Save all artifacts in a dated folder under Finance or RevOps shared storage with a consistent naming convention.
- Tag each artifact with the variable tested and the quarter so later teams can search it.
- Draft the next hypothesis within two weeks of the read-out while context is still fresh.
- Put the next experiment start date on the shared finance calendar before the current one is closed out.
Tip: Set a quarterly pricing review on the calendar whose only agenda is the artifact folder. If the folder is empty at the review, no one is running experiments, and pricing is drifting on gut again.