How-to guide

How to run a pricing experiment without blowing up the pipeline

Most B2B teams change pricing by committee: a slide at a quarterly review, a gut call from the CEO, and a Monday morning rollout to every new prospect in the funnel. A real pricing experiment is different. It is a 90-day, cohort-randomized A/B test on one variable at a time, with pre-registered metrics, guardrails on margin and win rate, and a stop rule you agreed to before you looked at the data. This guide walks the CFO or Finance lead, the CMO, and the RevOps operator through the exact plan so the output is a defensible answer, not a story.

Before you start

What you need.

Time: 90 days end to end (two weeks design, 60 days run, two weeks read-out)

  • A joint CFO or Finance lead and CMO sponsor who have agreed in writing on which single variable is being tested
  • A RevOps or analytics operator who can randomize cohorts, instrument the funnel, and pull the read-out cleanly
  • At least 90 days of baseline funnel data: traffic, trial starts, trial-to-paid conversion, ARPU, logo churn, and gross margin
  • A documented list price, discount guardrails, and current contract structure so the control arm is well defined
  • Legal and billing sign-off that the test variant can be quoted, invoiced, and renewed under existing terms
Run a 90-day B2B pricing experiment

Step by step.

  1. 1

    Pick exactly one variable and write the hypothesis down

    Pricing experiments fail when they try to test strategy. Strategy is the framework: packaging, segmentation, value metric, and positioning. An experiment is a controlled test on one lever inside that framework. Pick exactly one: seat price, tier threshold, or usage-overage rate. Not two. Not seat price with a trial-length change. One. Then write a single-sentence hypothesis in the form 'if we change X from A to B for a randomly selected cohort of new accounts, we expect metric M to move by at least D over 60 days of exposure, without degrading guardrail metric G.' Save the written hypothesis in the shared RevOps folder before randomization starts. It is the only defense you will have against hindsight storytelling in the read-out.

    • Choose one of three variables: seat price, tier threshold, or usage-overage rate.
    • Write the hypothesis sentence with named primary metric, minimum detectable effect, and guardrail metric.
    • Have the CFO and CMO both sign the hypothesis document before cohort assignment begins.
    • Pre-register the stop criteria and the read-out date in the same document.
    Tip: If the sponsors want to test three variables at once because 'we only get one shot,' walk away and run three sequential 90-day tests instead. A multi-variable test tells you nothing you can act on and burns the same quarter.
  2. 2

    Define cohort randomization and the control arm cleanly

    The quality of a pricing experiment is set the moment cohorts are assigned. The default pattern for B2B is a prospect-level random assignment at the point of first qualified touch, so a given account sees either the control price or the variant price for the full 60-day exposure window and never sees both. Randomize on a stable hash of the account identifier so the assignment is reproducible and auditable. Do not randomize on session, because the same buying committee will see both arms. Do not post-hoc segment by industry or size after the fact, because that is how teams invent findings that are not there. Document the control arm with the same precision as the variant: current list price, current discount policy, current packaging. If the control is ambiguous, the read-out will be too.

    • Randomize at the account level using a stable hash, not session or user.
    • Lock the control arm definition: list price, discount guardrails, and packaging as of day one.
    • Exclude existing customers, expansions, and partner-sourced deals unless the test is scoped to them.
    • Confirm the Strkr CRM is tagging every new opportunity with its cohort assignment at creation.
    Tip: A 50 or 50 split is almost always correct. Teams that split 90 or 10 to 'limit exposure' end up underpowered and reading noise. If the test is too risky at 50 or 50, the real problem is the hypothesis is too aggressive.
  3. 3

    Instrument the funnel for conversion, ARPU, and trial-to-paid

    Three metrics carry a pricing experiment: funnel conversion from qualified lead to closed-won, average revenue per account, and trial-to-paid rate for self-serve motions. Instrument each one by cohort before the test starts, not after. In Strkr, that means a cohort field on the opportunity, a billing export that joins to the cohort tag, and a dashboard that updates daily so the operator can watch guardrails. Pull a dry run against the last 60 days of baseline data and confirm the pipeline math reconciles to the finance system within one percent. If the dry run drifts, the live read-out will drift too, and no one will trust the result when it matters.

    • Add a cohort field on the opportunity and inherit it onto any resulting subscription or invoice.
    • Build a single daily dashboard that shows funnel conversion, ARPU, and trial-to-paid by cohort.
    • Reconcile the dashboard numbers against the billing system before day one of the live test.
    • Lock the dashboard view so no one can filter by segment mid-flight and reshape the story.
    Tip: If finance and RevOps cannot agree on how ARPU is calculated during the dry run, stop the experiment and fix the definition. You cannot resolve a definition dispute after you have a result.
  4. 4

    Set guardrails on margin, win rate, and logo quality

    The primary metric tells you whether the variant moved the needle. Guardrails tell you whether the move came at an unacceptable cost. For a B2B pricing experiment, three guardrails are standard: gross margin per deal in the variant arm cannot drop more than a pre-agreed percentage below control, win rate on competitive deals cannot drop more than a pre-agreed percentage below control, and the variant arm cannot overweight small or low-fit logos that will churn in year two. Each guardrail gets a numeric threshold in the hypothesis document, and each one has a named owner on the Finance side who signs off on the threshold before randomization. If any guardrail breaches during the run, the test stops and the sponsors meet before day sixty.

    • Set a numeric margin floor for the variant arm, with the CFO as the named owner.
    • Set a numeric win-rate floor on competitive deals, with the CMO as the named owner.
    • Set a logo-quality guardrail: cap the share of small or low-fit accounts the variant can absorb.
    • Agree the stop meeting is automatic the first time any guardrail breaches, no debate on the day.
    Tip: Guardrails without named owners are decorations. If a breach happens and no one has signing authority on the stop call, the test keeps running and the quarter is already lost.
  5. 5

    Launch, hold the design still, and run for the full window

    The hardest part of a pricing experiment is leaving it alone once it is live. Sales will ask to flex the variant price on a strategic deal. Marketing will want to re-skin the pricing page mid-flight. The CEO will see a bad week and ask to pause. Each of those is a design change that invalidates the read-out. Agree in advance that the only mid-flight actions are the pre-registered guardrail stops and nothing else. Hold the variant exactly as launched for the full 60 days of exposure. Communicate the lockdown to Sales leadership and Marketing in writing before day one so no one is surprised when a request is declined. The point of running an experiment is to earn a defensible answer, and defensibility only survives if the design survives.

    • Send a one-page briefing to Sales and Marketing leadership on day zero describing the lockdown rules.
    • Route any mid-flight discount, packaging, or page change request through RevOps for impact review.
    • Snapshot the pricing page, quote templates, and billing config on day zero for audit.
    • Hold a short weekly check-in to review guardrails only, not the primary metric.
    Tip: If the sponsors cannot resist peeking at the primary metric weekly, hide it. Give the operator a private dashboard and share only the guardrail view until day sixty. Peeking at primary metrics is how teams stop good experiments early and keep bad ones alive.
  6. 6

    Watch the pre-registered stop criteria, not your gut

    A disciplined experiment has two stop paths and only two. First, a guardrail breach: margin, win rate, or logo quality drops below the pre-registered threshold, and the test ends immediately pending a sponsor meeting. Second, a futility stop: by the midpoint, the primary metric is tracking so far below the minimum detectable effect that completing the window will not change the decision. Everything else is noise. A big win in week two is noise. A bad week after a holiday is noise. A loud deal that went sideways is noise. The pre-registration exists so the sponsors do not have to argue in the moment about which noise is real. The operator reports against the stop criteria at the weekly check-in and the sponsors ratify or override in writing.

    • At each weekly check-in, the operator reports yes or no against each pre-registered stop condition.
    • Any override of a stop condition requires both sponsors to sign the override in the shared document.
    • Record every override with a timestamp so the final read-out can account for design drift.
    • If a stop fires, pause new cohort assignment immediately and let in-flight deals complete under their arm.
    Tip: Teams that override their own stop criteria more than once per experiment do not have an experiment, they have a rolling opinion. Rewrite the pre-registration process before the next test.
  7. 7

    Read out the result and ship a decision in writing

    At day ninety, the operator produces a single read-out document with four sections: hypothesis as pre-registered, cohort sizes and exposure dates, primary metric and guardrails with confidence intervals, and the recommended decision. The decision is one of three: adopt the variant, keep the control, or run a follow-up test on a tightened hypothesis. The CFO and CMO sign the decision in the same document that holds the original hypothesis so the paper trail is complete. Then communicate the outcome to Sales, Marketing, and Finance in a short internal memo that names the lever tested, the result, and the effective date of any change. A quiet, documented decision is the point of all this. A loud roll-out without a signed read-out is how pricing changes get rolled back three months later.

    • Produce the read-out document with the four fixed sections, no appendices beyond the raw dashboard snapshot.
    • Decision is one of adopt, keep, or re-test with a tightened hypothesis. No 'adopt with modifications.'
    • Both sponsors sign the decision in the same file as the original hypothesis and the override log.
    • Send a one-page internal memo to Sales, Marketing, and Finance with the lever, result, and effective date.
    Tip: If the read-out is longer than one page plus a dashboard screenshot, the team is burying the finding. Make it short enough that the next pricing debate can start by quoting it directly.
  8. 8

    File the artifacts and schedule the next test

    The last step of a pricing experiment is the one most teams skip, and it is the reason the next experiment starts from scratch. Save every artifact in a durable location: the signed hypothesis, the cohort assignment log, the daily dashboard snapshots, the override log, the read-out document, and the final decision memo. Tag each artifact with the lever tested and the quarter so a future RevOps hire can find it. Then book the next pricing experiment on the finance calendar now, while the lesson is fresh, with a candidate hypothesis in draft form. B2B pricing power compounds across tests. A single experiment answers one question. A rolling cadence of one test per quarter, each one tightening the last, is how a company earns durable pricing leverage over three to five years.

    • Save all artifacts in a dated folder under Finance or RevOps shared storage with a consistent naming convention.
    • Tag each artifact with the variable tested and the quarter so later teams can search it.
    • Draft the next hypothesis within two weeks of the read-out while context is still fresh.
    • Put the next experiment start date on the shared finance calendar before the current one is closed out.
    Tip: Set a quarterly pricing review on the calendar whose only agenda is the artifact folder. If the folder is empty at the review, no one is running experiments, and pricing is drifting on gut again.
Avoid

Common mistakes.

  • Testing more than one variable at once because 'the quarter is tight,' which produces a result nobody can attribute to a specific lever and forces a re-test anyway.
  • Randomizing at the session or user level instead of the account level, so the same buying committee sees both arms and the result is contaminated before day one.
  • Skipping the written hypothesis and stop criteria, then arguing after the fact about whether the variant 'really worked,' which always ends in a political decision rather than an evidence-based one.
  • Letting Sales flex the variant price on strategic deals during the run, which collapses the design and leaves the read-out indefensible.
  • Peeking at the primary metric weekly and stopping the test early on a lucky two weeks or killing it on an unlucky two weeks, which guarantees the team learns the wrong lesson.
  • Shipping no written decision memo at the end, so the finding evaporates and the same hypothesis gets re-litigated at the next pricing offsite.
FAQ

Frequently asked questions.

How long should a B2B pricing experiment run?

Plan on 90 days end to end: two weeks to design and instrument, 60 days of live cohort exposure, and two weeks to read out and ship a decision. B2B sales cycles are long enough that shorter windows read as noise, and longer windows let the market drift under the test. If the cycle is longer than 60 days, extend the exposure window to cover at least one full median cycle, not more.

What is the right thing to test first: seat price, tier threshold, or usage overage?

Test the lever where the current price was set by gut and where a modest move would most change ARPU. For most early-stage B2B companies that is seat price. For companies with wide usage dispersion across customers it is usage-overage rate. For companies where the mid-tier is converting far below the top and bottom it is the tier threshold. Pick one, write the hypothesis, and queue the other two for the next two quarters.

How do we avoid tanking in-flight pipeline during the test?

Guardrails on win rate and margin, named owners on each guardrail, and an automatic stop meeting the first time any guardrail breaches. Randomize at the account level so a committee only ever sees one price. Exclude strategic or named-account deals from the test population if the risk on any single logo is too high. The experiment is a tool, not a dare, and the guardrails are what keep it that way.

Who owns a pricing experiment inside a B2B company?

Joint CFO or Finance lead and CMO as sponsors, with a RevOps or analytics operator doing the execution. The CFO owns margin and the structural pricing math, the CMO owns positioning and conversion, and the operator owns cohort assignment, instrumentation, and the read-out. Any experiment with only one of the three roles signing off will drift on the dimension the missing role would have defended.

How is a pricing experiment different from a pricing strategy review?

A pricing strategy review is a framework exercise: packaging, segmentation, value metric, positioning, discount policy. It happens annually and sets the lanes. A pricing experiment is a 90-day controlled test on one variable inside those lanes, with a pre-registered hypothesis and a signed decision at the end. Strategy answers 'what should we charge for and how.' An experiment answers 'is this specific number right.' Confusing the two is how teams change everything at once and learn nothing.

What does Strkr do during a pricing experiment?

Strkr is where the cohort assignment, opportunity tagging, pipeline math, and read-out dashboard live. RevOps tags every new opportunity with its cohort at creation, Strkr inherits the tag onto the resulting subscription or invoice line, and the daily guardrail dashboard reads directly from pipeline and billing joined through the cohort field. The point is that the operator never has to reconcile two spreadsheets by hand on the day of the read-out.

See it in Strkr

Related product surfaces.

Strkr platform features Strkr CRM

Run pricing experiments on evidence, not opinion

Strkr gives RevOps a single place to tag cohorts on every opportunity, join pipeline to billing, and watch guardrails daily so your next pricing test earns a defensible answer instead of another slide at the quarterly review.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.