# Pricing Experiments: How to Run One Without a Backlash

> Learn to design pricing experiments that isolate price, avoid the risk that shut down Instacart's tests, and turn results into a real pricing decision.

- URL: https://missiongrowth.io/blog/pricing-experiments
- Published: 2026-08-21 · Updated: 2026-09-24
- Author: Ömer Furkan Aktaş, Founder, Mission Growth
- Publisher: Mission Growth. Company facts: https://missiongrowth.io/llms.txt

A pricing experiment is a controlled test: a hypothesis, an isolated variable, a control group. Picking two prices and splitting your customers is only the start.

Two decisions actually decide whether pricing experiments are trustworthy: how you isolate price so nothing else contaminates the read, and how far you can personalize price before you cross the fairness and legal line that just shut down Instacart's own price-testing program.

Amazon's own design mechanic isolates price correctly where a simple customer split doesn't. The regulatory boundary below is dated and specific: a named case and an exact vote count. And one real example proves a price before a product ever reaches full production.

## What is a pricing experiment?

A pricing experiment is a controlled test that changes price and only price for one group of customers, then measures the difference against a group that saw the current price.

That's the whole method: a hypothesis, one isolated variable, a control group. Charm pricing, bundling and tiered plans are strategies to test, using this method. Run any of them through it and you find out whether it actually moves revenue, signups or retention for your customers, instead of assuming it does.

Ten distinct formats fit this method, from a simple A/B price test to an algorithm that reallocates traffic to the winning price on its own. The next section sorts all ten by what each one isolates.

## Why run a pricing experiment?

Pricing experiments answer four questions a spreadsheet can't:

- Is this price point too high or too low for the segment you're testing?
- Which customer segments will pay more without churning?
- Does a pricing model change protect retention, or quietly erode it?
- Is a new pricing model worth the engineering cost of building it?

Most teams stop at the first question. Just one of the eight pages currently ranking for this topic ties a pricing-experiment goal to retention or a full pricing-model change, and it's Stripe's own research.

Stripe's September 2024 survey of 2,000+ subscription business leaders found that 69% planned to launch new pricing models within the year. Testing whether a new model earns its build cost is already on most roadmaps, whether or not the roadmap calls it a pricing experiment.

## The 10 types of pricing experiments (and which ones fit your situation)

Pricing experiments fall into ten distinct formats, from a simple A/B price test to a multi-armed bandit that reallocates traffic to the winning price as data comes in, and no single guide in the current top results names more than seven of them (seven from one guide, plus three more found elsewhere: multi-armed bandit testing, multivariate testing and conjoint analysis).

| Type | What it isolates | Fairness exposure |
|---|---|---|
| A/B price test | A single price point | Lower |
| Discounts / promos | A temporary price cut | Lower |
| Psychological pricing (charm pricing) | How the price is presented, not the number itself | Lower |
| Dynamic / usage-based pricing trials | Price tied to usage or context | Higher |
| Segment-specific tests | Price by customer segment | Higher |
| Tier / multi-choice tests | Which plan configuration customers pick | Lower |
| Bundles vs. à la carte | Package composition | Lower |
| Multivariate testing | Several price elements tested at once | Moderate |
| Multi-armed bandit testing | Traffic allocation across price variants | Moderate |
| Conjoint analysis | Willingness-to-pay tradeoffs, no live price shown | Lower |

::figure{src="/blog/figures/pricing-experiments-2.svg" alt="Matrix of ten pricing experiment types against what each one isolates and its customer-fairness exposure, from A/B price tests to conjoint analysis" caption="The ten types of pricing experiments split by what they isolate, not by how popular they are." width="720" height="513"}

These types of pricing experiments split into three families by what they isolate: a single price point, a package, or an allocation method. Pick the family first, then the type inside it.

Start with an A/B price test when you're testing a single number. Test a bundle or a tier when you're testing what's included. Reach for a multi-armed bandit once you already know the winning price varies by traffic pattern: it reallocates traffic between price variants as data comes in, and the algorithm itself is a separate topic, covered in our [multi-armed bandit testing](https://missiongrowth.io/blog/multi-armed-bandit) guide.

The two types that personalize price by who's looking, segment-specific tests and dynamic or usage-based pricing trials, sit at the top of the fairness-exposure column above. The same product can show two different prices to two different people based on who they are, which is the exact pattern regulators are starting to restrict. More on that in the launch section below.

## How to design a pricing experiment that produces a trustworthy answer

A trustworthy pricing experiment starts with a falsifiable hypothesis, isolates price as the only variable, and assigns treatment at the level, customer or product, that won't contaminate its own result. Knowing how to design a pricing experiment starts with that order: hypothesis first, isolation second, treatment level third.

::figure{src="/blog/figures/pricing-experiments-3.svg" alt="A pricing experiment follows a fixed design order: write a falsifiable hypothesis, isolate price as the only variable, then choose the treatment level." caption="Design order, not the price picked, decides whether a pricing experiment's result can be trusted." width="720" height="233"}

Write the hypothesis before you touch a price. "If we raise the price from $X to $Y, sign-ups drop no more than Z% and revenue per user rises" is falsifiable. "Let's see what happens if we raise the price" is not; it can't be wrong, so it can't teach you anything.

Isolating "one variable" becomes a design problem once products interact with each other. The default assumption, splitting customers into two groups, isn't automatically safe.

Amazon's own Pricing Labs team randomizes its price tests at the product level instead of the customer level. Joe Cooprider and Shima Nassiri, in a 2023 Business Economics paper, explain why: related products interact, so a customer who never sees a discount can still be affected by it through substitution, and showing two customers two different prices on the identical listing at the same time would break its no-price-discrimination policy.

Randomizing by product sidesteps both problems. The price itself becomes the unit under test, not the person looking at it, and Amazon layers crossover experimental designs and demand-trend controls on top of that randomization to sharpen the estimate further.

If you're testing a SaaS tier change, the same logic applies. Split by account or plan cohort, not by individual seat inside one account, and check whether adjacent plans, a solo tier next to a team tier, can leak into each other's results before you trust the read.

For how many customers you need and how long to run it, use the sample-size math from your [growth experiment cadence](https://missiongrowth.io/blog/growth-experiment-cadence) or work the [minimum detectable effect](https://missiongrowth.io/blog/minimum-detectable-effect) directly.

## How to launch a pricing test without breaking the business (or the law)

Launching a pricing test safely means rolling it out gradually, keeping control and variant clean, and staying inside a fairness line that just cost Instacart its own price testing program. Five moves keep it safe:

- **Roll out gradually.** Begin with a limited slice of traffic or one market, then widen it once you trust the split.
- **Keep the groups clean.** Don't let a customer bounce between control and variant mid-test, and don't let support quietly override the price for someone who complains.
- **Protect the buying experience.** Too many price or plan options in front of one customer at once undermines the read and confuses the person you're trying to convert.
- **Run more than one test.** Don't converge on the first result. Queue the next price or format test before you lock in a decision, especially early on.
- **Know the fairness line before you cross it.**

That line got drawn in public over sixteen months. Consumer Reports opened a data-practices investigation into Kroger in May 2025; one shopper who requested their own data under a state privacy law got back a 62-page profile.

That investigation became the template for a second one, aimed at Instacart. Consumer Reports and the Groundwork Collaborative recruited 437 shoppers for live virtual shopping sessions across four cities, then used 193 of those submissions in the final analysis after data cleaning.

Instacart had been running its own price experiments on Eversight, an AI pricing company it acquired in 2022. The investigation found an average gap of 13% between the lowest and highest price for identical items, and differences as high as 23% for some products, a spread Consumer Reports estimated could add up to $1,200 a year at checkout.

::figure{src="/blog/figures/pricing-experiments-4.svg" alt="Instacart's price gap on identical items averaged 13% and reached as high as 23%, the gap behind its pricing-test shutdown." caption="By Consumer Reports and Groundwork's count, Instacart's price gap on identical items ran from a 13% average to a 23% high." width="720" height="182"}

Instacart ended the program on December 22, 2025. Its own statement said the price differences weren't based on individual shopper data, and that retail partners could still run their own promotions and discounts.

Keep those two claims separate: Instacart's defense is about what data caused the gap, and Consumer Reports' finding is about the gap itself.

Nine months later, Seattle's city council voted 7-2 on September 22, 2026 to adopt the Fair Pricing and Transparency Act, council bill CB 121267, the first U.S. city council to approve a ban on algorithmic grocery pricing. It still needs the mayor's signature before it becomes law.

The act bans using personal data, such as browsing history, location or income inferences, to set an individual grocery price, and requires uniform posted pricing instead. Maryland, Connecticut and New Jersey have already passed state laws restricting this kind of surveillance pricing, so the direction doesn't depend on one city's ordinance holding.

::figure{src="/blog/figures/pricing-experiments-1.svg" alt="Consumer Reports' Kroger investigation, Instacart's pricing-test shutdown and Seattle's algorithmic-grocery-pricing ban, timed across sixteen months." caption="A data investigation into one company's pricing became a citywide ban in sixteen months." width="720" height="184"}

None of this means don't test price. It means treating customer-level personalization as the highest-risk design choice on the menu, one to justify case by case.

Fold pricing tests into your regular [experiment cadence](https://missiongrowth.io/blog/growth-experiment-cadence) and check your state and city rules before you ship a segment-specific test.

## How to analyze pricing experiment results without fooling yourself

A pricing experiment's result sorts into one of four outcomes, and each one calls for a different next move rather than a blanket "ship the winner." That's the core of pricing experiment analysis: sort the result into an outcome, then decide, instead of stopping at a p-value.

| Result | What it means | Next move |
|---|---|---|
| Significant win, no downside | The change improved your primary metric without hurting anything else | Ship it |
| Significant win, retention cost | The change won on the primary metric but hurt retention in some segment | Roll out only to the segment where the win holds |
| Not significant | Price wasn't the binding constraint on the decision you're testing | Look elsewhere: packaging, positioning or the product itself |
| Significant loss | The change measurably hurt the metric you set out to test | Document why, and close the test |

Check significance once, against the analysis plan you set before the test started. If you need to check results mid-test without inflating your false-positive rate, that's a separate, peeking-safe method (see [sequential testing](https://missiongrowth.io/blog/sequential-testing)); if you want a faster, more precise read from the same sample size, [cuped](https://missiongrowth.io/blog/cuped) is the variance-reduction technique for that.

Whatever the result, translate it into a number finance will actually use. Run the resulting change in ARPU or retention through an [LTV:CAC ratio calculator](https://missiongrowth.io/tools/ltv-cac-calculator) to see whether the new price improves unit economics rather than top-line revenue alone.

It's the same finance-facing framing we use to calculate [SEO ROI](https://missiongrowth.io/blog/seo-roi): a test result only earns trust once it's stated in the room's own currency.

If the result is a loss, document it with the same trust-preserving discipline behind [SEO reporting on a bad month](https://missiongrowth.io/blog/seo-reporting): state what you expected, what happened, and why, instead of quietly dropping the test from the record.

## A pricing experiment that proved a price before the product shipped

First Build, GE Appliances' innovation studio, proved a price for the Opal Nugget Icemaker before building it at full production volume. It priced an Indiegogo campaign at $399 against an anticipated $499 retail price and let real money answer whether the market would pay it, in three moves:

1. **Read what the market already says about price, before you set one.** Forum posts and customer interviews put the existing countertop nugget-ice-maker category at $2,000-$3,000, a range shoppers consistently called too expensive.
2. **Set a real-money test price below the ceiling the market just told you about.** First Build listed a $399 early-bird price on Indiegogo, against an anticipated $499 retail price once the product left crowdfunding, both far under the $2,000-$3,000 range shoppers had rejected.
3. **Let real money answer the question, not a survey.** The campaign's $150,000 funding goal wasn't just met; within days it crossed $1.3 million, and with two weeks still left in the campaign it had passed $2 million from 4,800 backers, already more than 13 times the goal (a $2 million-plus pace divided by the $150,000 target).

David Bland, who worked on the campaign, and Crowdfund Insider's contemporaneous August 2015 reporting both confirm these numbers.

The sequence generalizes into a repeatable technique for any product with no existing price data: find the ceiling the market has already set in its complaints, price a real-money test under that ceiling, and let the result decide the number that ships. It's one of the clearest pricing experiment examples of a price proven before a product existed at retail scale.

## Tools for running and analyzing pricing experiments

Match the tool to the job it does at that stage of the experiment.

| Job | Tool category | When you need it |
|---|---|---|
| Run the split and calculate significance | Experimentation or feature-flagging platform | From the day the test launches |
| Price the variants without breaking checkout | Billing or CPQ infrastructure | Before launch, so the test price charges correctly |
| Set the test's duration before it starts | Sample-size / power calculator | Before launch, so you know how long to run it |
| Screen an idea before spending any live traffic | Willingness-to-pay survey (such as conjoint analysis) | Before you build or price anything live |

An experimentation or feature-flagging platform runs the split and calculates significance once the test is live. Billing or CPQ infrastructure prices the variants without breaking checkout, which matters more for a pricing test than any other kind: a broken checkout doesn't just cost you data, it costs you the sale.

A sample-size calculator sets the test's duration before it starts. A willingness-to-pay survey, such as conjoint analysis, screens an idea before it spends any live traffic at all. Once you have a result, run it through the LTV:CAC calculator from the analysis step above to see whether it actually improved unit economics.

A pricing experiment is trustworthy when it does two things at once: isolate price the way Amazon's own randomization mechanics require, and stay inside the fairness and legal line that just ended Instacart's price-testing program and got Seattle's city council to vote for a ban.

Pick one price change you're already considering, write the falsifiable hypothesis first, and check your state and city's surveillance-pricing rules before you decide how to run a pricing experiment on it.

## FAQ

### Is it legal to charge different customers different prices for the same product?

It depends on the data used and the jurisdiction. Maryland, Connecticut and New Jersey already restrict surveillance pricing by state law, and Seattle's city council passed a citywide ban on algorithmic grocery pricing in September 2026, pending the mayor's signature. Check your state and city rules before you personalize a price test by segment.

### What real company got in trouble for a pricing experiment?

Instacart. It ran price tests on Eversight, an AI pricing company it acquired in 2022, and a Consumer Reports and Groundwork Collaborative investigation found price variances as high as 23% for identical items. Instacart ended the program on December 22, 2025, saying the differences weren't based on individual shopper data.

### Does the .99 trick (charm pricing) actually work?

Treat it as testable, the same as any other price point. Charm pricing is one experiment type to run, not a fact to assume: run it through a control group and let the result show whether it actually changes signups for your product.

### How long should a pricing experiment run?

Long enough to cover at least one full purchase or billing cycle, so you don't mistake a seasonal or timing effect for a price effect, and long enough to hit the sample size your minimum detectable effect requires. The sample-size math lives in growth experiment cadence.

### What's the difference between a pricing experiment and a pricing strategy?

A pricing strategy, like charm pricing, bundling or tiered plans, is one thing you could test. A pricing experiment is the controlled method, a hypothesis, an isolated variable and a control group, that tells you whether that strategy actually works for your customers.
