Almost every monetization case study has the same shape.
The platform was installed. Revenue went up. The platform receives the credit.
Sometimes that conclusion is right. Sometimes demand improved across the market, a holiday began, a high-value country grew, a new user-acquisition campaign changed the audience, or reporting moved from estimated to finalized revenue.
“Before versus after” measures a change. It does not prove a cause.
To prove incremental lift, you need to answer a counterfactual question: what would the same inventory have earned during the same period without the AI actions?
The measurement hierarchy
| Method | What it compares | Confidence | Main weakness |
|---|---|---|---|
| Before/after | This month against last month | Low | Time changed along with the system |
| Matched period | Similar weekdays or season last year | Low–medium | Audience and market may still differ |
| Switchback | Treatment and control alternate over time | Medium | Time blocks can influence each other |
| Concurrent holdout | Similar users or requests at the same time | High | Requires clean randomization and enough traffic |
| Multi-cell experiment | Separate levers and combinations | Highest | Traffic-hungry and operationally complex |
Use the strongest design your traffic supports. Do not pretend a weak design is strong because the result is favorable.
Define “lift” before the test
Start with a primary metric that represents money, not a proxy.
For most ad-monetized apps:
Revenue per eligible request = total recognized ad revenue ÷ requests eligible for the experiment
Revenue per DAU or ARPDAU is useful when the treatment can change impression frequency. Revenue per thousand impressions is not enough because a system can raise eCPM by filling fewer impressions.
Choose guardrail metrics beside the primary metric:
- Fill or match rate.
- Impressions per DAU.
- Ad latency and error rate.
- Session length.
- Day-1 and Day-7 retention.
- IAP conversion and revenue.
- Crash-free users.
- Policy or complaint rate.
UndrAds’ breakdown of static floors versus AI floor pricing explains the classic trap: a higher floor can raise reported eCPM while reducing fill enough to lower total revenue.
Write the hypothesis precisely
Weak hypothesis:
AI will improve monetization.
Testable hypothesis:
Allowing the system to adjust rewarded-video floors by country and hour will increase recognized ad revenue per eligible request by at least 5% over 21 days, without reducing fill by more than 2% or Day-7 retention by more than 0.5 percentage points.
The precise version defines:
- The action.
- The population.
- The success threshold.
- The evaluation period.
- The acceptable cost.
Agreeing on those before the data arrives prevents the team from moving the goalposts afterward.
Build a concurrent holdout when possible
Randomly assign eligible users—or, where technically appropriate, ad requests—to two stable groups:
- Control: current configuration and operating cadence.
- Treatment: AI actions inside agreed guardrails.
Both groups run during the same hours, holidays, market changes and acquisition campaigns. The difference between them is therefore more likely to come from the treatment.
User-level versus request-level assignment
Use user-level assignment when the treatment can affect experience, frequency, retention or later behavior. A player should not jump between different ad-frequency policies from one request to the next.
Request-level assignment may work for narrow auction changes that have no memory or user-experience effect. Even then, confirm that demand partners cannot recognize and price the two streams differently for reasons unrelated to the treatment.
Keep assignment stable
Once a user enters treatment or control, keep them there for the test. Re-randomizing creates contamination: the effects of yesterday’s treatment may appear in today’s control behavior.
Also exclude or separately analyze:
- Internal and test traffic.
- Newly launched countries or platforms.
- Ad units with configuration changes outside the experiment.
- Days affected by outages or reporting failures.
- Users exposed to overlapping monetization tests.
Exclusions should be defined before results are read. Removing “bad days” only from the losing variant is not analysis.
Do not pick a universal test duration
Duration depends on three things:
- Traffic volume.
- Natural daily volatility.
- How long the treatment takes to affect the metric.
A floor change can affect auction revenue immediately. Retention requires days or weeks. A test can therefore show a revenue signal before it is safe to roll out.
Google recommends running mediation A/B tests for at least two weeks, changing one setting at a time and collecting at least 10,000 ad requests before determining a result. Its reporting compares scaled revenue, eCPM, match rate and impressions. Google AdMob testing guide, Google test analysis
Those are useful minimums, not guarantees of statistical power for every app.
Calculate the incremental dollars
Assume:
- Control revenue per 1,000 eligible requests: $8.00
- Treatment revenue per 1,000 eligible requests: $8.56
- Eligible monthly requests: 50 million
Then:
Relative lift = ($8.56 − $8.00) ÷ $8.00 = 7%
Monthly incremental revenue = ($0.56 ÷ 1,000) × 50,000,000 = $28,000
Now subtract incremental costs:
- Platform fee or revenue share.
- New demand fees.
- Engineering and implementation time.
- Additional latency or infrastructure.
- Any measurable IAP or retention loss.
The decision metric is incremental contribution, not gross uplift.
Segment only after reading the total
Country, platform, format and placement cuts are valuable, but they can manufacture winners. If you inspect enough segments, some will outperform by chance.
Use this order:
- Read the pre-declared total result.
- Check guardrails.
- Inspect pre-declared strategic segments.
- Treat unexpected segment findings as hypotheses for the next test.
This prevents a failed global test from being marketed as a win because Android rewarded video in Canada improved on Tuesdays.
Separate model learning from treatment
Some AI systems need a learning period. Decide how that period enters the evaluation.
Three defensible approaches:
- Shadow mode: the system reads data and proposes actions without executing them, then the experiment begins.
- Burn-in period: treatment runs, but the first predefined days are excluded from the primary analysis for both business and statistical reasons.
- Learning included: measure the full adoption experience, including early underperformance, because that is what a buyer will actually experience.
Do not decide which method to use after seeing the curve.
Watch for six false lifts
| False lift | How it fools you | Check |
|---|---|---|
| Seasonality | Demand rose after launch | Concurrent control |
| Audience shift | More Tier-1 users arrived | Compare geo and acquisition mix |
| Reporting lag | One source finalized sooner | Use recognized revenue on matched data windows |
| Fill trade-off | eCPM rose while fewer ads served | Revenue per request and fill |
| IAP cannibalization | Ads earned more but purchases fell | Total ARPDAU and payer conversion |
| Novelty or ramp | Early action differs from steady state | Predefined burn-in and longer observation |
The article Why Does My Ad Revenue Drop and Recover a Few Hours Later? is useful here because it separates predictable demand cycles from genuine configuration problems.
A result table worth publishing
A credible case study should show more than one percentage.
| Field | Control | Treatment | Difference |
|---|---|---|---|
| Eligible users or requests | |||
| Recognized ad revenue | |||
| Revenue per eligible request | |||
| Fill rate | |||
| Impressions per DAU | |||
| Day-7 retention | |||
| IAP revenue per DAU | |||
| Test dates | Same period | Same period | — |
Also disclose:
- Randomization method.
- Exclusions.
- Traffic allocation.
- Confidence interval or uncertainty range.
- Vendor fees excluded or included.
- Whether the vendor participated in analysis.
Decision rules
Define four outcomes before launch:
Ship
The primary metric clears the minimum lift, guardrails hold and the result is sufficiently precise.
Extend
The direction is promising but uncertainty remains too wide.
Restrict
The total result is mixed, but a pre-declared segment shows a credible benefit. Deploy only there and run a confirmation test.
Stop
The primary metric misses, a guardrail breaks or operating cost consumes the gain.
This is the part vendors rarely show. A trustworthy system should make it easy to conclude that it did not work for a property.
The shortest honest answer
You prove AI AdOps lift with a stable concurrent control, a money-based primary metric, retention and IAP guardrails, predefined duration and transparent cost calculation.
Anything weaker can still inform a decision. It should not be described as proof.
If you are deciding which functions to test first, see 7 Signs Your Ad Stack Needs AI Automation. If you are comparing vendors, use UndrAds’ guide to AI ad monetization platforms and ask each one to show its holdout design, not only its best before-and-after chart.
FAQ
Is a before-and-after comparison ever useful?
Yes, for directional diagnosis when a randomized test is impossible. Match weekdays and seasonal periods, control for audience changes and describe the result as an estimate rather than causal proof.
What is the best primary metric?
Use recognized ad revenue per eligible request when testing auction changes. Use total revenue or contribution per DAU when the treatment can change impression volume, retention or purchases.
How much traffic should go into control?
There is no universal split. Preserve enough control traffic to measure the minimum lift within a reasonable period. High-volume properties may need only a small holdout; low-volume apps may require a balanced split.
How long should an AI AdOps test run?
Long enough to cover normal weekday and weekend cycles and observe the slowest important guardrail. Two weeks is a practical minimum for many mediation tests, while retention effects may require longer.
Should the AI vendor analyze its own test?
It can contribute, but the publisher should retain access to raw assignment and revenue data, agree on the method beforehand and be able to reproduce the result independently.



