UndrAds is this year's Diamond Sponsor at Mobidictum 2026, Istanbul 💎 - Check out the news

UndrAds
AI Ad Ops

How Do You Prove an AI AdOps Platform Actually Increased Revenue?

Rashmita Behera
Rashmita Behera
Sep 9, 2026
How Do You Prove an AI AdOps Platform Actually Increased Revenue?

Almost every monetization case study has the same shape.

The platform was installed. Revenue went up. The platform receives the credit.

Sometimes that conclusion is right. Sometimes demand improved across the market, a holiday began, a high-value country grew, a new user-acquisition campaign changed the audience, or reporting moved from estimated to finalized revenue.

“Before versus after” measures a change. It does not prove a cause.

To prove incremental lift, you need to answer a counterfactual question: what would the same inventory have earned during the same period without the AI actions?

The measurement hierarchy

Evidence hierarchy for proving AI AdOps incrementality
MethodWhat it comparesConfidenceMain weakness
Before/afterThis month against last monthLowTime changed along with the system
Matched periodSimilar weekdays or season last yearLow–mediumAudience and market may still differ
SwitchbackTreatment and control alternate over timeMediumTime blocks can influence each other
Concurrent holdoutSimilar users or requests at the same timeHighRequires clean randomization and enough traffic
Multi-cell experimentSeparate levers and combinationsHighestTraffic-hungry and operationally complex

Use the strongest design your traffic supports. Do not pretend a weak design is strong because the result is favorable.

Define “lift” before the test

Start with a primary metric that represents money, not a proxy.

For most ad-monetized apps:

Revenue per eligible request = total recognized ad revenue ÷ requests eligible for the experiment

Revenue per DAU or ARPDAU is useful when the treatment can change impression frequency. Revenue per thousand impressions is not enough because a system can raise eCPM by filling fewer impressions.

Choose guardrail metrics beside the primary metric:

  • Fill or match rate.
  • Impressions per DAU.
  • Ad latency and error rate.
  • Session length.
  • Day-1 and Day-7 retention.
  • IAP conversion and revenue.
  • Crash-free users.
  • Policy or complaint rate.

UndrAds’ breakdown of static floors versus AI floor pricing explains the classic trap: a higher floor can raise reported eCPM while reducing fill enough to lower total revenue.

Write the hypothesis precisely

Weak hypothesis:

AI will improve monetization.

Testable hypothesis:

Allowing the system to adjust rewarded-video floors by country and hour will increase recognized ad revenue per eligible request by at least 5% over 21 days, without reducing fill by more than 2% or Day-7 retention by more than 0.5 percentage points.

The precise version defines:

  • The action.
  • The population.
  • The success threshold.
  • The evaluation period.
  • The acceptable cost.

Agreeing on those before the data arrives prevents the team from moving the goalposts afterward.

Build a concurrent holdout when possible

Concurrent treatment and control design for an AI AdOps test

Randomly assign eligible users—or, where technically appropriate, ad requests—to two stable groups:

  • Control: current configuration and operating cadence.
  • Treatment: AI actions inside agreed guardrails.

Both groups run during the same hours, holidays, market changes and acquisition campaigns. The difference between them is therefore more likely to come from the treatment.

User-level versus request-level assignment

Use user-level assignment when the treatment can affect experience, frequency, retention or later behavior. A player should not jump between different ad-frequency policies from one request to the next.

Request-level assignment may work for narrow auction changes that have no memory or user-experience effect. Even then, confirm that demand partners cannot recognize and price the two streams differently for reasons unrelated to the treatment.

Keep assignment stable

Once a user enters treatment or control, keep them there for the test. Re-randomizing creates contamination: the effects of yesterday’s treatment may appear in today’s control behavior.

Also exclude or separately analyze:

  • Internal and test traffic.
  • Newly launched countries or platforms.
  • Ad units with configuration changes outside the experiment.
  • Days affected by outages or reporting failures.
  • Users exposed to overlapping monetization tests.

Exclusions should be defined before results are read. Removing “bad days” only from the losing variant is not analysis.

Do not pick a universal test duration

Duration depends on three things:

  1. Traffic volume.
  2. Natural daily volatility.
  3. How long the treatment takes to affect the metric.

A floor change can affect auction revenue immediately. Retention requires days or weeks. A test can therefore show a revenue signal before it is safe to roll out.

Google recommends running mediation A/B tests for at least two weeks, changing one setting at a time and collecting at least 10,000 ad requests before determining a result. Its reporting compares scaled revenue, eCPM, match rate and impressions. Google AdMob testing guide, Google test analysis

Those are useful minimums, not guarantees of statistical power for every app.

Calculate the incremental dollars

Assume:

  • Control revenue per 1,000 eligible requests: $8.00
  • Treatment revenue per 1,000 eligible requests: $8.56
  • Eligible monthly requests: 50 million

Then:

Relative lift = ($8.56 − $8.00) ÷ $8.00 = 7%
Monthly incremental revenue = ($0.56 ÷ 1,000) × 50,000,000 = $28,000

Now subtract incremental costs:

  • Platform fee or revenue share.
  • New demand fees.
  • Engineering and implementation time.
  • Additional latency or infrastructure.
  • Any measurable IAP or retention loss.

The decision metric is incremental contribution, not gross uplift.

Segment only after reading the total

Country, platform, format and placement cuts are valuable, but they can manufacture winners. If you inspect enough segments, some will outperform by chance.

Use this order:

  1. Read the pre-declared total result.
  2. Check guardrails.
  3. Inspect pre-declared strategic segments.
  4. Treat unexpected segment findings as hypotheses for the next test.

This prevents a failed global test from being marketed as a win because Android rewarded video in Canada improved on Tuesdays.

Separate model learning from treatment

Some AI systems need a learning period. Decide how that period enters the evaluation.

Three defensible approaches:

  • Shadow mode: the system reads data and proposes actions without executing them, then the experiment begins.
  • Burn-in period: treatment runs, but the first predefined days are excluded from the primary analysis for both business and statistical reasons.
  • Learning included: measure the full adoption experience, including early underperformance, because that is what a buyer will actually experience.

Do not decide which method to use after seeing the curve.

Watch for six false lifts

False liftHow it fools youCheck
SeasonalityDemand rose after launchConcurrent control
Audience shiftMore Tier-1 users arrivedCompare geo and acquisition mix
Reporting lagOne source finalized soonerUse recognized revenue on matched data windows
Fill trade-offeCPM rose while fewer ads servedRevenue per request and fill
IAP cannibalizationAds earned more but purchases fellTotal ARPDAU and payer conversion
Novelty or rampEarly action differs from steady statePredefined burn-in and longer observation

The article Why Does My Ad Revenue Drop and Recover a Few Hours Later? is useful here because it separates predictable demand cycles from genuine configuration problems.

A result table worth publishing

A credible case study should show more than one percentage.

FieldControlTreatmentDifference
Eligible users or requests
Recognized ad revenue
Revenue per eligible request
Fill rate
Impressions per DAU
Day-7 retention
IAP revenue per DAU
Test datesSame periodSame period
Blank fields are for you to fill.

Also disclose:

  • Randomization method.
  • Exclusions.
  • Traffic allocation.
  • Confidence interval or uncertainty range.
  • Vendor fees excluded or included.
  • Whether the vendor participated in analysis.

Decision rules

Define four outcomes before launch:

Ship

The primary metric clears the minimum lift, guardrails hold and the result is sufficiently precise.

Extend

The direction is promising but uncertainty remains too wide.

Restrict

The total result is mixed, but a pre-declared segment shows a credible benefit. Deploy only there and run a confirmation test.

Stop

The primary metric misses, a guardrail breaks or operating cost consumes the gain.

This is the part vendors rarely show. A trustworthy system should make it easy to conclude that it did not work for a property.

The shortest honest answer

You prove AI AdOps lift with a stable concurrent control, a money-based primary metric, retention and IAP guardrails, predefined duration and transparent cost calculation.

Anything weaker can still inform a decision. It should not be described as proof.

If you are deciding which functions to test first, see 7 Signs Your Ad Stack Needs AI Automation. If you are comparing vendors, use UndrAds’ guide to AI ad monetization platforms and ask each one to show its holdout design, not only its best before-and-after chart.

FAQ

Is a before-and-after comparison ever useful?

Yes, for directional diagnosis when a randomized test is impossible. Match weekdays and seasonal periods, control for audience changes and describe the result as an estimate rather than causal proof.

What is the best primary metric?

Use recognized ad revenue per eligible request when testing auction changes. Use total revenue or contribution per DAU when the treatment can change impression volume, retention or purchases.

How much traffic should go into control?

There is no universal split. Preserve enough control traffic to measure the minimum lift within a reasonable period. High-volume properties may need only a small holdout; low-volume apps may require a balanced split.

How long should an AI AdOps test run?

Long enough to cover normal weekday and weekend cycles and observe the slowest important guardrail. Two weeks is a practical minimum for many mediation tests, while retention effects may require longer.

Should the AI vendor analyze its own test?

It can contribute, but the publisher should retain access to raw assignment and revenue data, agree on the method beforehand and be able to reproduce the result independently.

Never Miss New Updates

Subscribe to our weekly newsletter and stay ahead with latest updates.