How YAPP measures podcast ad load

YAPP processes podcast episodes to identify the structure of each one, including where the advertising sits. A by-product of that work is a measurement: how much of a given episode, show, or genre is advertising. We publish those measurements on individual show pages, and monthly in the Podcast Ad Load Index.

This page explains exactly what we count, how we count it, and where the method's limits are. It is the reference for every number we publish. If you are deciding whether to cite our figures, read this first.

What counts as an ad

We count advertising and promotional segments inserted into or recorded within an episode's audio.

Counted

Not counted

The boundary case that matters most is the host-read sponsor segment that begins mid-sentence out of editorial content. We count these, and they are also where our boundary timings are least precise. See Known limitations.

What we measure

Every figure we publish is one of the following, defined here once and used consistently throughout.

MetricDefinition
Ad seconds detectedTotal duration of segments classified as advertising in a single episode
Ad minutes per hourAd seconds detected ÷ total episode duration, expressed per 60 minutes of runtime. Our primary comparative metric, because it is not distorted by episode length
Ad share of runtimeThe same quantity as a percentage of total episode duration
Ad breaks per episodeCount of discrete advertising segments, where two spots separated by less than 5 seconds of non-ad audio are treated as one break
Average episode durationMean total runtime of the episodes in the sample

A show's figure is time-weighted: its total detected advertising over its total runtime. Averaging per-episode rates instead would let a show that publishes both short segment clips and long full episodes be dominated by the clips, which carry proportionally more advertising, and would disagree with the per-episode averages printed beside it.

Genre and monthly figures are unweighted means of per-show values. A show with 200 episodes and a show with 12 episodes each contribute one value to their genre's mean, so a single prolific publisher does not dominate a genre figure. Where we report a median or a distribution instead, the chart says so.

How detection works

Each episode passes through four stages.

  1. Transcription. Episode audio is transcribed by a commercial speech-to-text service with speaker diarization, producing a transcript aligned to timecodes.
  2. Classification. A large language model labels transcript segments as advertising or content, working from the transcript together with its speaker and timing structure rather than from audio alone. A second model runs as fallback.
  3. Boundary refinement. Classification gives approximate boundaries, and transcript timing is not precise enough on its own. We refine each segment's start and end using acoustic fingerprinting, which recognises the repeated audio of a given spot across episodes and across shows and locates its exact edges.
  4. Inclusion. Not every processed episode is fit to publish. We exclude an episode when its duration is unknown or under 5 minutes, when it is a standalone trailer, when no advertising segment is detected, or when the detected advertising exceeds 60% of runtime — at that level the measurement has failed rather than the show being unusual. Individual segments shorter than 3 seconds are treated as detection artefacts and dropped.

The fingerprinting stage is what makes the measurements precise rather than approximate. The same 30-second spot appears across thousands of episodes; once its audio signature is known, its boundaries are exact wherever it recurs, regardless of what the transcript says. It does not reach every segment — see Known limitations for what that costs.

On vendors and models. This page describes the method but not the specific services or models behind each stage. That is a deliberate choice. Nothing in the figures depends on knowing them; what the figures do depend on is here — the definitions above, the inclusion rules, and the limitations we state. Where a component change could affect comparability between reporting periods, it is recorded in the changelog rather than left to be inferred.

Why there is no single ad load for an episode

This is the most important limitation on this page, and it applies to every figure we publish.

Most podcast advertising is dynamically inserted: ads are stitched into the audio when a listener requests the episode, not baked into a fixed file. Two people downloading the same episode can receive different ads, different numbers of ads, and different total ad durations, depending on their location, their device, the time of day, and what inventory the network has sold.

An episode therefore does not have one true ad load. It has a distribution of possible ad loads.

What we publish is ad load as observed by YAPP's processing pipeline at the time each episode was fetched. Our vantage point is a single set of requests from our own infrastructure. A listener in a different market may hear materially more or less advertising in the same episode.

Our figures are sound for the comparisons we make with them — relative ad load between shows and genres, and change over time, all measured from a consistent vantage point. They are not sound as a claim about what any particular listener hears. We do not report an episode's ad load as a property of the episode, and we ask anyone citing these numbers not to either.

Baked-in advertising, common on independent and older shows, does not have this problem. We do not currently report baked-in and dynamically inserted segments separately.

What is in the sample

Our sample is shaped by which shows YAPP's users listen to and request, so it is not a random sample of all podcasts. It skews toward shows with active listenership, and toward English-language shows.

Monthly figures are grouped by the month an episode was published, not the month we happened to process it. Processing runs in batches and reaches back through a show's archive, so grouping by processing date would describe our own schedule rather than the state of podcast advertising.

Inclusion thresholds, applied to every published figure:

Each published month states its sample size — shows, episodes, and total hours — alongside the headline figure. Where a genre's sample is small enough that the figure should be read cautiously, we mark it on the table.

Accuracy

We have not yet evaluated detection against an independently labelled corpus, so we do not publish precision or recall figures.

What we can state: episodes and segments that fail the inclusion rules above are excluded from published statistics, and boundary timings for recurring spots are derived from acoustic fingerprinting rather than transcript timing alone. We are building a labelled validation set and will publish measured accuracy figures, and the set itself, when we have them.

Known limitations

Beyond dynamic insertion, the following affect our figures.

Corrections, claims, and opt-out

If you produce a show and our figures look wrong to you, tell us and we will check them. If you would rather not be included, we will remove you.

Submit a correction or opt out: yapp.at/creators

Opt-out takes effect immediately: the show's page stops resolving as soon as you confirm, and the show is excluded from the directory, the sitemap, rankings, and all future datasets. Corrections to a published monthly dataset are appended as dated notes rather than applied silently; archived months are never edited in place.

Smart Skip runs on the listener's device. YAPP does not modify, re-host, or redistribute your audio, and does not remove anything from your feed. Each listener's app decides locally, for that listener, what to play.

Licence and citation

The Podcast Ad Load Index datasets are published under Creative Commons Attribution 4.0 (CC BY 4.0). Use them commercially, republish them, build on them — attribution is the only condition.

Each month's page carries its own citation line, in this form:

YAPP Podcast Ad Load Index, September 2026. https://yapp.at/ad-load-index/2026-09/

Full monthly data is available as CSV and JSON from each month's page.

Changes to this methodology

Methodology changes are versioned and listed here, because a figure is only comparable across months if the method behind it is known.

VersionDateChange
1.018 September 2026Initial publication
1.118 September 2026Added the minimum-shows threshold for a month to be reported, after backfilling surfaced months with one or two qualifying shows
1.218 September 2026A show's figure is now time-weighted across its runtime rather than an unweighted mean of per-episode rates. Shows publishing a mix of short clips and long episodes read materially lower; the August 2026 dataset was recomputed and carries a correction note

Where a change affects comparability with earlier months, we say so on the affected Index pages as well as here.