AIREITER

AI Image

FLUX.2 ProGPT-Image 2Wan 2.7 Image ProGPT 4o ImageSeedream 5.0 ProSeedream V5 liteSeedream V4.5More

AI Video

Kling 3.0 Motion ControlSora 2 ProKling 3.0 TurboSora 2Kling 3.0Grok Imagine 1.5Veo 3.1More

LLM

Gemini 3.6 FlashGemini 3.1 ProKimi K3Gemini 3 ProGemini 2.5 ProClaude Opus 5Claude Fable 5More
Coming soonSeedance 2.5
Super ResolutionLyric Video GeneratorGPT Image 2 1K GeneratorGPT Image 2 Product Mockup GeneratorUse GPT-5.6 Online
API DOCSPRICING
BlogUpdatesLLM API GuideClaude API GuideKimi K3 API Guide
TEMPLATES
  • AIReiter
  • Blog
  • How to Read an Ad's Second-by-Second Retention Curve: Three Curves and One Percentile

How to Read an Ad's Second-by-Second Retention Curve: Three Curves and One Percentile

Last Updated: 2026-07-31 06:12:19

Most people study competitor creative like this: sort the creative center by play count, open the top performer, watch it three times, and pull out a gut summary like "the pacing is fast up front" or "it nails the pain point," then shoot their own version of it. The result lands somewhere lukewarm, so they open the next hit and repeat.

The problem with this routine isn't diligence, it's that they're reading the wrong thing. Play count and likes tell you this creative won, but not a word about which second it won at or what it won with. And those two things are exactly what you can carry into your next asset.

First, where the data comes from: every curve and every percentile in this piece is from each platform's public ad library and creative center, visible with your own account, no capture, no signatures, nothing bypassed. The creative center is something the platform built for advertisers in the first place, the data is sitting there, and the only issue is that most people don't know which few numbers to look at.

Play count is the least useful number of all

Play count is a number heavily polluted by the platform's distribution logic. A high play count could be because the content is good, or just because it was launched early, launched hard, and caught a cheap-traffic window. You can't separate "good creative" from "big budget" out of play count itself.

Likes are worse. A like is an emotional reaction, and it often runs in a different direction from the conversion you actually care about. A piece that makes viewers laugh gets plenty of likes, but they laugh and swipe away, and conversion is zero. Optimize toward likes and what you optimize is entertainment value, not purchase intent.

Both numbers share one fatal flaw: they're scalars with no time dimension. A piece of creative is something that unfolds along time, viewer behavior at second 2, second 5, and second 8 is completely different, and one total play count flattens that timeline entirely. Take a flattened number to guide work that is fundamentally about "what happens when," and the information was gone from the start.

The data that actually carries a time dimension is right there in the creative center, on that creative's detail page, and nobody clicks it.

Three curves, three inflection points

Open a top creative and, beyond the video itself, the platform gives you three per-second curves almost nobody studies: per-second retention, per-second clicks, per-second conversion. The x-axis is which second of the video (the platform tells you the total length), and the y-axis is, respectively, how many people are still watching at that second, how many clicked, how many converted. The platform even marks a few highlight frames for you, but marking them isn't reading them for you. A highlight frame tells you where there's a swing, not what the swing means.

The key to reading a curve isn't whether it's high overall, it's which second the inflection point is at. Each of the three curves' inflection points says one thing:

The retention curve's cliff is whether your opening held. Every retention curve starts at 100% and falls, and what matters is the shape of the fall. A steep cliff at second 2 or 3 means the opening didn't give a reason to stay in the first two seconds, and the people who swiped are the ones it didn't hook. The cliff's position maps directly to that one line, that one shot in your script. Cliff at second 3, the problem is the frame at second 3.

The click curve's peak is which second your selling point actually landed. Clicks don't happen evenly, they burst in one or two seconds, and whatever the frame is doing that second (showing the price, demonstrating the result, calling out the promise) is this creative's real hook. Plenty of people assume the hook is up front, and the click peak often tells you it's in the middle.

The conversion curve's lag is the most counterintuitive and most valuable one. Conversion almost never syncs with clicks, it always lags. The viewer is won and clicks at second 5, but doesn't convert until second 12 after seeing the benefit through. That lag (how many seconds between the click peak and the conversion climb) tells you how long a persuasion stretch your creative needs to go from "interested" to "willing to buy." Too long a lag means you laid the benefit out too slowly and viewers left before it landed.

Stack the three curves and a creative's whole story appears: hold them in a few seconds, win them at a certain second, then spend a few seconds turning the win into conversion. The second-marks of those three inflection points are the structure you can carry verbatim into your next script.

The CTR percentile: the absolute value is meaningless

Beyond the curves, the platform gives you one more number: this creative's CTR percentile.

The beginner's easiest mistake is reading CTR as an absolute value, "this one's CTR is 5%, not bad." Whether 5% is high or low can't be judged apart from a peer benchmark: in an impulse category 5% might be dead last, in a high-ticket category 5% might be top. The absolute CTR carries no information, the percentile does. Percentile 0.9 means it clicks better than 90% of the creatives in its industry, and that's a sentence you can decide from.

The percentile carries an easily-overlooked parameter too: the time window. The same creative can have very different percentiles computed over the last 7 days, 30 days, or 180 days. A creative that only recently gained volume sits high in a 7-day window and gets diluted in a 180-day one. Which period you're comparing over decides what the number means, so align the window to your launch cadence.

There's one more easily-missed curve: the creative's popularity over days, whether it's climbing, plateaued, or already turned down into decline. A creative with a high CTR percentile but a popularity curve that's already pointing down, if you copy it now, is chasing the last wave of traffic. What you should actually copy is a decent-percentile one whose popularity is still climbing. That curve decides not "copy or not," but "is it too late to copy now."

Feed a batch of curves to the model

For one creative's three curves, a human brain can still read them. But to get a reusable conclusion, you have to look at a batch, the dozen or two that ran well in one industry, and find what their inflection points have in common. Did they all hold retention before second 2? Are the click peaks all pinned to some fixed frame type? A human brain can't do this, because you'd have to align the point-by-point values of dozens of curves in your head at once, and this is exactly where the model comes on.

The right way is two steps, not one shot (this two-step method is the same thinking as the ad-copy clustering piece):

Step one, extract structured fields per creative. For each one, extract the inflection second of each of the three curves, the value drop around each inflection, and what the frame was doing at that second, into one structured record. This is high-concurrency mechanical work, one call per creative, hundreds to thousands per batch.

Step two, feed the whole table to the model at once for attribution. The prompt has to pin the output format. Don't ask "why are these creatives good" (you'll get a pile of "engaging opening" mush), have it give three parts for each common structure:

You'll receive per-second curve summaries for N high-performing creatives in
one category, each with three inflection seconds, drops, and a frame
description at that second. Find the recurring structures, and output each
one in the format below, no prose:

1. Inflection second
   Which second of the video this structure reliably appears at (give a
   range, not a single point)

2. What the frame is doing
   The specific action/element at that second, pointing to specific
   creative IDs, no vague descriptions allowed

3. Reusable script structure
   Abstract it into one instruction you can write into the next script
   (e.g. "an after-use result frame must appear before second 2")

4. Counterexample
   Are there creatives in this batch that break this structure and still
   perform well? If so, list them. If the structure doesn't actually hold,
   say so directly

Step one suits the cheap batch tier, and step two needs to hold a whole batch of curve values at once, which is the long-context tier's natural ground: the point-by-point arrays of dozens of creatives stack up to a few hundred thousand tokens easily. Two tiers together, and both cost and precision line up.

Attribution has to become a script structure

The third part of that prompt is where the whole thing earns its keep. An attribution that can't become a script instruction for the next asset is mush.

"These hits all have good rhythm" isn't attribution, it's a feeling, and you can't shoot to "good rhythm." "The creatives that hold viewers all show an after-use result frame, not the product itself, before second 2" is attribution, because it's a directly executable shooting instruction: next script, before second 2, put up a result frame.

To judge whether an attribution is useful, ask one thing: can it become a specific variable in the script? The second is a variable, the frame type is a variable, the length of the persuasion stretch is a variable. Keep the ones that become variables, cut the ones that don't. What you want in the end isn't a competitor analysis report, it's a variable table for the next batch of assets: what to put at which second, how many seconds of persuasion to leave, which frame the hook sits on.

One step past that variable table is the script, and one more is the finished spot. Script to finished spot is on-site image and video generation, and the full "keyword to curve attribution to script to finished spot" loop is in the keyword-to-finished-ad piece. This piece only handles the half that reads curves into a variable table.

The trap: don't let the model take correlation for causation

Batch attribution has one nearly unavoidable trap: the model leans hard toward calling correlation causation.

It sees a cat at the opening of ten high-retention creatives and tells you "the cat is the key to retention." But the truth might be that this batch happens to come from one team that likes cats, and the cat and the retention are both just results of that team's style, with no causation between them. Believe that attribution, force a cat into the next batch, and retention won't move.

This trap is the same disease as the model "recognizing an algorithm family and stopping the close look" in JS reverse engineering: once the model gives a self-consistent explanation, it stops doubting itself. The reverse-engineering workflow solves it by forcing the model to output a counter-evidence section in the prompt, and here it carries straight over: the fourth part of the prompt above, the counterexample, isn't optional, it's the test of whether an attribution can be trusted.

The use is direct: if the model claims "the result frame at second 2 causes high retention," push it to answer "are there creatives in this batch that put a result frame at second 2 and had poor retention." If there are, that causation is false, and something else sits between the frame and the retention. If the model searches the whole batch and finds no counterexample, the structure is worth entering into your variable table. A model that can't produce a counterexample, you can't use a single one of its attributions to guide creation, or you'll shoot to a pile of correlations and not understand why it doesn't work.

Which model for which step

This four-step pipeline asks completely different things of a model, and running one model for the whole thing either burns money or burns precision:

Step

Capability it needs

Pick

model id

Extract fields per creative (inflections + frame for N creatives)

Cheap, hundreds to thousands of calls at high concurrency

Claude Sonnet 5

claude-sonnet-5

Batch attribution (hold a whole batch of curves at once, find commonality)

Long context, a few hundred thousand tokens per input

Kimi K3

kimi-k3

Causation check and counterexamples (pick at the commonality)

Strong reasoning, willing to argue against itself

Claude Opus 5

claude-opus-5

Single-creative difference attribution (why does this one lag on conversion)

Mid reasoning, explains against a specific second

GPT-5.6 Sol

gpt-5.6-sol

The one most worth calling out is the counterexample step. It tests the same thing as the counter-evidence section in algorithm-family identification: whether the model will keep picking at a conclusion it just produced. A weaker model writes the counterexample section as a restatement of the conclusion ("the structure is validated across most creatives"), a stronger one actually digs out the creative that contradicts it. The quality of the counterexample section decides directly how many false causations you shoot to.

Don't take my word for the difference, test it. The protocol:

  1. Pull 10 to 15 top creatives in one industry, take each one's three per-second curves plus CTR percentile.

  2. Extract them into a structured table on the cheap tier (inflection seconds plus the frame at that second).

  3. Feed the whole table to claude-opus-5 and gpt-5.6-sol with the counterexample prompt from part 4 above.

  4. Look at one thing: does a creative with beautiful retention but poor conversion get pulled out as a counterexample by the model on its own. The tier that offers counterexamples unprompted is the one whose attribution you can write scripts from.

One round shows you the difference, more directly than any benchmark.

The friction isn't picking a model, it's the switching cost

Four models from three vendors, three SDKs, three auth schemes, three error formats. To use different models at different steps, the naive move is to wire three clients, and most people run the numbers, decide it isn't worth it, and end up on one model the whole way, using a tier with weak counterexamples at the batch-attribution step, shooting a pile of junk to false causation without knowing why.

AIReiter flattens that layer: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.

# Batch attribution + counterexamples: the reasoning tier
curl https://aireiter.com/api/v1/chat/completions \
  -H "Authorization: Bearer $AIREITER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-5",
    "messages": [{"role": "user", "content": "<attribution prompt + N curve-summary table>"}]
  }'

# Extract fields per creative: change the model field, leave the rest
#   "model": "claude-sonnet-5"
# Long context to hold a whole batch of curves:
#   "model": "kimi-k3"

If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.

On price, Claude models run at 30% off list and GPT models at half. For this flow the discount lands right on the main cost: per-creative field extraction is the most call-dense step, hundreds to thousands per batch, running on Claude Sonnet at 30% off; single-creative difference attribution runs on GPT-5.6 Sol at half. The batch-attribution step's long-context input is a few hundred thousand tokens, on Kimi K3, callable on the same key. The discount sits right on the two most expensive parts.

  • Get an API key

  • Try it without signing up: feed a few curves by hand first, compare whether the two models offer counterexamples, then decide which to wire in.

In closing

Reading creative by play count and likes is taking a scalar with the timeline flattened out and using it to guide work that is about time from end to end. The information was gone from the start.

The creative center has been showing you time-dimensioned data all along: the three per-second curves' inflection points tell you the seconds to hold, to win, and to convert, and the CTR percentile tells you whether this creative is any good among its peers, where the absolute value counts for nothing. Hand a batch of creatives' curves to a model for batch attribution and what you get isn't a "it's good" feeling, it's a variable table you can write into the next script, provided you force the model to give counterexamples and filter its assumed correlations into real causation one by one.

The model's place here is specific: it's the attributor that can align dozens of curves at once, find the commonality, and (well-prompted) is willing to contradict itself. It doesn't shoot the spot for you, and it doesn't decide causation for you. Counterexamples filter real causation, the variable table becomes the script, and the finished spot is the next piece's job.

>_AIReiter Model Directory

Fast API access to models related to this guide

Kimi K3

Chat

A long-context reasoning model for coding, writing, analysis, and agent workflows.

moonshotGet API Key >

Claude Opus 5

Chat

A premium Claude model for complex reasoning, coding, and long-context professional work.

anthropicGet API Key >

Claude Sonnet 5

Chat

A balanced Claude model for advanced reasoning, coding, and everyday work.

AnthropicGet API Key >

GPT-5.6 Sol

Chat

A premium GPT-5.6 text model for demanding coding, reasoning, and long-form agent work.

OpenAIGet API Key >

Claude Fable 5

Chat

A premium Claude model for deep reasoning and complex long-form work.

AnthropicGet API Key >

Recent Posts

GPT-5.6 Price Cut: What Luna and Terra Really Cost Now

2026-07-31

Invalid API Key: Diagnose 401 and 403 Before You Fix

2026-07-31

Fix OpenRouter 429: Provider Error or Rate Limit?

2026-07-31

DeepSeek V4 Flash vs GLM-5.2: 0731 Update Tested

2026-07-31
AIREITER

Questions? Contact us at
[email protected]

LLM

Gemini 3.6 FlashGemini 3.1 ProKimi K3Gemini 3 ProGemini 2.5 Pro

AI Video

Kling 3.0 Motion ControlSora 2 ProKling 3.0 TurboSora 2Kling 3.0

AI Image

FLUX.2 ProGPT-Image 2Wan 2.7 Image ProGPT 4o ImageSeedream 5.0 Pro

Blog

View All →

Company

Privacy PolicyTerms of ServiceRefund Policy

© 2026 AIReiter. All rights reserved.