How many ad concepts can your budget honestly test on Meta and TikTok?
Count concepts, not ads, and size the count by conversions. The arithmetic, a decision table and a weekly creative testing method for Meta and TikTok.
- Capacity is test budget divided by target CPA, divided by 25. That is concepts per cycle.
- The common rule of 2 to 3 times CPA per creative buys two or three conversions, which cannot rank anything.
- More concepts at a fixed budget raises the odds that the winner is luck.
- Meta's Creative diversity column is a visual check labelled as in development. Meta has published no similarity threshold.
The number of ads you can produce is no longer the limit. The limit is how many conversions you can afford to spend on finding out whether an idea works. That number is small, and most testing advice ignores it. This article sizes a creative test from conversions upward, for Instagram Reels and TikTok, and ends with a decision table and a weekly method.
Count concepts, because the platforms already do
An ad is a file. A concept is an idea: a distinct opening, person, format and message. Ten cuts of one talking-head video are ten ads and one concept.
Meta's delivery system has a retrieval stage, called Andromeda, that narrows tens of millions of candidate ads to a few thousand before ranking, according to Meta's engineering post of December 2, 2024. Meta's standing advice is to diversify creative. In late August 2026 it put a measure of that in front of advertisers. A Creative diversity column in Ads Manager rates each ad set Low, Medium or High. Common Thread Collective dates the launch to August 26, and Meta labels the metric as estimated and in development.
Two details matter. Novoads, reading Meta's help page on September 17, reports that the metric compares images and video thumbnails inside an ad set. For video that is one still frame. And UAMaster reports that campaigns mixing statics, videos with different hooks, UGC and carousels sometimes still received Low. Treat the column as a prompt to look, not a score to chase.
TikTok publishes no similarity rating. Its help centre asks for three to five different creatives per ad group and says large differences work better, especially in testing. Smart+ now accepts up to 50 creative assets in one ad and delivers the best combinations automatically, so on TikTok the platform does much of the variation work. Your job is the concepts.
What Meta has not said about similarity
Many guides state that ads above roughly 60% similarity are merged and that you should stay under 40%. Those figures appear in vendor and agency articles, such as one from the UGC platform Billo, with no Meta source behind them. Meta has not published them.
Marketing Brew asked Meta on September 28, 2026. A spokesperson said the company does not discuss how its ranking models represent individual creative assets and that its guidance is deliberately about outcomes. The useful detail came from an agency. Olivia Cripps of Billion Dollar Boy said two videos of the same person that differ only in on-screen text would count as the same creative.
So our working definition is practical, not numeric. A concept is different when a viewer would notice the difference in the first seconds with the sound off. Change at least two of these: the opening frame, the person on camera, the format, the message angle.
Does more creative volume win?
The sources disagree.
The case for volume rests on two datasets. Motion analysed 578,750 creatives from 6,015 Meta ad accounts, covering $1.29 billion in spend between September 1, 2025 and January 1, 2026. It defines a winner as a creative that spends at least 10 times its account's median and at least $500. Only about 4% of creatives qualify in accounts under $10,000 a month, and a little over 8% in accounts above $1 million. If winners are that rare, you need many attempts. Common Thread Collective adds that brands in its portfolio of more than 170 testing 20 or more new ads a month had 65% higher ROAS than those testing fewer than 10.
The case against is about dilution. Jon Loomer's July 2026 brief recommends one campaign and one ad set, because splitting budget makes the roughly 50 optimised events a week that end the learning phase harder to reach. In his words: "More ads won't get you there. Genuinely diverse ads will." Foxwell Digital notes that raising volume often lowers ad performance. AppsFlyer's September 30, 2026 report on 1.2 million creative variations from 1,400 apps found that the top 2% of gaming video creatives took 53.5% of spend in the first quarter of 2025 and 60.7% by the second quarter of 2026. Spend concentrated further even as output grew at the top spending tier.
Both sides are right, at different budgets. Motion states that its benchmarks describe associations, not causes, and its winner is defined by spend, not profit. Common Thread Collective has not published its method. Volume and budget travel together. Motion's median account under $10,000 launches about 3 creatives a week. Its median account above $1 million launches about 19, and has roughly double the hit rate. A large account can give each creative a readable sample. A small one cannot. The constraint is conversions per concept.
How many conversions before you can call a winner?
Published rules range widely. SparkUGC suggests spending 2 to 3 times your target CPA per creative. RocketShip HQ calls anything under 30 conversions per variant mostly noise. Stackmatix asks for 100 per variant and only two variants. Foxwell Digital puts the ideal at 50 and admits it is rarely practical.
The platforms give two anchors. Loomer applies the old rule of 50 optimised events in a week to Meta's creative testing tool. TikTok's learning phase FAQ also says 50 conversions, but its main learning phase page, updated in June 2026, says volatility starts to decline after about 25 results or 7 days.
The arithmetic explains the spread. Conversions arrive like random counts, so a count of n carries a 95% range of roughly plus or minus 2 divided by the square root of n. These are our calculations, not platform figures.
| Conversions on a concept | Approximate 95% range on its CPA |
|---|---|
| 3 | Not usable |
| 10 | Plus or minus 62% |
| 25 | Plus or minus 39% |
| 50 | Plus or minus 28% |
At 3 times CPA, a concept that truly converts at target still shows five or more conversions 18% of the time and zero or one 20% of the time. That rule is fine for one job: removing a concept that has spent 3 times CPA with no conversions. An on-target concept does that only 5% of the time.
For winners, 25 conversions with CPA at least 30% below your account average is a workable bar. An average concept clears it about 2% of the time. It is the same sample-size logic we apply to AI answers that keep changing.
A worked example with invented numbers
This brand is hypothetical. It spends $20,000 a month on Meta with a target CPA of $40. It reserves 20% for testing, the ceiling Meta recommends for its creative testing tool according to Loomer. That is $4,000, or 100 test conversions a month.
We simulated four ways to spend it, 400,000 runs each. A false winner means every concept is truly equal, yet the best one shows a CPA at least 30% better. A true winner found means one concept really converts 40% more per dollar and finishes first.
| Plan | Spend per concept | Expected conversions each | False winner | True winner found |
|---|---|---|---|---|
| 40 concepts | $100 | 2.5 | Almost always | 6% |
| 8 concepts | $500 | 12.5 | 50% | 46% |
| 4 concepts | $1,000 | 25 | 9% | 77% |
| 2 concepts | $2,000 | 50 | Under 1% | 96% |
The honest capacity is four concepts a month, one a week. Forty concepts feels productive and tells the brand almost nothing. The formula is simple: test budget divided by target CPA, divided by 25.
How many concepts fit your budget: a decision table
| Test conversions you can buy per week | What the budget supports | How to call it |
|---|---|---|
| Under 10 | No separate test. Add one new concept a month to the main ad set | Watch four weeks. Or test on a cheaper event such as add to cart |
| 10 to 25 | One new concept per week or fortnight against your best current ad | 25 conversions and 30% better CPA |
| 25 to 50 | One or two concepts per week | Same bar. Confirm in the main ad set |
| 50 to 125 | Two to five concepts, the range of Meta's creative test | Same bar, then 50 conversions to confirm |
| Over 125 | Five or more, in separate tests | Add a second test of hooks inside each winning concept |
On TikTok the structure differs. A split test compares two ad groups on a divided audience and names a winner only at 90% confidence. TikTok recommends at least 7 days and a budget that gives 80% power. Use it for the one comparison that matters most. Use ordinary ad groups of three to five creatives for screening.
A weekly testing method in seven steps
- Compute capacity. Test budget divided by target CPA, divided by 25. If the answer is below one, lengthen the cycle to two or four weeks. Do not add concepts.
- Write concepts, not variations. Each must differ from every live ad on at least two of opening frame, person, format and message. On Meta, check the Creative diversity column afterwards as a rough visual audit.
- Launch with even spend. On Meta, use the creative testing tool, which compares two to five ads at similar spend. On TikTok, use a split test for a head-to-head or one ad group per concept.
- Apply one kill rule. Pause a concept that spends 3 times target CPA with no conversions. Touch nothing else for 7 days.
- Read at the end. Winner: 25 or more conversions and CPA 30% better than the account average. Loser: 25 or more and 30% worse. Everything between is no result. Judge every concept on the same attribution setting.
- Confirm in live delivery. Move the winner into the main ad set and review after two more weeks. Loomer warns that once a test ends, Meta may not give the test winner the most budget.
- Log by concept. Record wins and losses against the idea, not the file. Hook and edit variations come after a concept wins.
Where this stops
A test tells you which idea earns spend today. It does not tell you how long it lasts. Meta's analytics team found in 2023 that the likelihood of a conversion was about 45% lower by the fourth repeated exposure to the same creative, so a winner starts wearing out as it scales. That is the reason to keep one honest test running every week, at the size your conversions allow.
Questions
How many conversions does an ad concept need before I can call it a winner?
Plan for about 25 conversions per concept and a cost per acquisition at least 30% below your account average. At two or three conversions the result is mostly chance. TikTok's help centre says volatility starts to decline after about 25 results, and Jon Loomer uses 50 optimised events in a week for Meta's creative test.
What counts as a different creative on Meta now?
Meta has not published a rule. Its Creative diversity column, added to Ads Manager in late August 2026, rates each ad set Low, Medium or High and is reported to compare images and video thumbnails. Agencies say the same person on camera with different text overlays is treated as one creative.
Is there an official creative similarity score threshold?
No. The figures often quoted, that ads above roughly 60% similarity are grouped and that you should stay under 40%, appear in vendor and agency articles with no Meta source. A Meta spokesperson told Marketing Brew in September 2026 that the company does not discuss how its models represent creative assets.
How is creative testing on TikTok different from Meta?
TikTok's split test compares two ad groups on a split audience and declares a winner at 90% confidence. It asks for at least 7 days and enough budget for 80% power. Meta's creative testing tool compares two to five ads at similar spend but does not report a confidence level, so you apply your own threshold.
Sources
- Engineering at Meta, Meta Andromeda, Supercharging Advantage+ automation with the next-gen personalized ads retrieval engine engineering.fb.com
- Analytics at Meta, Creative Fatigue, How advertisers can improve performance by managing repeated exposures medium.com
- Marketing Brew, How advertisers are adjusting to Meta's ad creative diversification best practices marketingbrew.com
- Common Thread Collective, Meta's New Creative Diversity Score Is Live in Ads Manager commonthreadco.com
- Novoads, Creative Diversity in Meta Ads, Why Meta Treats Look-Alike Ads as One Creative novoads.ai
- UAMaster, Meta Ads, Why Creative Diversity Matters More Than Creative Quantity uamaster.net
- Billo, Creative diversity, Meta's algorithm buckets look-alike UGC billo.app
- Motion, Creative Benchmarks 2026 motionapp.com
- Motion, Methodology and definitions, Creative Benchmarks 2026 motionapp.com
- Motion, Key Benchmarks and Insights motionapp.com
- Common Thread Collective, Meta Andromeda Killed Your ROAS. Here's the Only Way to Get It Back. commonthreadco.com
- AppsFlyer, State of Creative Optimization 2026 appsflyer.com
- Jon Loomer Digital, Meta's Creative Testing Tool, Setup, Strategy, and Results jonloomer.com
- Jon Loomer Digital, The Master Brief, My Complete Approach to Meta Ads jonloomer.com
- Foxwell Digital, Meta Creative Testing Frameworks Winning in 2026 foxwelldigital.com
- TikTok Business Help Center, About Split Testing ads.tiktok.com
- TikTok Business Help Center, Split Test Best Practices ads.tiktok.com
- TikTok Business Help Center, About Learning Phase ads.tiktok.com
- TikTok Business Help Center, Learning Phase FAQs ads.tiktok.com
- TikTok Business Help Center, Creative best practices for performance ads ads.tiktok.com
- TikTok Business Help Center, Smart+ Upgraded Experience ads.tiktok.com
- SparkUGC, How Much to Spend Per Creative on Meta Ads to Get a Real Read sparkugc.com
- RocketShip HQ, How to test ad creatives on Meta in 2026 rocketshiphq.com
- Stackmatix, A Structured Meta Ads Creative Testing Framework for Growth Teams stackmatix.com
See where your brand stands.
Your Digital Authority Score, who AI recommends for your target questions, and the placements that close each gap. Your first report is free.
Get your free report →Google's site reputation policy now depends on where the searcher is. What the EEA split changes
Since 30 August 2026 Google's site reputation manual action no longer affects results in the EEA. What changed, what did not, and an eight-step audit.
→2026.10.07 · 8 MINCreator handle or brand handle? Who should publish the collab post and run the ad
The same creator video can run from the creator's handle or yours. What two 2026 datasets show for Instagram collab posts, partnership ads and Spark Ads.
→2026.10.07 · 8 MINChatGPT now searches brand domains directly. A readiness audit for your own pages
In August 2026 ChatGPT began scoping many background searches to named domains. What the trackers agree on, where they differ, and a seven-step audit.
→