NEXSUM_LABS
  1. Home
  2. Work
  3. A skincare brand built a creative-testing cadence that stopped depending on one winning ad
Book a call

[ Case study ]

BeautyMeta AdsConversions API with dedup verificationCreative testing matrixShopify reconciliation

A skincare brand built a creative-testing cadence that stopped depending on one winning ad

One UGC ad carried 70% of account revenue. When it fatigued, revenue dipped for weeks while the team scrambled — there was no testing system, no decision rules, and creative debates were settled by whoever felt strongest.

CLIENT a direct-to-consumer skincare brand — FOCUS Set the cadence and the kill rules

Facebook & Meta AdsPaid MediaFacebook & Meta AdsBeautyRepresentative example
Client
a direct-to-consumer skincare brand
Industry
Beauty
Engagement
8 weeks — growth pod — paid social specialist + creative strategist
Service
Paid Media / Facebook & Meta Ads
Headline outcome
Revenue share of the single best ad, replaced by a portfolio of five proven concepts: 70% → 34%, read from Ads manager reconciled with Shopify orders

Representative examplesEvery case study in this library is an illustrative composite of the kind of engagement we deliver — written to show our method and standards, not to name clients.

Where they started

The skincare brand sells a focused range direct to consumers, built almost entirely on Meta with a single content creator who shoots product and lifestyle together. Revenue had grown on the strength of one user-generated ad the account kept feeding; everything else in the library was leftover. Orders live in Shopify, margins are healthy enough to spend on acquisition but not to waste, and creative decisions had always been argued rather than decided — loudest voice in the room wins.

What it was costing

One UGC ad carried 70% of account revenue. When it fatigued, revenue dipped for weeks while the team scrambled — there was no testing system, no decision rules, and creative debates were settled by whoever felt strongest.

What they could see

  • One user-generated ad accounted for most account revenue; when its frequency climbed, weekly revenue sagged until someone manually refreshed it.
  • Creative debates ran for days — which concept to scale, which to kill — settled by whoever argued hardest, with no number to check.
  • Meta's reported purchases ran ahead of Shopify's order count by a margin nobody could attribute to anything specific.
  • New concepts launched occasionally, spent without a threshold, and were abandoned by mood rather than decision.

The constraints we worked inside

  • The brand's content creator was one person; the test volume had to fit one shoot day a fortnight.
  • Margins were healthy but not infinite — kill rules had to prevent emotional spend on losing concepts.
  • Meta attribution needed reconciling against the platform's orders before any 'win' was declared.

What had been tried before

The team ran occasional A/B tests whenever a creative debate grew heated enough.
Tests had no kill thresholds or budget discipline, so losing variants ran for weeks and the results were argued about rather than acted on.
A freelance editor was brought in to refresh the winning ad's visuals.
Re-cutting one concept into new skins extended its life briefly but deepened the dependence on a single idea instead of building a portfolio.
Budget was raised on the winning ad whenever weekly revenue sagged below plan.
Feeding a fatiguing ad spends more for less each week; the cure accelerated the fatigue while the library stayed empty behind it.

What we proposed

We proposed a fortnightly testing cadence sized to one creator's shoot day: three concepts per batch — hook variant, format variant, offer variant — with kill thresholds agreed before launch, so losers die at the data line instead of in debate. Winning concepts graduate into the rotation with iteration budget rather than retiring. Attribution was reconciled weekly against Shopify orders, with the Conversions API's deduplication verified before any win was declared — a 'win' means the P&L sees it.

Just as important is what we ruled out, and why:

  • Outsourcing creative production to an agency benchThe brand's voice lived with its one creator and her instinct about what fits the feed — an external bench multiplies volume while severing the voice the tests depend on.
  • Fully automated rules-based scaling toolsKill rules were the discipline the account lacked, but a rules engine can't reconcile attribution against Shopify orders — automation on unverified numbers accelerates bad decisions.
  • Broadening audiences to find new winnersThe constraint was creative supply, not audience supply; broader audiences burn budget faster against the same handful of concepts.

How the work ran

01Set the cadence and the kill rules

A fortnightly batch of three concepts (hook-variant, format-variant, offer-variant) with pre-committed kill thresholds — losers die at the data line, not in debate.

02Verify attribution before celebrating

Meta's reported results were reconciled against Shopify orders weekly, with deduplication checked via the Conversions API setup — a 'win' means the P&L sees it.

03Graduate winners into the rotation

Winning concepts get iteration budget (new hooks on proven bodies) rather than retired — the rotation keeps the account from single-point-of-failure.

Delivered by the growth pod — paid social specialist + creative strategist over 8 weeks, with working increments reviewed with the client every week.

The stack, and the reasoning

Meta Ads
The audience and the creative supply both live on Meta; the account needed a system, not a platform change, and Meta's testing tools fit one creator's volume.
Conversions API with dedup verification
Server-side events with verified deduplication meant browser loss couldn't distort the testing reads — a kill rule is only as good as the number it kills on.
Creative testing matrix
The matrix enforced concept-versus-variant discipline — hook, format, and offer tested separately so a loss says what died, not just that something did.
Shopify reconciliation
Weekly order-data checks kept platform wins honest; a concept graduates only when the P&L sees it too.

What went wrong

Obstacle

From the first reconciliation, two sources of truth disagreed: Meta's reported purchases ran ahead of the payment gateway's settled orders, and the gap split almost entirely along browser events that fired twice — once in the browser, once through the Conversions API.

Handled: We deduplicated at the gateway side — one event ID per order, shared by the browser and server events — re-based the account's history against the gateway's settled orders, and only then ran the first testing batch on clean numbers.

Obstacle

The creator's first shoot day produced fewer usable concepts than the matrix called for — half the footage worked as texture but not as a testable hook.

Handled: We re-cut the batch plan around hook variants on proven bodies, which the one hour could support, and moved full-concept production to every second batch instead of every one.

Obstacle

The first kill rule triggered on a concept the founder personally loved, and the initial instinct was to override the threshold and spend another week proving it wrong.

Handled: We showed the pre-committed threshold beside the concept's numbers, ran one controlled extension with a hard cap, and the account accepted the rule as real when the extension confirmed the kill.

How we worked together

Cadence
A fortnightly review aligned to each shoot batch — results of the last matrix read first, next batch's brief agreed second, every time in the same order.
Client side
The founder attended both meetings; the creator owned production and a part-time marketing hire kept the testing log between batches.
Decisions
Kill thresholds were pre-committed in writing per batch, so most decisions made themselves; the founder reserved one override per month and rarely used it.
They provided
The creator's shoot hours, Shopify read access for reconciliation, and product margins in ranges so kill thresholds could be set against real economics.

What changed

The headline: revenue share of the single best ad, replaced by a portfolio of five proven concepts70% → 34%, read from Ads manager reconciled with Shopify orders. A second check: weeks where revenue dipped on creative fatigue since the cadence started at 0.

Creative meetings got shorter because the numbers close the argument. The creator briefs hooks now, not vibes, and describes her shoot days in variants produced rather than ads finished. Revenue no longer swings on one ad's fatigue — the rotation carries weeks the old account would have dipped through. The founder describes the testing log as the first marketing document she actually reads. Decisions that used to take a meeting now take a threshold.

The result was read from Ads manager reconciled with Shopify orders against the pre-engagement baseline over the stated window, with a guardrail check on weeks where revenue dipped on creative fatigue since the cadence started. Where platform-reported numbers and business outcomes differ, this record says which layer it is quoting.

What they own now

  • The testing matrix template — concept definitions, variant types, thresholds per batch.
  • The kill-rule sheet with pre-committed thresholds, reusable for every batch.
  • A verified CAPI setup with dedup checks documented for the developer.
  • The weekly Shopify reconciliation routine with the signed-off comparison format.
  • The rotation schedule showing which concepts hold iteration budget this quarter.

What we would do differently

We would have defined 'concept' vs 'variant' before the first batch — the first test matrix conflated them and proved nothing.

Paid MediaFacebook & Meta AdsBeautyMeta Ads

Next case study

A manufacturer's Meta lead gen stopped producing leads the sales team ignored