[ Case study ]
Donation campaigns ran on conviction and argument — leadership debates about messaging were settled by whoever argued longest. The donation page converted at 11%, and nobody knew which story actually moved donors.
CLIENT a regional food-security nonprofit — FOCUS Pre-commit the test plan
Representative examplesEvery case study in this library is an illustrative composite of the kind of engagement we deliver — written to show our method and standards, not to name clients.
Regional food security is this nonprofit's whole work, and its standing in the community is the asset every campaign spends and rebuilds. Individual giving runs through seasonal pushes — a winter appeal, a spring match, a giving-day effort — planned by a leadership team and executed by a two-person fundraising staff. Donation pages had been redesigned over the years by whoever had capacity, and each campaign's messaging reflected the last strong opinion voiced in the room. Fundraising meetings spent as much time relitigating old wording as planning the next appeal.
Donation campaigns ran on conviction and argument — leadership debates about messaging were settled by whoever argued longest. The donation page converted at 11%, and nobody knew which story actually moved donors.
Pre-commit the test plan before the season opens: three message hypotheses — impact-per-dollar, local visibility, beneficiary voice — each written down with a sample size and an end date, agreed by leadership in advance. Variants then run inside real campaign traffic, with dynamic text matching each campaign's audience, so results come from actual donors rather than synthetic visits. Every outcome, including the losers, goes into the campaign retro. The discipline is the point: once leadership signs the plan, the argument moves from the meeting to the field.
Just as important is what we ruled out, and why:
Three message hypotheses (impact-per-dollar, local visibility, beneficiary voice) were written down with sample sizes and end dates before the season opened.
Variants ran against actual campaign traffic with DTR matching each campaign's audience, not synthetic test traffic.
The winning message and the rejected ones went into the campaign retro — including the hypothesis the leadership team liked least.
Delivered by the growth pod — strategist over 6 weeks, with working increments reviewed with the client every week.
Obstacle
Mid-season, a local news story drove a traffic surge that unbalanced the variant split and threatened to decide the test by accident.
Handled: We held the test's end date and reported the surge's donors separately — their money was welcome, but a news spike should no more pick the message than the loudest meeting voice should.
Obstacle
The beneficiary-voice hypothesis needed a featured family's consent, and that conversation ran longer than the campaign calendar allowed.
Handled: The test launched with two variants and the third joined a week late with its window extended; the pre-committed plan bent without breaking.
Obstacle
Seasonal traffic thinned before the third variant reached its planned sample, leaving its question honestly unanswered.
Handled: We closed it on schedule and recorded 'inconclusive' in the retro — the plan earned its credibility precisely by refusing to declare a winner it did not have.
The headline: visit-to-donation conversion on the tested campaign versus the prior-year equivalent — 11% → 17%, read from Donation platform records. A second check: hypotheses resolved within the pre-committed windows at 2 of 3.
The argument pattern changed: a claim now arrives with 'that's testable' attached, and the season's plan holds the answer. Leadership stopped relitigating old wording because the retro documents why each choice was made, including the ones that did not work. The fundraising staff describe running experiments like professionals instead of defending instincts like staff. And the losing hypotheses stay valuable — the rejected stories are written down, so the organization remembers what it already learned about its donors instead of relearning it every winter.
The result was read from Donation platform records against the pre-engagement baseline over the stated window, with a guardrail check on hypotheses resolved within the pre-committed windows. Where platform-reported numbers and business outcomes differ, this record says which layer it is quoting.
What we would do differently
We would have started the third test earlier — seasonal traffic ran out on the last variant, and the honest answer is 'inconclusive', not a winner.
[ Related service ]
[ Related builds ]
2.9% 7.8%Replay-to-enrollment conversion across the next two cohorts versus the two prior
38 96Mobile Lighthouse performance score (lab diagnostic), with p75 LCP moving from 4.2s to 1.3s in field data
[ Next step ]
Next case study