Learn. Build. Test. Repeat.
The simple loop behind how I market, teach and build, and why a fast feedback loop beats waiting for certainty. Here is how to run it on your own marketing, from hypothesis to decision.

Quick answer
A marketing experimentation framework is a repeatable loop: learn what the market is telling you, build the smallest test that could prove you wrong, run it long enough to trust the result, then decide using rules you set in advance. Its advantage is speed of learning, not certainty. Teams that close the loop faster compound insight while others wait for perfect data.
Most marketing plans are written as if the author already knows what will work. I have written plenty of them, for B2B firms, for agency clients at Paradigm Media Networks, and for my own products, and after seven years in growth, SEO and performance marketing the uncomfortable truth is that the confident plans were rarely the ones that won. The wins came from loops: small bets, measured carefully, repeated without ego.
So this is not a post about a clever framework. It is about a habit I call Learn. Build. Test. Repeat. It shapes how I run campaigns, how I teach, and how I build products. If you treat certainty as the prerequisite for action, you will always be late. If you treat a fast feedback loop as the strategy, being wrong becomes cheap, and being right becomes repeatable.
Key takeaways
- The edge in marketing is how fast you learn, not how much you know before you start.
- Every test needs a written hypothesis that can fail, a single primary metric and a decision rule set before launch.
- Minimum viable tests answer one question cheaply; they are not miniature versions of the full campaign.
- Stopping a test the moment it looks good inflates false positives dramatically. Fix your sample or duration in advance.
- An experiment log turns individual tests into institutional memory, and that is where compounding comes from.
Why certainty is the most expensive thing in marketing
Waiting feels responsible. You want one more month of data, one more stakeholder review, one more competitor audit. But in a market where algorithms, platforms and buyer behaviour shift every quarter, the information you are waiting for usually expires before it arrives.
The evidence that experts cannot reliably predict winners is not anecdotal. In their Harvard Business Review piece on online experiments, Ron Kohavi and Stefan Thomke describe a Bing employee's idea for changing how ad headlines were displayed. It was judged low priority and shelved for over six months. When someone finally tested it, it lifted revenue by 12%, which they estimated at more than $100 million a year in the US alone. Their point was blunt: even experts struggle to assess new ideas. A separate paper from the same Microsoft experimentation team notes that many ideas move key metrics by around 1% and are not well estimated in advance.
If Microsoft's product managers, with all their data, could not tell a nine-figure idea from a low-priority one, your gut feeling about a landing page headline deserves some humility too.
Certainty is a luxury. A faster feedback loop is a strategy.
The mistake to avoid here is confusing speed with recklessness. A fast loop is disciplined: it is small, measured and reversible. Recklessness is a big, unmeasured, irreversible bet made quickly. The two look similar from the outside and produce opposite results.
A tactical example: a founder wants to know whether a webinar will generate qualified leads. The slow route is to build the full funnel, design assets and buy a month of ads. The loop route is a simple registration page, one ad set and a short email to the existing list, with a pre-agreed threshold for registrations from the target audience. Within days you know whether the topic has pull, before you have spent on production.
The loop, stage by stage, and where it comes from
Learn. Build. Test. Repeat. is my own phrasing, but the idea stands on well-tested shoulders. Eric Ries's Lean Startup popularised the Build-Measure-Learn feedback loop, with the minimum viable product as the vehicle and "validated learning" as the unit of progress. Earlier, US Air Force Colonel John Boyd developed the OODA loop (observe, orient, decide, act) in the 1970s, arguing that whoever cycles through it faster can get inside an opponent's decision cycle and gain the advantage.
I put Learn first deliberately. Most marketers I work with do not suffer from a lack of building; they suffer from building before they have looked properly at what is already working: search queries, sales call notes, comment sections, competitor ads, their own analytics.
| Model | Stages | Built for | What it adds to marketing |
|---|---|---|---|
| Build-Measure-Learn (Lean Startup) | Build, measure, learn, then pivot or persevere | Startups under extreme uncertainty | The MVP mindset and the explicit pivot-or-persevere decision |
| OODA (John Boyd) | Observe, orient, decide, act | Competitive, fast-moving situations | Tempo matters, and orientation (how you interpret data) shapes every decision |
| Controlled experimentation (A/B testing) | Hypothesis, randomised test, analysis | Products and sites with meaningful traffic | Statistical discipline: sample size, significance, no peeking |
| Learn. Build. Test. Repeat. | Learn, build, test, repeat | Marketers, teachers and builders working with limited budgets | Research before building, and documentation as a non-negotiable part of "repeat" |
The operational reality is that the stage most teams skip is Repeat. They run a test, glance at the result and move to the next shiny idea without asking what the result changes. Repeat is not "do it again". It is "carry the lesson into the next cycle", which is why it pairs so naturally with building marketing systems over one-off campaigns.
Here is my first contrarian observation: most small and mid-sized businesses should not be running classic A/B tests on most things. They lack the traffic. What they should be running are bold, clearly framed experiments with large expected differences, measured against a sensible baseline, and accepted as directional evidence rather than proof. Pretending a 400-visit split test is scientific is worse than admitting you made an informed judgement call.
Start with a hypothesis you could lose
A hypothesis is not "let's try video ads". It is a falsifiable statement that names a change, an audience, an expected effect and a reason. Without one, every result is interpretable as a success, which means nothing was learned.
The format I use and teach is simple: Because we observed [evidence], we believe that [change] for [audience] will cause [measurable effect]. We will know this is true if [primary metric] moves by [threshold] within [time or sample]. The "because" clause matters most. It forces the Learn stage to happen and it tells you what to revisit if the test fails.
The common failure is stacking variables. A new headline, new image, new offer and new audience in one variation produces a result you cannot explain. That is why disciplined tests change one variable at a time. If you want to understand why creative so often decides paid social performance, I have covered it in why Meta ads creative is the targeting.
What most people miss
A hypothesis that fails is not a wasted test if the "because" clause was specific. It tells you your reading of the evidence was wrong, which is often more valuable than a lucky win you cannot explain or repeat.
A tactical example from SEO: "Because Search Console shows our service page ranking on page two for comparison-style queries, we believe adding a comparison table and FAQ section will lift clicks from those queries. We will judge it on clicks to that URL over 28 days against the previous 28, controlling for seasonality with a similar untouched page." It is not a randomised trial, and I say so in the log. It is still a far better decision than rewriting forty pages on a hunch. The groundwork for tests like this sits in your technical SEO foundations; if crawling or indexing is broken, the experiment measures the plumbing, not the content.
Minimum viable tests: build only what answers the question
The Build stage is where enthusiasm does the most damage. A minimum viable test is the cheapest artefact that can answer one question. It is not a scaled-down version of the full campaign; it is a different object designed for learning.
Ask what you need to believe before you invest properly. If the question is "will anyone want this offer?", you need a page and some traffic, not a finished product. If the question is "does this angle beat our current one?", you need two ads with the same budget and audience, not a rebrand. Usually the riskiest assumption is demand, not design, so test demand first.
| Question you need answered | Over-built approach | Minimum viable test |
|---|---|---|
| Is there demand for this service? | Full website, brochure, launch campaign | One landing page, one offer, a small paid budget and direct outreach |
| Which message resonates? | Three complete campaigns with new creative suites | Two or three ad variations differing only in the hook |
| Will this content format rank or get cited? | A 50-post content calendar | Three posts in the format, tracked for impressions, clicks and AI citations |
| Can AI speed up this workflow? | A company-wide tool rollout | One person, one recurring task, timed before and after |
That last row is where many teams are right now. I wrote about it in more depth in AI in marketing workflows: start with one task and measure it, rather than buying a platform and hoping.
Pro tip
Before building anything, write down what result would make you stop. If you cannot name a result that would kill the idea, you are not testing; you are decorating a decision you have already made.
Sample sizes, patience and the peeking problem
Fast loops do not mean impatient tests. This is the paradox people struggle with: you want many cycles per quarter, but each cycle must run long enough to be trusted.
The classic warning comes from Evan Miller's essay on how not to run an A/B test. If you check results continuously and stop the moment you see significance, your real false positive rate can climb far above what the dashboard implies. In his worked example, checking after every observation pushes a nominal 5% false positive rate to 26.1%. His remedy is to fix the sample size in advance and not act on the result until it is reached.
Platform guidance points the same way. Google Ads' custom experiments documentation recommends a 50% split to give the best comparison between original and experiment campaigns, and for audience-list experiments with cookie-based splits it recommends at least 10,000 users in the list. That is a useful reality check for small accounts: if your audience is a few hundred people, you are running a directional test, not a statistically robust one, and your decision rule should reflect that.
Common mistake
Calling a winner after three days because one ad has a lower cost per lead. Early results swing with day-of-week effects, learning phases and small numbers. Set a minimum duration that covers at least one full weekly cycle and a minimum conversion count before anyone is allowed to look at the scoreboard.
In practice, I handle this with two numbers written on every test card before launch: a minimum run time and a minimum volume of the primary event. Whichever comes later is when we decide. For low-volume businesses, I move the primary metric up the funnel (landing page conversion instead of closed deals) and accept that I am testing a leading indicator, then confirm with the lagging one over the following cycles.
Decision rules and the experiment log
The test is not finished when the data arrives. It is finished when a decision is made and recorded. This is my second contrarian view: running more tests is not an experimentation culture. Deciding is. An organisation that runs fifty tests and acts on five has a reporting habit, not a learning habit.
Kohavi's team at Microsoft described having over 200 concurrent experiments running on a given day at Bing. You do not need anything like that scale. You need the same discipline in miniature: every test has an owner, a rule and a written outcome. Here is the card I use for every experiment, whether it is an ad, a page, an email or a lesson format.
- HypothesisThe "because, we believe, will cause" statement, including the evidence it rests on.
- Primary metric and guardrailsOne metric decides the test. One or two guardrail metrics (such as lead quality or refund rate) can veto it.
- Minimum run and minimum volumeWritten before launch. No decision is made before both are met.
- Decision ruleScale if the primary metric beats control by the agreed threshold and guardrails hold; kill if it underperforms; iterate if the result is inconclusive but the "because" still looks right.
- Outcome and lessonWhat happened, what you decided, and the one sentence a future teammate needs to know.
- Next loopThe follow-up hypothesis this result suggests. This is the "repeat" that most teams leave out.
The log itself can be a spreadsheet or a Notion database. What matters is that it is searchable and honest, including the failures. Over a year, it becomes the most valuable strategy document a business owns, because it records what this market did, not what a blog post claimed markets generally do.
The mistake to avoid is editing history. When a test fails, there is a temptation to quietly re-label it as "brand awareness" or drop it from the log. Do not. A documented loss is worth more than an undocumented win, because the loss stops someone repeating it in six months.
Your first loop this week
This habit shapes my teaching as much as my client work. When I plan a workshop, I treat the format as a test: one change at a time, one question I want answered, and notes I review before the next session. The same loop that improves a Google Ads account improves a lesson plan or a product listing. It is a way of working, not a marketing trick.
If you want to start, pick one thing this week. Write one hypothesis with a "because" clause. Build the smallest version that could prove it wrong. Set your minimum run and volume. Decide using the rule you wrote, log it, and write the next hypothesis before you close the tab. Then do it again next week. It is not a clever framework. It is a habit, and habits compound.
Frequently asked questions
What is a marketing experimentation framework?
It is a repeatable process for testing marketing ideas before committing full budget. A good framework covers how you generate hypotheses from evidence, how you design the smallest test that can answer one question, how long you run it, which metric decides the result and how you record what you learned. The aim is not to prove you were right but to learn faster than competitors, so that each cycle makes the next decision better informed.
How is Learn. Build. Test. Repeat. different from Build-Measure-Learn?
It borrows the core idea from Eric Ries's Build-Measure-Learn loop but starts with Learn on purpose. Many marketers build before looking closely at what already works: search data, sales conversations, competitor ads and their own analytics. It also treats Repeat as carrying a documented lesson into the next hypothesis, not simply running another test. That makes it better suited to marketers, educators and small teams working with tight budgets rather than funded startups.
How long should a marketing test run?
Long enough to meet two thresholds you set before launch: a minimum duration and a minimum volume of the primary conversion event. As a working rule, cover at least one full weekly cycle so day-of-week swings even out. Do not stop early because one variation looks ahead; repeatedly checking and stopping at the first sign of significance can inflate false positives far above the level your dashboard reports. Decide only once both thresholds are met.
Can small businesses run A/B tests without much traffic?
They can run experiments, but they should be realistic about what those experiments prove. With low traffic, test bold changes that are likely to produce large differences, measure a metric higher in the funnel such as landing page conversion, and treat results as directional evidence. Google Ads, for example, recommends at least 10,000 users for some audience-list experiments. If you are far below that, combine test data with qualitative signals and confirm over several cycles.
What should an experiment log include?
At minimum: the hypothesis with its supporting evidence, the primary metric and any guardrail metrics, the minimum run time and volume, the decision rule, the actual outcome, the decision taken and one sentence on what was learned. Add the next hypothesis the result suggests. Keep failures in the log. Over time it becomes a record of how your specific market behaves, which is far more useful than generic best practice.
Where does the OODA loop fit into marketing?
John Boyd's OODA loop (observe, orient, decide, act) was developed for fast-moving competitive situations. Its lesson for marketers is that tempo matters and that orientation, meaning how you interpret what you observe, shapes every decision. Applied to marketing, it reminds you to keep reading the market between tests, update your assumptions when platforms or buyer behaviour shift, and avoid acting on stale mental models simply because a plan was approved months ago.
Sources
- The Surprising Power of Online Experiments — Ron Kohavi and Stefan Thomke, Harvard Business Review, 2017
- Online Controlled Experiments at Large Scale — Kohavi, Deng, Frasca, Walker, Xu and Pohlmann, KDD, 2013
- The Lean Startup Methodology: Principles — The Lean Startup (Eric Ries)
- How Not To Run an A/B Test — Evan Miller, 2010
- Set up a custom experiment — Google Ads Help
- OODA loop — Wikipedia

