Where AI actually helps a marketing team (and where it doesn't)
AI is excellent at some marketing work and confidently wrong at the task next door. Here is how to decide what to hand over, what to protect, and how to build guardrails that hold.

Quick answer
AI in marketing workflows earns its place on work that is repeatable and easy to check: research synthesis, first drafts, repurposing, ad variations, data clean-up and reporting. It is weak at strategy, judgement, brand voice and factual accuracy. Build it in with three guardrails: a written brand context, a named human reviewer at each stage, and a source check before anything is published.
Most marketing teams I audit are not short of AI tools. They are short of a decision about which work AI should touch. Over seven years in growth marketing, SEO and performance work, first inside B2B firms and now running Paradigm Media Networks, I have watched every new tool get the same treatment: bought with enthusiasm, used for whatever task was nearest, and judged on whether the output "looked good". AI is no different, except the output looks good far more often than it is good.
The useful question is not whether to use AI. It is which work to hand over, and which work to protect. Get that split right and AI shortens the busywork while widening your options. Get it wrong and you produce more content, more variations and more reports, all slightly off-brand and occasionally untrue, at a speed no reviewer can keep up with.
Key takeaways
- AI performance is uneven: field research with BCG consultants found big gains on tasks inside the model's capability and worse results on tasks just outside it.
- Hand AI the work that is repeatable and easy to verify; keep people close to anything that is hard to check.
- Strategy, taste, brand voice and factual accuracy stay human, because they depend on context the model does not have.
- The teams seeing results redesign the workflow around AI rather than bolting a chatbot onto the old one.
- Three guardrails do most of the work: a brand context document, named human review gates and a source check before publishing.
Why "use AI more" is the wrong instruction
The most useful piece of research on this topic is not a marketing study. In 2023, researchers from Harvard Business School, Wharton, MIT and elsewhere ran a field experiment with 758 BCG consultants. On tasks that sat inside what GPT-4 could do well, consultants using it completed 12.2% more tasks, finished 25.1% faster and produced work rated more than 40% higher in quality than a control group. On a task deliberately chosen to sit just outside the model's capability, consultants using AI were 19 percentage points less likely to reach a correct answer than those working without it.
The authors called this the "jagged technological frontier". AI is not uniformly good or uniformly bad. It is excellent at some tasks and confidently wrong at neighbouring ones, and from the outside the two look the same. That is the whole problem for marketers. A blanket instruction to "use AI more" sends people straight across the frontier without telling them where it is.
The paper describes two styles that worked: "centaurs", who split tasks cleanly between themselves and AI, and "cyborgs", who wove it in at a finer level. In both, a person decided where the line was.
Common mistake
Judging AI output on how polished it reads. Fluency is the one thing these models never lack. A confident, well-structured paragraph containing a wrong product claim is more dangerous than a clumsy one, because nobody stops to question it.
Where AI earns its place in a marketing team
When I map a team's week, the tasks that suit AI share two traits. They repeat, and a competent person can tell quickly whether the output is right. Here is how that plays out across the work most teams do.
| Task | What AI does well | What the human still owns | How easy to check |
|---|---|---|---|
| Research synthesis | Summarising reviews, forum threads, call transcripts and competitor pages into recurring themes | Choosing the sources and spotting what is missing | Medium: spot-check against the raw material |
| Briefs | Turning notes and a template into a structured first brief | The objective, the audience and the one thing the piece must achieve | Easy |
| First drafts | Outlines, meta descriptions, email skeletons, landing-page section drafts | Argument, examples, voice and every factual claim | Medium |
| Repurposing | Turning a webinar or article into posts, emails and carousel scripts | Deciding what is worth repeating and for whom | Easy: the source is right there |
| Ad variations | Generating headline and hook angles at volume for testing | Picking which angles reflect a real customer motive | Easy: the market tests it |
| Data analysis | Cleaning exports, categorising keywords, flagging anomalies | Interpreting why something moved and what to do about it | Easy if you keep the raw data |
| Reporting | Drafting the narrative around numbers you supply | The recommendation and the honesty about what did not work | Easy |
| Strategy | Listing options and stress-testing your reasoning | Almost everything, especially what not to do | Hard |
Two of these deserve more attention than they get. Ad variations are where AI pays for itself fastest, because the feedback loop is built in: you do not need to judge which hook is best, the auction does. I cover the creative side of that in why Meta ads creative is now the targeting. Research synthesis is the other. Feed a model fifty customer reviews and ask for the five objections that recur, and you get a usable starting point in minutes. The skill is in choosing what goes in. Give it only five-star reviews and it will tell you everyone loves you.
A tactical example. For a service business, I will take a month of sales-call notes, strip out personal details, and ask a model to group every objection by theme and quote the exact phrasing customers used. A strategist then reads the raw notes for the top two themes before anything goes into ad copy. The model does three hours of sorting; the person spends twenty minutes confirming it is real.
Where AI falls short, and why it will keep falling short
The weaknesses are not bugs waiting for the next model release. They come from what these systems are.
Factual accuracy
OpenAI's own research, published in September 2025, explains that language models learn to predict the next word without true-or-false labels, and that standard training and evaluation reward guessing over admitting uncertainty. Specific, low-frequency facts, such as your pricing, your service area or the date of your last product change, are exactly what pattern-matching cannot reliably produce. Google's guidance on generative AI content says the same thing in plainer terms: outputs may contain inaccuracies, so check them before publishing, and that includes titles, meta descriptions, structured data and alt text, not just body copy.
Brand voice
A model writes in the average voice of everything it has read unless you give it something better. That average is fluent, upbeat and interchangeable. Brand voice lives in the choices a brand refuses to make: the words it never uses, the claims it will not overstate, the jokes it does not tell. Those refusals have to be written down, or the model cannot know them.
Strategy and judgement
Strategy is mostly deciding what not to do, and that depends on context the model does not have: your margins, your team's capacity, the client who will churn if you change the offer, the channel that failed last year for reasons nobody wrote down. AI can produce ten options. Choosing the right one, and killing the other nine, is still your job.
Here is my first contrarian view. The task most teams hand to AI first, "write a blog post about X", is one of the worst fits. It sits right on the jagged frontier: long, fact-heavy, voice-dependent and hard to check quickly. Start with the dull work at the edges and earn trust before AI goes near anything public. If you care about how that content performs in AI search too, read why GEO is not the new SEO; the fundamentals have not changed as much as the tooling.
Use AI to widen the options and shorten the busywork. Keep the decisions human.
The checkability rule: a simple test for every task
The rule from the original version of this post still holds, and I use it in every workflow review: if a task is repeatable and the output is easy to check, automate it. If the output is hard to check, keep a person close to it.
"Easy to check" is the part teams get wrong. Ask how long it would take a competent person to confirm the output is correct, and what it costs if they miss something. Categorising 2,000 keywords into intent groups is easy to check: scan a sample, fix the rules, rerun. A comparison page making claims about competitors' pricing is hard to check and expensive to get wrong. Both can be drafted by AI. Only one can be shipped with light review.
In practice I sort tasks into three lanes. Automate: repeatable, cheap to verify, low cost of error. Assist: AI drafts, a person rewrites and approves. Protect: the person does the thinking, and AI is at most a sparring partner. Most teams put far too much in the first lane and almost nothing in the third.
What most people miss
Review time is the real cost of AI, not the subscription. If a draft takes ten minutes to generate and an hour to make accurate and on-brand, you have not saved time on that task. Track review time per task for a month and the right lane for each one becomes obvious.
How to build AI into marketing workflows
McKinsey's 2026 State of AI survey found that nearly nine in ten respondents report regular AI use in at least one business function, yet only 37% attribute any EBIT impact to it. The gap between the two is the workflow. Among the survey's high performers, nearly three-quarters say they have fundamentally redesigned workflows because of AI, against roughly a quarter of everyone else. Gartner's late-2025 research points the same way: only 5% of marketing leaders who use generative AI solely as a tool report significant gains on business outcomes.
Using AI as a tool means individuals open a chat window when they feel like it. Building it into a workflow means the process itself specifies where AI runs, what it receives and who checks it. This is the framework I use.
- Map the work before the toolsList every recurring task for one channel, such as SEO content or paid social, and note how often it happens and how long it takes. You cannot decide what to hand over until you can see it.
- Sort each task into a laneApply the checkability rule: automate, assist or protect. Be strict; when in doubt, start one lane more cautious and loosen later.
- Write the brand context onceOne document covering audience, offer, proof points you are allowed to cite, banned claims, tone and words you never use. Every prompt starts from it, so nobody improvises the brand.
- Standardise the promptsSave the prompt that works for each task in a shared library, with an example of good output. A prompt that lives in one person's chat history is not a process.
- Name a reviewer for every gateEach AI step gets a named person who approves it, not "the team". Put the gate where an error would become expensive, usually before anything is published or sent to a client.
- Check sources before publishingEvery number, quote, price and claim gets traced to a source someone has opened. If it cannot be traced, it comes out.
- Review the system monthlyLook at review time, error rates and what slipped through. Move tasks between lanes based on evidence, not enthusiasm.
This is the same thinking behind building marketing systems over campaigns: the value is in the repeatable process, not in any single output.
The three guardrails that hold under pressure
Frameworks fail when deadlines arrive. These three guardrails are the ones I insist on because they survive a busy week.
1. Brand context the model always sees
The brand context document from step three is the single highest-leverage asset in any AI workflow. My second contrarian view: which model you use matters far less than what you give it. Teams spend weeks comparing tools and minutes writing the brief. Reverse that. A mid-range model with a sharp brand document will outperform the best model working from "write in a friendly, professional tone".
2. Human review with a name attached
"Someone will check it" means nobody will. A named reviewer, a checklist for that task type and the authority to send work back are what make review real. For anything public, the reviewer should be someone who knows the business well enough to notice a claim that is fluent but false.
3. Source checks as a hard stop
Given what OpenAI and Google both say about hallucination, treat every factual statement in AI output as unverified until a person has opened the source. Google's guidance also warns that generating many pages without adding value for users can breach its spam policy on scaled content abuse. A source check slows you down just enough to stay on the right side of that line.
Pro tip
Ask the model to list every factual claim in its draft as a separate checklist, with "source needed" beside each one. It turns a vague review into a concrete task, and it often exposes claims you would have read past.
The operational reality is that guardrails feel slow for a month and invisible after that, because nobody is rewriting the same mistakes.
Where to start this week
Do not start with a tool. Start with one channel and one week of tasks. Map them, sort them into the three lanes, write a one-page brand context and pick two "automate" tasks to systemise first. Keyword grouping and report drafting are good candidates because they are easy to check and low risk. Measure review time. Expand only when the numbers say the process is working. That loop of learning, building, testing and repeating is the same one I describe in learn, build, test, repeat.
If you want the templates ready-made or an outside view on your team, both options are below, or browse all my products.
Frequently asked questions
What marketing tasks should AI handle first?
Start with tasks that repeat often and are easy to verify: grouping keywords by intent, cleaning data exports, formatting briefs from a template, drafting report narratives around numbers you supply, and generating ad headline variations for testing. These give you quick time savings with low risk, because a reviewer can confirm the output in minutes. Leave long-form public content, competitor comparisons and anything with pricing or legal claims until your brand context, prompt library and review process are in place and working.
Can AI replace a marketing strategist?
No. AI can list options, summarise research and stress-test an argument, which makes it a useful sparring partner. Strategy, though, is mostly deciding what not to do, and that depends on context a model does not have: your margins, team capacity, customer history and what has already failed. Research with BCG consultants found that people using AI on a task outside its capability were less likely to reach the correct answer than those working alone. Keep strategic decisions with a person who owns the outcome.
How do I keep AI content on brand?
Write a brand context document and make it the starting point of every prompt. Include your audience, offer, the proof points you are allowed to cite, claims you never make, tone, and specific words to avoid. Add two or three examples of on-brand writing. Then have a named reviewer who knows the brand approve anything public. Without a written context, the model defaults to a generic, upbeat voice that sounds like every other business in your category.
Does Google penalise AI-generated content?
Google's guidance focuses on quality rather than how content is produced. It does warn that using generative AI to create many pages without adding value for users may violate its spam policy on scaled content abuse. It also advises checking AI output for accuracy before publishing, including titles, meta descriptions, structured data and alt text, and considering telling readers how content was created. Useful, accurate, reviewed content is the safe path, whoever or whatever drafted it.
Why does AI make up facts?
Language models are trained to predict the next word, not to look up facts. OpenAI's 2025 research explains that specific, rarely repeated facts cannot be reliably predicted from patterns, and that common training and evaluation methods reward guessing over admitting uncertainty. For marketers, the risk is highest with exactly the details that matter most: prices, dates, product specifications, statistics and quotes. Treat every factual claim in AI output as unverified until someone has opened a source.
How do I measure whether AI is helping my team?
Track time per task including review, not just generation. A draft that takes minutes to create but an hour to correct is not a saving. Also log errors caught at review and anything that slipped through to publication. Compare these against the same tasks before AI. After a month, you will see which tasks belong in the automate lane, which need a person rewriting, and which should stay human. Then tie the time saved to outcomes such as output volume or test velocity.
Sources
- Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality (Dell'Acqua et al.) — Harvard Business School Working Paper / SSRN, 2023
- The state of AI in 2026: On the road to ROI — McKinsey & Company, 2026
- Gartner Survey Finds 65% of CMOs Say Advances in AI Will Dramatically Change Their Role in the Next Two Years — Gartner, 2025
- Why language models hallucinate — OpenAI, 2025
- Google Search's guidance on using generative AI content on your website — Google Search Central, accessed 2026

