Between the chat window and the publish button, most lean marketing teams have exactly one step: copy, paste, skim. There’s no source check, no claims review, and no point where a named person says this is fine to put our name on. That isn’t a discipline problem. Nobody ever drew the workflow, so there’s nothing to skip.
The bill arrives later, as a post that reads fine but says nothing only your company could have said, a statistic in paragraph four that nobody traced to a source, and a blog that has quietly converged on the same register as everyone else in the category. Teams reach for prompt fixes when this happens, because the prompt is the part they can see. The work that changes the output happens around it.
The numbers say most teams are stuck in exactly that spot. In a survey of 980 B2B marketers fielded in mid-2024 and published in October 2024 by the Content Marketing Institute, MarketingProfs and The MX Group, 81% said their teams were using generative AI tools, up from 72% the year before, but only 19% said it was integrated into daily workflows, 54% described their use as ad hoc or experimental, and 45% had no formal AI usage guidelines at all. That’s self-reported data from marketers who volunteered to answer, so treat it as a picture of how teams describe themselves rather than a measurement of what they do. Even discounted, the shape is hard to miss: adoption ran well ahead of process.
An AI content system is a pipeline with named handoffs: research and sourcing stay human, drafting and reformatting can be AI-assisted, and every draft clears a fixed QA gate before it publishes. Three artifacts hold it together: a prompt library, a voice guardrail file, and a checklist one person signs.
The Five Stages, and Who Owns Each One
Before deciding where AI goes, name the stages. Most small teams have never written the pipeline down, which is why the handoffs are invisible and therefore skippable.
- Brief and angle. What this piece argues, who it’s for, what it has to do. Human.
- Research and sourcing. Finding the evidence, opening the pages, recording the dates. Human.
- Outline. Turning the brief and the research into a structure. Mixed.
- Draft. First-pass prose from the outline and the sources you supplied. AI-assisted.
- Edit and QA. Fact checks, claims review, voice, formatting, approval. Human.
Then a publish gate that nothing gets through without a signature.
This split is our framework, not a studied standard. It’s built on what the evidence below supports and on what professional publishers already do, and the rest of this article is the reasoning behind it. Two things are worth saying now. Stage two is where hallucinated citations get in, so it can’t be the model’s job. Stage five is where the piece either becomes yours or stays generic, so it can’t be skipped when the week gets short.
If you’re building this from nothing, our content marketing work starts at the strategy layer that sits above the pipeline.
Where AI Belongs, and Where a Human Has to Stay
Newsrooms have been running this experiment in public for a couple of years, with more legal exposure than you have, and their answer is consistent.
In a Reuters Institute survey of 314 news executives across 56 countries, fielded in November and December 2023 and published in January 2024, back-end automation such as tagging, transcription and copyediting was rated “very important” by 56%, against 28% for content creation with human oversight. Asked separately which area carried the greatest reputational risk from AI, 56% named content creation, against 11% for back-end automation. That’s a strategic sample of digital leaders reporting their own perceptions of importance and risk rather than measured outcomes, so it tells you where experienced publishers are placing their bets, not whether the bets are right. Read together, the two answers describe the same instinct: automate the plumbing, keep hands on the prose.
None of that means AI is useless on the writing itself. The best measurement we have says the opposite. In a preregistered experiment published in Science in July 2023, 453 college-educated professionals were given occupation-specific writing tasks, and the half given access to ChatGPT finished roughly 40% faster and scored roughly 18% higher on output quality. The detail that matters most for pipeline design is buried in the same paper: the tool restructured the work towards idea generation and editing, and away from rough drafting. This was an online experiment on mid-level tasks like press releases and short reports, with participants in specific occupations, so it isn’t a marketing team in the wild. But the direction is the useful part. The gain shows up when a human’s time moves from producing first drafts to shaping and checking them.
That’s roughly where the Associated Press landed. Its August 2023 staff guidance, reported by Poynter, tells journalists to treat any AI output as unvetted source material, apply normal sourcing standards before publication, and never use generative AI to create publishable content directly. AP’s own framing is that AI is “a tool we can use, but does not replace the journalistic smarts, experience, expertise” the work needs, and the stated concerns were hallucination, the ease of producing disinformation, and privacy in whatever staff type into a chat window. “Unvetted source material” is a useful phrase to steal. It puts the output in the same bucket as a tip from a stranger.
The same Reuters Institute report describes what that looks like when it’s actually built. Le Monde used AI to help translate around 30 stories a day into its English edition, but paired it with several human checks and customized the software to follow the paper’s own style book. Aftonbladet added AI-assisted bullet-point summaries to articles, with an on-article note saying the summary was made with AI support and quality assured by Aftonbladet staff. Both of those are publisher-reported descriptions from late 2023 and early 2024 rather than independently audited workflows, and the human-check steps are the publishers’ own account. Still, the pattern is the one worth copying: a narrow, well-defined task for the machine, a named human step, and a disclosure where a reader would want one.
Build a Prompt Library, Not Better Prompts
Rewriting a prompt every time you write a post is how teams end up with output that varies by mood. The fix is boring: a small set of reusable prompts you keep, edit, and reuse, with the things that shouldn’t change written into them.
Most of the Library Is Examples
The reason to keep passages rather than instructions is mechanical. OpenAI’s 2020 paper on few-shot learning showed that a large language model can take on a new task purely by conditioning on a handful of examples placed in the prompt, with no fine-tuning and no retraining, though the same paper is candid that few-shot performance still struggles on some datasets. Anthropic gives the same advice for its own product, recommending “realistic and specific examples” of inputs and ideal outputs including edge cases, in a February 2024 post that also says plainly there is “no single best technique for prompt engineering.”
Both of those describe how the tool responds to input, not how good the output will be. The operational takeaway is narrow and worth taking literally: the examples file earns its place in the library, and adjectives mostly don’t.
Personas Are the First Lever Everyone Reaches For, and the Weakest
“You are a world-class B2B copywriter” is the most common opening line in marketing prompts, and there’s reason to doubt it does much. A November 2023 study tested 162 personas, spanning six kinds of interpersonal relationship and eight expertise domains, across four families of models on 2,410 factual questions. Adding a persona to the system prompt did not improve performance over using no persona at all, and the effect of any given persona was largely unpredictable. The authors note that if you could pick the best persona for each individual question you would see a real gain, but predicting which one that is performs no better than choosing at random.
Read that carefully before you throw out your system prompt. The study measured accuracy on factual questions, not voice, register, or style. It does not show that a persona has no effect on how writing sounds. What it does undercut is the assumption that a job title at the top of a prompt is doing serious work. If the persona line is the only steering in your prompt, you have decoration where you need instructions.
What Actually Goes in the Library
Here we’re out of research and into recommendation, and we’d rather say so than dress it up. There’s no study we could point to on how a marketing team should structure a prompt library, because as far as we can tell nobody has run one. This is what we’d put in it:
- A task prompt per content type. Blog post, case study, product page, newsletter. One each, not one universal prompt with switches.
- An examples file. Two to four passages of your own published writing that you’d be happy to see imitated. Include a before and after if you have one.
- A constraints block. What you never claim, what you never invent, what you never say about a competitor, and your formatting rules.
- A sourcing rule. Something close to: use only the sources pasted below, and if a claim needs a source that isn’t here, flag it instead of finding one.
- The voice guardrails. Covered next.
- A plain-text home and a change log. A doc, a repo, a wiki page. Somewhere two people can both see it, with a note when it changes and why.
The change log matters more than it sounds. When output quality shifts, you want to know whether the prompt moved. Treating prompts like any other shared asset your team versions is the whole idea, and it costs nothing.
That structure is close to what we build for clients in our brand voice and AI content systems work, and none of it requires a tool you don’t have.
Voice Guardrails Are One Input, Not the Whole System
Two studies point at the symptom that probably brought you here. In a July 2024 experiment published in Science Advances, 293 participants wrote short stories, some with AI-generated story ideas available and some without, and 600 evaluators produced 3,519 ratings. Stories written with AI help were rated more creative individually, but collectively they were more similar to each other. The authors are explicit that the task was constrained in length, medium and output type, and that results “may not generalize to other less-constrained creativity tasks.” A separate 36-participant study from February 2024 found a similar pattern in idea generation: people using ChatGPT as a creativity aid produced ideas that were less distinct from each other than people using a non-AI alternative. That’s one small lab study against one specific comparison tool, framed by its authors as testing a hypothesis.
Neither study says AI makes your writing worse. Both suggest something more specific and more relevant: unguided AI assistance makes different people’s output converge. If everyone in your market is prompting the same tools with the same kind of instruction, sounding like everyone else is the default outcome, not bad luck.
Guardrails are the counterweight, and they’re a real piece of work with their own method. What goes in the guardrail file, how to derive it, and whether you need a style guide or a message blueprint first are the subject of their own article. For pipeline purposes, treat the guardrails as one input to the prompt library alongside constraints and sourcing rules, keep them in the same place, and load them into every session rather than remembering them selectively.
The Editing Layer, and What It’s Actually Checking For
The failure modes are documented well enough that you can design the check around them instead of guessing.
Arithmetic that looks fine. In an AI-written CNET explainer on compound interest, the copy stated that depositing $10,000 at 3% annual interest would yield “$10,300 at the end of the first year,” which confuses the account balance with the interest earned. The earnings were $300. Futurism reported it in January 2023 as one of three errors in that single article, with a possible fourth also flagged, and CNET corrected all three after being contacted. Nothing about that sentence looks wrong at a skim, which is exactly the danger.
Fabricated people. Sports Illustrated published articles under author bylines that did not correspond to real people, illustrated with profile photos sold on a site that sells AI-generated headshots. CNN reported in November 2023 that the articles were removed and that publisher Arena Group attributed the content to a third-party contractor. Arena Group disputed the framing, saying the contractor had writers use pen names to protect author privacy. Worth being precise about what’s established here: the fabricated bylines and photos are the confirmed part, not that the article text itself was machine-written.
Citations that don’t exist. This is the one that should change how you work. A Tow Center study published by the Columbia Journalism Review in March 2025 gave eight AI search tools 1,600 queries, each one an excerpt from a real article, and asked them to name the article, the publisher, the date and the URL. The tools answered incorrectly more than 60% of the time. The most accurate tool tested was still wrong 37% of the time, and the least accurate was wrong 94% of the time. The researchers note their design “may not reflect typical user behavior” and that each query was run only once against systems that answer dynamically, so treat it as one careful snapshot rather than a fixed rate. Even at its most generous, it means the same thing for your pipeline: if you ask a chat tool where a quote came from, it is wrong more often than it is right. Sources go in from a human. They do not come out of the model.
Volume without review. The scale of unreviewed output is measurable in places you’d expect to be well defended. An October 2024 analysis using detectors calibrated to a 1% false-positive rate found that over 5% of newly created English Wikipedia articles were flagged as AI-generated, and that flagged articles tended to be lower quality and more often self-promotional. The authors are careful to call this a lower bound rather than a measurement, given the false-positive rate baked into their threshold. The developer Simon Willison proposed a name for this category in May 2024: slop, meaning AI content that is “mindlessly generated and thrust upon someone who didn’t ask for it.” His own qualifier is the useful part. Not all AI-generated content is slop. The distinguishing feature is that nobody reviewed it.
What the Rules Actually Say, and What They Don’t
Most teams arrive at this topic braced for a penalty that doesn’t exist, and miss the exposure that does.
On search, Google’s position has been public and stable since February 2023. Its guidance on AI-generated content states: “Our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years.” There’s no ban on AI-assisted writing. The condition attached is about intent, and Google restated it plainly in March 2024: “Our long-standing spam policy has been that use of automation, including generative AI, is spam if the primary purpose is manipulating ranking in Search results.”
That March 2024 update introduced a policy worth reading closely if you’re planning to scale output. Scaled content abuse, in Google’s words, “is when many pages are generated for the primary purpose of manipulating Search rankings and not helping users,” and it applies “no matter whether content is produced through automation, human efforts, or some combination of human and automated processes.” Production method is not the test. Volume of low-value pages is. A team publishing four genuinely useful posts a month with AI help is not the target. A team publishing four hundred thin ones is, whoever or whatever typed them.
The exposure people underestimate is the advertising kind. In September 2024 the FTC announced Operation AI Comply, five law enforcement actions over AI-related deceptive claims. One involved an AI writing tool whose subscribers could mass-generate consumer reviews “potentially containing false information,” with some generating “hundreds, and in some cases tens of thousands” of them. None of that is an adjudicated finding of wrongdoing. Three of the five are lawsuits still to be decided by a federal court. The action against the writing tool was a settlement, finalized in December 2024 on a 3 to 2 Commission vote with two commissioners dissenting. Separately, the FTC’s revised Endorsement Guides, effective July 2023, broadened “endorser” to cover “what appear[s] to be an individual, group, or institution,” language the Commission said was meant to reach “the writers of fake reviews and non-existent entities that purport to give endorsements.” Two honest caveats: the Guides describe themselves as advisory administrative interpretations, and they never use the word “AI” anywhere. The application to AI-generated endorsements is an inference from wording that never names the technology.
Disclosure is a judgment call rather than a requirement. Google’s own guidance frames it situationally: “AI or automation disclosures are useful for content where someone might think ‘How was this created?’ Consider adding these when it would be reasonably expected,” while noting that giving AI an author byline “is probably not the best way” to be clear with readers. Reader attitudes point the same direction. The Reuters Institute’s Digital News Report 2024, a 28-market survey published in June 2024, found 19% of respondents comfortable with news made mostly by AI with human oversight against 36% comfortable with human-made news that had AI assistance. The report calls these early-days attitudes that are still forming, so don’t build policy on the gap alone. What it does suggest is that readers care less about whether a tool was involved than about whether a person was.
You Can’t Outsource the Check to a Detector
The obvious shortcut is to run the draft through an AI detector and ship it if it passes. That shortcut does not work, and it has a specific failure mode worth knowing.
A 2023 study by Liang and colleagues ran seven widely used AI-text detectors against 91 real TOEFL essays written by non-native English speakers. The detectors misclassified them as AI-generated an average of 61.22% of the time, and 19.78% of those essays were flagged by all seven. The same detectors classified 88 essays by native English speakers accurately. The authors attribute the gap to the detectors’ reliance on text perplexity: writing with a narrower vocabulary and grammar range looks statistically similar to machine output. That’s a narrow sample of two essay sets, not a general audit of every detector on the market, but the mechanism is the concerning part, because it means the tool is measuring the wrong thing.
The Markup reported in August 2023 on international students being falsely accused of cheating on that basis, and noted that OpenAI had shut down its own AI-text classifier at the end of July 2023 because of low accuracy. Turnitin’s chief product officer disputed the bias finding for Turnitin specifically, while acknowledging the tool “ends up learning that more complex writing is more likely to be human.” The article leaves that tension unresolved, and so should you.
The deeper problem is that a detector answers a question you don’t actually need answered. It cannot tell you whether a claim is true, whether a cited source exists, or whether the piece sounds like your company. Those are the three things that decide whether the post should go out, and all three are checklist items.
The Human-in-the-Loop QA Checklist
This is the artifact. It’s our framework rather than a standard, and we have no evidence that running it improves any metric, because as far as we can find nobody has measured that. What we can say is that every gate below maps to a documented failure mode from earlier in this article.
One rule before you start. The person who ran the prompt is not the person who approves the post. If your team is one person, put a day between drafting and approving, then run the checklist cold. This is the only part of the system that can’t be automated.
Gate 1: Sourcing. Roughly 10 minutes.
- Every source in the draft was supplied by a human. If the model named a source, that’s a lead to verify, not a citation.
- Every link opens and lands on the page you expect, not a homepage redirect.
- Every source’s publish date is visible on the page itself, not inferred from a search result.
- No claim that depends on timing rests on an undated page that can be quietly rewritten.
- Nothing is attributed to “studies show,” “research indicates,” or “experts say” without a named study.
Gate 2: Factual accuracy. Roughly 15 minutes.
- Every number appears on the page it’s cited to. Not a rounded version, not a version from a summary. Open the page.
- Every number has the right unit, base and direction. A percentage of what, out of what.
- Every proper noun is spelled correctly and belongs to a real person, company or product. Check bylines specifically.
- Any arithmetic the model performed has been redone by a human.
- Anything you can’t verify is deleted, not softened.
Gate 3: Claims. Roughly 10 minutes.
- No performance, revenue, client, pricing, capacity or guarantee claim about your own company that hasn’t been approved.
- No claim about a competitor you couldn’t defend in writing.
- No testimonial, review, quote or endorsement that wasn’t given by a real named person.
- Every third-party statistic carries the source’s own qualifier in the same paragraph: sample size, lab conditions, self-reported, vendor-run, projection rather than measurement.
- No legal, financial, medical or compliance claim has been introduced without a human who knows the subject.
Gate 4: Brand voice. Roughly 10 minutes.
- Read the first paragraph out loud. If it could open a competitor’s post on the same topic, rewrite it.
- The piece contains at least one thing only your company would say: a real example, a constraint you’ve actually hit, an opinion.
- Your banned constructions are gone. Keep a list. Ours starts with announcer openings, manufactured urgency, and any sentence whose only job is to introduce the next one.
- Nothing contradicts what you’ve published elsewhere.
Gate 5: Formatting and structure. Roughly 5 minutes.
- Headings say what the section argues, not what it’s about.
- Every list has a reason to be a list.
- Links have descriptive anchor text and point at pages that exist.
- Images have alt text that describes the image.
- Title and meta description are inside your character limits and were written by a person.
Gate 6: Approval. Roughly 2 minutes.
- One named person has read the whole thing start to finish, in one sitting.
- That person’s name and the date are recorded somewhere.
- If a paragraph would embarrass you quoted out of context, it doesn’t go out.
- Disclosure decision made and recorded: would a reader of this specific piece reasonably want to know how it was made?
That’s roughly 50 minutes a post by our estimate, and the estimate is the point. A checklist nobody has costed gets abandoned in week three. Fifty minutes against a draft that took nine seconds is a trade most teams will take once they’ve seen what gate 2 catches.
Common Questions About AI Content Workflows
Does Google penalize content just because it was written with AI?
No. Google’s February 2023 guidance states that “our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years.” The condition is intent: Google restated in March 2024 that “use of automation, including generative AI, is spam if the primary purpose is manipulating ranking in Search results.”
If we increase output tenfold with AI, will we get hit by the spam policies?
Not for the volume by itself. Google’s scaled content abuse policy, introduced in March 2024, targets cases where “many pages are generated for the primary purpose of manipulating Search rankings and not helping users,” and Google states it applies “no matter whether content is produced through automation, human efforts, or some combination of human and automated processes.” Low value at scale is the risk, and a human author is no defense against it. Ten times the output with the same standard is a different thing from ten times the output with no standard.
Can we just run AI drafts through a detector before publishing?
No, and it can do real harm. A 2023 study found seven widely used detectors misclassified 91 real TOEFL essays by non-native English speakers as AI-generated an average of 61.22% of the time, while classifying 88 native-speaker essays accurately; the authors attribute this to detectors relying on text perplexity, and the sample is narrow. The Markup reported in August 2023 that OpenAI shut down its own classifier that July for low accuracy. A detector also can’t tell you whether a claim is true or a source exists, which is what actually matters.
We don’t have a full-time writer. Who’s supposed to catch factual errors?
Whoever runs your pre-publish QA checklist, and it doesn’t have to be a writer. The errors that get published are usually verifiable in a browser tab: an AI-written CNET explainer said $10,000 at 3% would yield “$10,300 at the end of the first year” when the earnings were $300, one of three errors Futurism found in that piece and CNET corrected. The Associated Press’s August 2023 guidance frames the posture well: treat AI output as unvetted source material and apply normal sourcing standards. Checking is a role anyone on the team can own.
Do we have to disclose that a post was AI-assisted?
There’s no general legal or search requirement, and nobody should tell you otherwise. Google’s guidance calls disclosures “useful for content where someone might think ‘How was this created?’” and suggests adding them “when it would be reasonably expected,” while noting an AI author byline “is probably not the best way” to be clear with readers. The Reuters Institute’s Digital News Report 2024 found 19% of respondents comfortable with news made mostly by AI with human oversight against 36% for human-made news with AI assistance, and calls these early-days attitudes still forming. Inventing an author is a separate and much clearer problem.
How do we tell whether our content sounds like us or like everyone else?
Read your opening paragraph out loud and ask whether a competitor could have published it. The convergence is real and measured: a July 2024 Science Advances experiment with 293 writers found stories written with AI-generated ideas were rated more creative individually but were collectively more similar to each other, though the authors note the task was constrained in length and format and may not generalize. The fix is documented voice loaded into every session, which is its own piece of work.
Will more quotes and statistics get us cited in AI answers?
Adding substance a model can quote is reasonable practice, and it’s also just good writing, but we’d rather not hand you a number here. Getting cited in AI answers is a different discipline from producing good content, with its own measurement problem, and the first useful step is finding out whether you’re being cited at all before changing anything. Our AI search visibility audit walks through that.
Start With Three Things This Week
Write down the five stages and mark who owns each one at your company. Not who should, who does. That single page usually shows you the gap without any further diagnosis.
Then save two of your own best published paragraphs into a file and paste them into your next prompt. Examples are the cheapest steering you can give the tool, and the input format it was built to learn from.
Then print gate 1 and gate 2 and run them on your next post, before it goes out. If they catch nothing, you’ve lost 25 minutes and gained some confidence. If they catch something, you’ll know why the rest of the checklist exists.
If you’d rather not build the pipeline from scratch, our brand voice and AI content systems work covers voice extraction, the prompt library, the workflow, and the editorial QA checklist as one build.



