Skip to content

Your Brand Voice Is Now a Data Asset: Voice Systems for AI-Assisted Content

Your Brand Voice Is Now a Data Asset: Voice Systems for AI-Assisted Content

The draft comes back in forty seconds. It’s clean. Grammar’s right, the facts hold up, the structure is sensible. And it reads exactly like the post your closest competitor published last month about the same thing.

That’s the complaint we hear most often from marketing leads running content with AI assistance, and it usually gets diagnosed wrong. The writer isn’t the problem, and mostly neither is the prompt wording. What’s thin is what you gave the model to work from. You asked for “friendly and professional,” which describes most of the B2B internet, and that’s what came back.

The awkward part is that most teams are already here. When the Content Marketing Institute and MarketingProfs published their fifteenth annual B2B survey in October 2024, 81% of B2B marketing teams said they were using generative AI tools, up from 72% the year before, and only 17% rated the AI-generated content as excellent or very good. Those are self-reported answers from 980 B2B marketers drawn from CMI and MarketingProfs’ own audience and fielded between June and August 2024, so read them as what practitioners say about their own work rather than as an independent quality audit. Even allowing for that, adoption is close to universal and confidence in the output isn’t, which is an odd place for a whole discipline to be sitting.

A voice system is your brand voice written down as inputs a model can use: voice attributes stated as contrast pairs, a position on each tone dimension, worked before-and-after examples, a reusable system prompt, a prompt library, a review checklist, and a named owner. A style guide describes the voice. A voice system feeds it.

Why AI-Assisted Writing Drifts Toward the Same Voice

The sameness is a documented property of these tools rather than a failure of your team. Two studies make the point.

In a controlled writing experiment, researchers Vishakh Padmakumar and He He had participants write argumentative essays three ways: unassisted, with a base language model, and with a model tuned on human feedback. Co-writing with the feedback-tuned model produced a statistically significant reduction in content diversity. Essays by different authors became more similar to each other, and the effect traced mainly to the model’s own contributed text being less varied, not to individual writers becoming repetitive. Two limits matter and belong right here: the effect was specific to the feedback-tuned model, since the base model showed no statistically significant reduction, and this was a lab study on argumentative essays rather than marketing content.

A second study, presented at the ACM Creativity and Cognition conference in 2024, found the same shape in a different task. Across 36 recruited participants, 33 of whom were analyzed, people brainstorming with ChatGPT produced ideas that were less semantically distinct from one another than people using an alternative creativity tool, even though each individual generated more ideas and more detailed ones. The authors attribute the homogenization to the model proposing similar ideas to different users, and they caution against generalizing past this small, time-boxed lab setting.

Neither study measured marketing copy, so neither is proof about your blog. Read them as mechanism rather than verdict: ask a model a loosely specified question and you get the loosely specified answer, which means the convergence shows up between teams rather than inside any one team’s work. That is exactly the symptom marketing leads describe.

None of this argues for switching the tools off, because the productivity effect is real and was measured properly. In a randomized controlled experiment published in Science, Shakked Noy and Whitney Zhang assigned 453 college-educated professionals occupation-specific writing tasks such as press releases and grant cover letters, and access to ChatGPT cut average completion time by 40% and raised human-graded quality by 18%, with the largest gains going to people whose unaided writing was rated lower. The authors note two limits: they didn’t evaluate factual accuracy, and real work usually carries company-specific context their tasks lacked, which could shrink the gains in practice.

Faster and better on average, and more like everyone else. That’s the trade you’re making, and the way out of it isn’t to stop using the tools. It’s to stop leaving the voice underspecified.

Voice, Tone, and Message Are Three Different Things

Most “our AI content sounds generic” conversations stall because three separate things get called the same thing. The cleanest formulation belongs to 18F, the US federal digital services team, whose public content guide puts it in one line: your voice is a constant, but your tone is a variable. Voice is the distilled character of the organization. Tone is the emotional register of a particular piece. That’s 18F’s own editorial house rule rather than an industry standard, but it’s the distinction that makes the rest of this tractable.

Tone is documentable, and Nielsen Norman Group gave it structure back in 2016. Their framework breaks tone into four independent dimensions: funny versus serious, formal versus casual, respectful versus irreverent, and enthusiastic versus matter of fact. Each one is a continuous spectrum with a viable neutral middle, not a three-way switch, and NN/g notes that a style sitting at the outer limit of any dimension rarely works for business purposes.

Four spectrums with a marked position on each is a specification. “Approachable but authoritative” is a mood board. A model can act on the first one.

The third thing is message, and it isn’t voice at all. Voice rules govern how you say it. A message blueprint governs what you claim, to whom, and with what proof. Hand a model your voice guide when it needed your message architecture and you get beautifully on-brand copy about the wrong thing. We drew that line in detail in our comparison of the two documents, which also covers NN/g’s companion study on how tone shifts brand perception. Worth settling before you automate anything.

One more boundary, and this one is judgment rather than evidence: a voice system can’t rescue positioning that hasn’t been decided. If three people at your company would describe what you sell three different ways, encoding a voice just makes the disagreement sound more polished. That’s a different piece of work, and it comes first.

Two part diagram. The top band shows voice as a constant and tone as a variable, attributed to the 18F Content Guide. Below it, four horizontal spectrums from the Nielsen Norman Group tone framework, each with a labeled pole at either end and a marked neutral midpoint: funny versus serious, formal versus casual, respectful versus irreverent, and enthusiastic versus matter of fact.
Voice holds still, tone moves by context. The four dimensions come from Nielsen Norman Group and the constant-versus-variable framing from 18F's content guide. Both are editorial frameworks, not research findings about your brand.

Adjectives Don’t Steer a Model. Examples Do.

This part comes out of how these systems actually work.

The foundational result is now old enough to be boring. OpenAI’s 2020 paper on few-shot learning established that large language models can perform a task from a handful of demonstrations placed directly in the prompt, with no retraining and no gradient updates, and that this few-shot performance improves substantially as models scale. The paper is honest that few-shot learning still struggles on some datasets and flags methodological concerns that come with training on large web corpora. But the mechanism held, and it’s the reason the single most useful thing in your voice guide is a set of passages rather than a set of adjectives.

Model providers give the same advice in operational terms. Anthropic’s guidance on prompt engineering for business recommends giving the model “realistic and specific examples of the inputs and ideal outputs you’re hoping to see” to handle edge cases and complex formatting or tone requirements, and describes a company that used examples specifically to move output away from another tool’s “wordiness, stilted tone, and overall lack of cohesion.” That’s a vendor’s qualitative customer story, not a controlled comparison, and no before-and-after number attaches to the tone example. Take it as a recommendation from people who build the systems, which is what it is.

Be clear about what nobody has published: there’s no head-to-head study showing that worked examples beat adjective-based instructions for steering brand voice. What exists is the mechanism plus the consistent advice of the people who build these models. That’s enough to act on and not enough to cite as proof, so we won’t pretend otherwise.

The complement to examples is the contrast pair, and Mailchimp’s public style guide is the clearest demonstration. Its voice section is written as paired opposites rather than standalone adjectives: weird but not inappropriate, smart but not snobbish, never condescending or exclusive. That’s Mailchimp’s own editorial choice rather than a research finding, but look at what the second half of each pair is doing. “Friendly” has no boundary, so a model will happily produce chummy, or fawning, or the kind of exclamation-point enthusiasm that makes a technical buyer close the tab. “Friendly but not chummy” has an edge, and an edge is something a model can respect.

The Tooling Already Assumes You Have This

The tooling moved on this while most teams were still arguing about whether to allow AI at all.

OpenAI launched custom instructions for ChatGPT in July 2023, letting people set persistent tone and style preferences that carry into future conversations without being retyped, with a 1,500-character limit per box. It launched in beta for paying subscribers and wasn’t available in the EU or UK at the time. Jasper announced Jasper Brand Voice in May 2023, a memory bank of company facts, tone, and style rules meant to keep generated content on brand at scale. HubSpot’s Content Hub relaunch included a Brand Voice feature described as maintaining consistent tone across social posts, blogs, and email. Those last two are product announcements relayed by press release and trade press, so they tell you what the vendors built and claim, not how well it works.

The field is already in your tools. Somewhere in your stack there’s a box waiting for a description of how your company sounds, and most teams have nothing good to put in it. That 1,500-character limit is a useful reminder of the shape of the answer: compressed, specific, written once.

What to Actually Write Down

Here’s the framework we use. It’s ours, built from the mechanisms above and from doing this work, and it’s worth saying plainly that no study validates this particular list. What the research supports is that examples steer models and adjectives underspecify them. The seven-part structure is our judgment about what a small team can maintain.

Start from whatever brand style guide you already have, then produce these:

  1. Voice attributes as contrast pairs. Three to five, each written as “X but not Y.” The “not” half is doing the work.
  2. A position on each tone dimension. Use NN/g’s four spectrums, mark where you sit on each, and note where tone legitimately shifts by context. An error message and a launch announcement should not share a register.
  3. Worked examples. Before-and-after rewrites and passages your team already agrees sound right. This is the highest-value artifact on the list and the one teams skip, because writing a page of adjectives feels like real work and digging up six good passages doesn’t. It’s the other way round. Adjectives are fast to write and weak to act on. The passages are the part that steers.
  4. A reusable system prompt. Items 1 through 3, compressed into something short enough to paste at the top of every session. Treat the character limit as a forcing function rather than a constraint.
  5. A prompt library and templates by content type. So nobody on your team starts from a blank box, and so two people writing the same kind of asset start from the same place.
  6. A review checklist. Named things a human is checking for, covered in the next section.
  7. A named owner. Someone whose job includes keeping this current. When Shopify rewrote its Polaris voice guidelines around the company’s first brand campaign, the work pulled in marketing strategists, support, and incident-response staff, not just the content team. That’s one company’s account of its own process rather than a general finding, but the instinct is right: the people who write in your voice under pressure know things the guide doesn’t.

The through line is that every item is an input to something rather than a description of something. That’s the work behind our Brand Voice and AI Content Systems engagements.

Bar chart of Content Marketing Institute figures published October 2024. Eighty one percent of B2B marketing teams use generative AI, up from seventy two percent the prior year. Nineteen percent describe it as integrated into daily workflows. Fifty four percent call their use ad hoc or experimental. Forty five percent have no generative AI usage guidelines, down from sixty one percent. Seventeen percent rate AI generated content as excellent or very good.
Adoption ran well ahead of governance. All five figures come from one source, the CMI and MarketingProfs B2B survey published October 2024, self-reported by 980 B2B marketers from the publishers' own audience and fielded June to August 2024. Not an independently measured dataset.

The Review Loop Is the Part Everyone Skips

Nearly half of the organizations in that same CMI survey, 45%, reported having no generative AI usage guidelines at all, though that was down from 61% the year before. Same self-reported survey, same 980 B2B respondents, same caveat about it being practitioners describing their own organizations. CMI content strategist Erika Heald, commenting in that report, tied the gap directly to voice: very few companies have comprehensive content governance in place, she said, “starting with defining their unique brand voice.” That’s a named strategist’s interpretation in the report, not a number the survey measured.

Three published models are worth borrowing, and none of them require a compliance department.

The Associated Press put its guidance in the Stylebook, and as Poynter’s Alex Mahadevan reported in August 2023, the framing is the useful part: AI output is unvetted source material that still has to clear the same editorial judgment and sourcing standards as anything else. That’s AP describing its own policy rather than an outside audit of whether AP follows it, but as a mental model for a marketing team it’s close to perfect. You wouldn’t publish a paragraph a stranger emailed you without checking it. Same rule.

WIRED went further and published its policy openly. As Nieman Lab reported in March 2023, editor-in-chief Gideon Lichfield said the outlet does not publish stories with AI-generated text, not whole stories and not snippets, extending to editorial newsletters, while marketing emails may use AI with disclosure. By his account the draft went through several revisions, was discussed with the whole newsroom, and was run past leadership, and he expected the stance to change as the tools did. It’s one publication’s policy, not a standard. Copy the habit rather than the rules: write it down, argue about it with the people it governs, expect to revise it.

If you want a scaffold instead of an example, NIST’s Generative AI Profile for its AI Risk Management Framework (document NIST AI 600-1, published July 2024) is organized around four functions: govern, map, measure, manage. It’s voluntary guidance aimed at organizations larger than yours, but those four verbs are a decent skeleton for a two-page internal policy.

And then there’s the case that explains why “we review everything” isn’t a review process. CNET began quietly publishing AI-assisted finance explainers in November 2022, and by January 2023 had run roughly 75 of them. After Futurism reported on the practice, a CNET spokesperson said the company was reviewing all its AI-assisted pieces to make sure no further inaccuracies got through, noting that humans make mistakes too. That’s a defensive statement to press rather than an independent audit finding. But one of the published pieces had stated that a $10,000 deposit at 3% annual interest would earn $10,300 in interest in the first year. The right number is $300.

That error survives a proofread. It does not survive a checklist with “recompute every number in the piece” on it. That’s the entire distinction between a review step that exists and a review step that works, and it’s why item six on the list above is a document rather than an intention.

How You’d Know If It’s Working

Nobody has good instrumentation for this yet, and you should be suspicious of anyone selling you a brand-voice score.

The measurement gap is broader than voice. In a vendor-run survey of 503 marketing leaders reported by Jasper in March 2025, 51% said they can’t measure the ROI of their AI investments, with another 22% planning to start tracking it. Jasper sells an AI content platform and drew its sample from its own database plus a research firm, so it isn’t a representative sample of marketers. The same report names data privacy and AI output quality as the biggest barriers to adoption, and ties the output-quality concern to pressure on brand integrity, accuracy, and consistency at scale.

Three things we’d actually do, all of them reasoning rather than anything a study supports:

Run a periodic sample audit. Pull five recent pieces at random and score them against your voice attributes. Not a rubric, just a read with the contrast pairs in front of you.

Do a blind read. Strip the logos off three of your posts and three from a competitor, hand them to someone who knows the category, and ask which are yours. If they can’t tell, you’ve learned something specific and cheap.

Watch how much editors are changing. Not as a target, since there’s no published guidance on how much a human should rewrite an AI draft and any percentage you’ve seen quoted is almost certainly untraceable. As a trend line, though, it’s informative. If the same editor is rewriting less over time on the same content type, the inputs are getting better.

Treat It as Source, Not as a Deliverable

The reason a voice system beats a voice PDF has nothing to do with format. It’s that a PDF has no update path.

Look at how the guides that survived were built. When GOV.UK launched its editorial style guide in 2012, the team built it from user research rather than arbitrary preference, mandated plain English across the site, and required acronyms to be spelled out on every page because people arrive from search with no prior context. They shipped it as an alpha and said outright it would keep changing. Mailchimp put its content style guide in a public GitHub repository as plain Markdown under a CC BY-NC 4.0 license, created in July 2015, so its history is a commit log rather than a version number in a filename. Shopify rewrote its guidelines when the company’s positioning moved.

Notice what all three have in common. They’re maintained, they’re versioned, and somebody owns them. That was good practice when the audience was human writers who could interpolate around a stale rule. It’s closer to mandatory when the consumer is a model that will apply a stale rule literally, at volume, every day, without ever mentioning that it looked wrong.

Your voice system will go out of date. That’s the tradeoff for building something a model actually consumes, and it’s why we’d rather put content on a system than file it as a document. A document that never needs updating is usually one nobody’s using. Plan for the edit rather than the launch.

Where to Start

Don’t commission anything yet. Four steps, all of which you can do with people you already have:

Pick three contrast pairs and argue about the second half of each until you agree. Collect six passages your team already thinks sound right, from anywhere, including a founder’s email. Compress those into a system prompt short enough to paste. Write the checklist of what a human verifies before anything ships, and put “recompute every number” on it.

That gets you most of the value. The rest is maintenance and the discipline to keep examples current as the company’s story changes.

If you’d rather not build it in-house, that’s the work we do: extracting the voice from what you’ve already written, building the prompt library and templates, and setting up the review checklist so the speed you gain doesn’t cost you the thing that makes you sound like yourselves.

Questions Marketing Leads Keep Asking

What’s the difference between brand voice and brand tone? Voice is constant, tone changes with context. 18F’s content guide puts it as your voice is a constant, but your tone is a variable, which is that team’s editorial house rule rather than an industry standard. Nielsen Norman Group makes tone documentable with four independent spectrums: funny versus serious, formal versus casual, respectful versus irreverent, and enthusiastic versus matter of fact. Document the voice once, then note where on each spectrum a given content type should sit.

Can I just tell an AI tool to write in a friendly, professional voice? You can, and you’ll get the average of everything on the internet written in a friendly, professional voice. Adjectives underspecify. What models are documented to use well is demonstrations: OpenAI’s 2020 research established that these systems can pick up a task from a handful of examples placed in the prompt, with no retraining, though the paper notes few-shot learning still struggles on some tasks. Anthropic’s business guidance likewise recommends giving the model realistic and specific examples of inputs and ideal outputs, which is a vendor recommendation rather than a controlled comparison. No published head-to-head study compares examples against adjectives for brand voice, so treat this as the mechanism the model builders point to, not as proof.

How much should a human editor change an AI first draft before we publish it? There’s no defensible number, and any percentage you’ve seen quoted is worth checking before you repeat it. Use a principle instead. The Associated Press treats AI output as unvetted source material that must clear normal editorial standards, as Poynter’s Alex Mahadevan reported in August 2023, describing AP’s own internal policy rather than an outside evaluation. The practical version: verify every fact and recompute every number, then edit for voice until it sounds like you. That’s a checklist, not a percentage.

We’re a five-person team with no brand manager. Is this worth doing? Probably more than it is for a large team, because you have less slack to absorb rework. In CMI and MarketingProfs’ B2B survey published October 2024, 45% of organizations reported no generative AI usage guidelines at all, down from 61% the prior year, based on self-reported answers from 980 B2B marketers in the publishers’ own audience. The minimum viable version is one page: three contrast pairs, six examples, one system prompt, one checklist.

Is a one-time brand voice document enough? No, and the organizations that got this right treated it as maintained source from the start. GOV.UK shipped its style guide as an alpha in 2012 and said outright it would keep evolving. Mailchimp keeps its content style guide in a public GitHub repository as versioned Markdown under a CC BY-NC 4.0 license. A model applies a stale instruction literally and at volume without flagging that it looks wrong, which makes the update path more important than it used to be, not less.

Does AI-assisted content hurt our search visibility? Not by virtue of being AI-assisted. Google’s Search Central blog stated in February 2023 that its “focus on the quality of content, rather than how content is produced” is what guides its ranking systems, and that automation becomes a spam policy violation when it’s used primarily to manipulate rankings. That’s Google’s own characterization of its systems rather than an independently verified claim, and the line it draws is about intent and quality. The practical risk isn’t a penalty. It’s that undifferentiated content has nothing worth citing in it, which matters for answer engine optimization as much as for rankings.

Voice or messaging first? Messaging, if what’s inconsistent is the claims rather than the register. A voice system makes your copy sound like you; a message blueprint decides what the copy says. If two of your pages describe the product differently, no amount of voice documentation fixes that, and the distinction is worth settling properly before you automate either one.

Ready to put these insights into action?

Let's discuss how Triaza can help your business grow.

Talk with us

Resources

Go To The Blog