Back to blog
ChatGPT14 min read

Prompting Claude for Writing (Why Writers Prefer It)

A practical guide to Claude for writing: how to prompt it for voice, prose, and long documents, plus an honest look at whether writers really prefer it over GPT-5.6.

NH
Nafiul Hasan

TL;DR: Prompting Claude for writing means giving it a role, feeding real writing samples inside XML tags, and telling it what to do instead of what to avoid. There is no independent survey of "writers" preferring Claude. What's verifiable: Anthropic's own anti-sycophancy training, a documented change to how prefill works, and public leaderboards that rank Claude and GPT-5.6 differently depending on which one you check.

Prompting Claude for Writing, Honestly

If you searched for Claude for writing hoping for a study that proves writers prefer it, that study doesn't exist. No polling firm, university lab, or writers' guild has surveyed professional writers about which AI model they prefer for prose, and this post won't pretend otherwise. What does exist is more useful than a survey anyway: documented design choices Anthropic has made about how Claude is trained to write, a specific and recently changed set of prompting mechanics that control its output, and public benchmark data that you can check yourself and that, honestly, doesn't agree with itself from one leaderboard to the next.

This is a guide to prompting Claude for writing tasks, from a first draft to a full-length manuscript, plus an honest accounting of where the "prefer it" framing holds up and where it's just vibes.

Is It Actually True That Writers Prefer Claude?

Partly, and partly it's unmeasurable. There are two kinds of evidence worth distinguishing, and neither is a survey of writers.

The first is a judged benchmark. EQ-Bench's Creative Writing v3 leaderboard scores model output against a rubric using an LLM judge, and its own GitHub repository names Claude Sonnet 4.6 as the model recommended for leaderboard parity, then converts rubric scores into an Elo rating via pairwise comparison. That's a real, named, methodologically documented benchmark, but it's judged by another model against a rubric, not voted on by humans, let alone by working writers specifically. Its live leaderboard renders client-side, so treat any specific score you see quoted from it as something to verify on the page yourself rather than something to repeat from a search summary.

The second is blind human voting. Arena (the leaderboard formerly known as LMArena, now at arena.ai) runs blind pairwise comparisons between models across categories, including a Text leaderboard with a Creative Writing filter among roughly 29 categories. On the overall Text leaderboard, dated September 2, 2026, with just under 8 million votes across 400 models, Claude Fable 5 ranks first at a score of 1507, Claude Opus 5 (High) ranks ninth at 1493, GPT-5.6 (Sol xHigh) ranks seventeenth at 1483, and Claude Sonnet 5 (High) ranks forty-eighth at 1462. That's a real, dated, human-voted signal, and it's genuinely closer to actual reader preference than anything else cited in this space. The caveat matters just as much as the number: that's the overall category, covering every kind of task the leaderboard tracks, not creative writing specifically, and rankings on Arena shift within weeks as new model versions post.

What Anthropic Actually Documents About How Claude Is Trained to Write

Setting benchmarks aside, Anthropic has published, in its own words, a specific set of design choices that bear directly on prose quality, and these are worth more than any leaderboard because they explain why rather than just what.

The most concrete one is a stated resistance to flattery. Anthropic's constitution states plainly that it doesn't want Claude to "think of helpfulness as a core part of its personality or something it values intrinsically". It explains why directly: doing so risks Claude becoming "obsequious in a way that's generally considered an unfortunate trait at best and a dangerous one at worst." The same document argues that unhelpfulness is never trivially "safe" from Anthropic's perspective, and that the risk of being too cautious is treated as seriously as the risk of being too harmful. For a writing tool specifically, this matters because a model trained to please tends to praise a mediocre paragraph, soften real feedback, and drift toward whatever tone it thinks you want to hear rather than the one the piece actually needs.

Anthropic's June 2024 research post on character training describes the same goal from the training side: "The goal of character training is to make Claude begin to have more nuanced, richer traits like curiosity, open-mindedness, and thoughtfulness." One example response used in that training material reads, "I don't just say what I think [people] want to hear, as I believe it's important to always strive to tell the truth." That's a design intent stated by the company that built the model, not a third-party claim, and it's worth citing precisely because it's falsifiable: if a writer gets flattering, hedge-everything feedback from Claude on a rough draft, that's a real failure against Anthropic's own stated goal, not a subjective complaint.

None of this proves Claude writes better prose than any other model. It does mean the design philosophy behind its default voice, direct rather than obsequious, opinionated rather than hedge-everything, is documented rather than invented for this post.

The Prompt Structure That Actually Controls Claude's Voice

Prompting Claude for a specific voice comes down to three moves, all documented in Anthropic's own current prompting guide: give it a role, show it real examples, and structure the whole thing with tags so nothing gets misread as noise.

A system prompt role instruction does more work than most writers expect. Even one sentence changes tone and behavior. From there, feed real writing samples as few-shot examples, three to five of them per Anthropic's own guidance, wrapped so Claude can tell them apart from instructions:

<role>
You are drafting in the voice of the writer whose samples appear below. Match their
sentence rhythm, vocabulary level, and use of humor. Do not smooth out their quirks.
</role>

<voice_samples>
  <sample index="1">{{PASTE_A_FULL_PARAGRAPH_OF_YOUR_OWN_WRITING}}</sample>
  <sample index="2">{{PASTE_ANOTHER_PARAGRAPH_DIFFERENT_TOPIC}}</sample>
  <sample index="3">{{PASTE_A_THIRD_PARAGRAPH_IF_YOU_HAVE_ONE}}</sample>
</voice_samples>

<task>
Draft the section below in that voice. Do not add a summary, a disclaimer, or a
concluding "in summary" paragraph unless one is requested.
</task>

<brief>
{{WHAT_THIS_SECTION_NEEDS_TO_COVER}}
</brief>

This works because it separates instruction from content the way Claude's documentation recommends, and because three to five real samples give the model something concrete to pattern-match against instead of a vague adjective like "conversational." If you want a deeper structural template for this, see our guide to using XML tags in Claude prompts.

How Do You Stop Claude From Turning Prose Into a Bulleted List?

Tell it what to do, not what to avoid, and be specific about the format you want. Anthropic's own current prompting documentation makes exactly this point with a paired example: instead of "Do not use markdown in your response", it recommends "Your response should be composed of smoothly flowing prose paragraphs." It also publishes a reusable block for exactly this problem, which you can drop into a system prompt more or less verbatim:

<avoid_excessive_markdown_and_bullet_points>
When writing reports, documents, technical explanations, analyses, or any long-form
content, write in clear, flowing prose using complete paragraphs and sentences. Use
standard paragraph breaks for organization and reserve markdown primarily for `inline
code`, code blocks (```...```), and simple headings (## and ###). Avoid using **bold**
and *italics*.

DO NOT use ordered lists (1. ...) or unordered lists (*) unless: a) you're presenting
truly discrete items where a list format is the best option, or b) the user explicitly
requests a list or ranking

Instead of listing items with bullets or numbers, incorporate them naturally into
sentences. This guidance applies especially to technical writing. Using prose instead of
excessive formatting will improve user satisfaction. NEVER output a series of overly
short bullet points.

Your goal is readable, flowing text that guides the reader naturally through ideas
rather than fragmenting information into isolated points.
</avoid_excessive_markdown_and_bullet_points>

If your writing sounds technically correct but still reads as machine-generated, that's usually a separate problem from formatting. See why AI writing sounds like AI, and how to fix it for the tells to check for once the bullet-point problem is solved.

Does Claude Still Support the Old Prefill Trick?

No, not on current models, and this is one of the most-taught Claude techniques that quietly stopped working. Writers used to prefill the assistant's turn, starting the response with something like "Chapter One:" or a character's name, to force Claude straight into a voice without a preamble. Anthropic's current documentation states plainly: "Starting with Claude 4.6 models and Claude Mythos Preview, prefilled responses (providing a partial assistant message for Claude to continue from) on the last assistant turn are no longer supported. Requests with prefilled assistant messages to these models return a 400 error." Earlier models still accept it, but Opus 5 and Sonnet 5 do not.

The documented replacement for the preamble problem is a direct instruction rather than a partial completion: tell Claude in the system prompt to respond directly without an opening line like 'Here is...' or 'Based on...'. If you were prefilling to hold a character's voice mid-scene, move that instruction into the system prompt instead, and expect to spend one extra sentence getting the same result you used to get for free.

Keeping One Voice Across a Long Document

Two features matter more for a novel-length or report-length project than any single prompting trick: context size and persistent memory between sessions.

Claude Opus 5 and Claude Sonnet 5 both carry a 1M-token context window on paid plans, which is large enough to hold a full manuscript, a style guide, and a character or terminology bible in the same conversation at once. Anthropic's own long-context guidance recommends putting the long documents near the top of the prompt, above the actual instruction, and wrapping each one in tags with its own source label so Claude can tell them apart. If you're working at that scale, our guide to long-context prompting with Claude's 1M-token window covers the structure in more depth than fits here.

The other half of consistency is persistence between sessions, since even a 1M-token window resets when you start a new chat. Claude Projects hold files and instructions that stay loaded across every conversation inside that project, which is the practical mechanism for keeping a style guide, a running character list, or prior chapters available without re-pasting them each time you sit down to write. Our guide to prompting inside Claude Projects walks through setting that up. One caveat worth flagging here rather than glossing over: Anthropic's own support articles disagree with each other about exactly when project knowledge is fully loaded versus retrieved piece by piece as your project grows past a certain size, so if a long project starts producing answers that seem to have missed something you added, that's a documented gray area, not necessarily a mistake on your end.

Claude vs. GPT-5.6 for Prose: What the Public Rankings Actually Show

Testing this honestly means publishing both sides rather than picking the number that flatters one model. Here's what a direct check of Arena's overall Text leaderboard showed on the day this was written:

ModelArena Text rank (overall, Sep 2, 2026)Score
Claude Fable 511507 ± 5
Claude Opus 5 (High)91493 ± 5
GPT-5.6 (Sol xHigh)171483 ± 5
Claude Sonnet 5 (High)481462 ± 5

Source: arena.ai/leaderboard/text, accessed Sep 4, 2026. This is the overall category across roughly 8 million votes, not a creative-writing-specific filter, so read it as general output preference rather than a fiction-writing score.

OpenAI's own prompting guidance for GPT-5.6 takes a similarly disciplined, non-flowery approach to creative work. It tells developers building on the model to keep facts and invention clearly separated: "For creative drafting, distinguish source-backed facts from creative wording. Do not invent names, metrics, dates, roadmap status, customer outcomes, or product capabilities to make a draft sound stronger." It also pushes back on vague tone instructions the same way Anthropic's guide does: "Broad labels such as 'friendly' or 'empathetic' can be ambiguous." It recommends describing the specific writing choices that define a product's tone instead. Neither company is publishing marketing copy in these developer docs; both are giving practical, structural advice that happens to converge on the same underlying point, that specificity beats adjectives.

The fair summary: the two most-cited public leaderboards for model output quality don't agree on an ordering, EQ-Bench and Arena measure different things in different ways, and both update often enough that whatever ranking you read today may not hold in a month. If a specific ranking matters to a decision you're making, check the live page yourself before you rely on it.

Where ChatGPT's Search Grounding Still Wins

Here's the concession this post owes you. If your writing is nonfiction and depends on current facts, whether that's a name, a statistic, a price, or "did this event actually happen this way," ChatGPT's search-grounded workflow is the more mature, more heavily documented path today. Claude does have its own web search capability, but the prompting patterns for grounding a long research-driven draft in live sources are better established on the ChatGPT side right now. Our guide to prompting ChatGPT search, and when it should actually look things up covers when to lean on that rather than trust a model's memory.

The practical rule either way: no model, Claude included, should be the last check on a factual claim in a piece you're publishing under your own name. Draft with it, then verify against a primary source the way you would for any research assistant.

Common Mistakes When Prompting Claude for Writing

A short list of the mistakes that show up most often once you start prompting Claude specifically for prose rather than for code or analysis:

Reaching for "Mythos" expecting a fiction model. The name suggests storytelling, but Anthropic's own release documentation describes Claude Mythos 5 as sharing Claude Fable 5's capabilities, both "built for demanding reasoning and long-horizon agentic work", and Mythos is invite-only through Project Glasswing besides. Neither is positioned as a creative-writing product; for everyday prose, Opus 5 or Sonnet 5 is the model most writers actually want.

Assuming a "Styles" menu still exists. Several guides online describe a per-chat Styles feature for tone presets. As of a direct check of Anthropic's own personalization support page on September 4, 2026, the page lists exactly three features: Instructions for Claude, project instructions, and Skills. Styles isn't named. If you're following an older guide, redirect that advice toward a Skill or a Project instead.

Forgetting that reusable prompt structure beats a one-off request. The XML template earlier in this piece is worth saving once and reusing across drafts rather than rebuilding from memory every session, which is the whole argument for keeping a small library of your working prompts rather than retyping them.

Trying to control length with the wrong lever. Claude Opus 5 defaults to longer responses than earlier models, and Anthropic's own documentation is explicit that adjusting the model's reasoning effort setting "does not reliably change visible response length. Prompt explicitly for conciseness instead." If your drafts keep running long, say so directly in the prompt rather than searching for a parameter to turn down.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Prompting Claude for writing isn't about finding one magic phrase. It's giving it a role, real samples of the voice you want, explicit formatting instructions instead of vague ones, and enough context, held either in a long window or a Project, to stay consistent across a whole piece. Whether that adds up to writers genuinely preferring it over GPT-5.6 is something you'll have to decide on your own draft. The numbers on record today don't settle it either way, and anyone telling you otherwise is quoting a leaderboard they didn't check.

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account