Back to blog
Engineering13 min read

Prompting for Technical Specs and Design Docs

Technical spec prompts that borrow structure from RFC 2119, Google's design-doc format, and Architecture Decision Records, so a model drafts the shape while you supply the constraints it can't know.

NH
Nafiul Hasan

TL;DR: Technical spec prompts work when you supply what a model cannot know: your system's real constraints, what you already decided, and why. Ask for structure and prose, not content. Borrow RFC 2119's MUST/SHOULD/MAY for testable requirements, Google's design-doc shape for context and tradeoffs, and one Architecture Decision Record per decision instead of one sprawling document.

What can a model actually write in a technical spec?

Roughly half of it, and the half it's good at is genuinely useful. Give a model your rough notes and a decision you've already made, and it will turn them into a goals-and-non-goals list, tighten a design section into readable prose, and format a requirements table without complaint. That is real time saved on a document nobody enjoys writing.

The other half isn't in your notes and never was. What your system currently does, which constraints are real (a latency budget, a compliance rule, a headcount, a deadline), what your team already tried and rejected, and which alternatives were seriously weighed rather than invented to pad a section. None of that is recoverable from a request that only says write a design doc for X. A model asked for it anyway produces something plausible, because plausible is what it does when a fact is missing. That's the same mechanism behind why models make things up in any other context, and a spec is an especially bad place for it to happen, because a fabricated alternatives considered section reads exactly as confident as a real one.

JobFrom your notes aloneWhat you must supply
Structure the documentYesNothing
Write clean prose for a decision you already madeYesThe decision itself
List alternatives consideredNoWhat you actually tried and rejected, and why
State current-system constraintsNoThe real budget, compliance rule, deadline, or headcount
Tag requirements as binding or optionalPartly, once you tell it the ruleWhich ones are genuinely non-negotiable
Draw an architecture diagramNoA text description, or an actual diagramming tool

Which structure should your spec actually follow?

There's no single global standard everyone follows, but three conventions have earned enough real use that borrowing their shape beats inventing your own headers from scratch.

Google's informal design doc. Malte Ubl, describing the practice from his own time at Google, writes that "One of the key elements of Google's software engineering culture is the use of design docs for defining software designs." He's careful to note it isn't a rigid template: "Design docs are informal documents and thus don’t follow a strict guideline for their content." What has held up across projects is a recurring section shape: Context and scope, Goals and non-goals, the actual design, Alternatives considered, and Cross-cutting concerns such as security and privacy. On length, Ubl reports that "The sweet spot for a larger project seems to be around 10-20ish pages." A much shorter version is fine for small changes.

Architecture Decision Records (ADRs). Michael Nygard proposed this format in 2011, and the argument for it is really an argument against the single mega-document. "Nobody ever reads large documents, either", he writes, adding that on more than one project he'd seen "the specification document was larger (in bytes) than the total source code size." An ADR captures one architecturally significant decision at a time, in five parts: Title, Context, Decision, Status, and Consequences. Nygard adds that "The whole document should be one or two pages long." You accumulate ADRs as you go instead of maintaining one document that has to represent every decision at once.

RFC 2119 requirement language. The IETF's 1997 standard defines MUST, SHOULD, and MAY so a requirements list stops being a list of intentions and starts being something a reviewer, and later a test suite, can check. MUST means the definition is "an absolute requirement of the specification." SHOULD means "there may exist valid reasons in particular circumstances to ignore a particular item, but the full implications must be understood and carefully weighed before choosing a different course." MAY means the item is genuinely optional. The standard also warns against overusing them: imperatives of this kind "must be used with care and sparingly." They should not be stamped on every line for emphasis.

FrameworkWhat it's actually forSource
Google-style design docThe whole narrative: context, goals, tradeoffsMalte Ubl, Design Docs at Google, 2020
Architecture Decision RecordOne architecturally significant decision at a timeMichael Nygard, Documenting Architecture Decisions, 2011
RFC 2119 keywordsTagging individual requirement lines as binding or optionalIETF RFC 2119, 1997

How do you stop a model from inventing your system's constraints?

By never sending a prompt that asks for a spec without also sending the facts. The single highest-leverage habit here is a context dump that states what's real and explicitly marks what isn't known, so the model has nothing left to guess at.

Do not write the spec yet.

Here are the facts about the current system. Anything not listed here is
UNKNOWN, and you must never assume a value for an UNKNOWN:

CURRENT BEHAVIOUR: <what the system does today, in your own words>
CONSTRAINTS: <latency budget, compliance rule, deadline, team size, whatever is real>
ALREADY TRIED: <what was attempted before and why it was rejected, or write NONE>
MY DECISION: <what you have already decided, or write UNDECIDED>

Read this, then list back every UNKNOWN you still need from me before you
draft anything. Do not fill a gap with a plausible guess.

This mirrors a workflow already covered for a different audience: turning meeting notes into a spec works the same way, because the failure mode is identical whether the missing input is a product decision or an engineering constraint. The model cannot ask a question it doesn't know it should ask, so making the gaps visible before drafting is the actual fix, not a nicer-sounding prompt.

What's a base prompt for turning notes into a first draft?

Once the facts are in hand, this is the one to save. It borrows the section shape from Google's convention and leaves the decision itself entirely to you.

You are drafting a technical spec from my notes and decisions below. Do not
add a decision, alternative, or constraint I did not give you.

CONTEXT AND SCOPE (facts about the landscape, not requirements):
<paste>

GOALS:
<paste>

NON-GOALS (things that could reasonably be goals but aren't, on purpose):
<paste>

MY DECISION (the design itself, in my own words):
<paste>

ALTERNATIVES I ACTUALLY CONSIDERED (or write NONE):
<paste>

CONSTRAINTS:
<paste>

OUTPUT — these sections, in this order:
## Context and scope
## Goals and non-goals
## Design
Tighten my decision into clear prose. Do not add reasoning I did not give you.
## Alternatives considered
Only the ones I listed. If I wrote NONE, write "No alternatives were formally
evaluated" rather than inventing any.
## Open questions
Anything you had to mark UNKNOWN above, listed as questions for me.

RULES:
- Every sentence must trace to something I gave you.
- No adjectives about quality: not "robust," not "clean," not "scalable"
  unless I used that word myself.

The No alternatives were formally evaluated fallback matters more than it looks. A model left to fill that section on its own will invent two or three plausible-sounding options and then explain why yours beats them, which reads exactly like real analysis and is entirely fabricated.

How do you turn one decision into an ADR without rewriting the whole doc?

By scoping the prompt to exactly one decision, the way Nygard's format intends. This is the prompt worth reaching for mid-project, when a decision needs recording but the surrounding document isn't finished yet.

Write one Architecture Decision Record for a single decision. Use exactly
this format and nothing more:

Title: <short noun phrase, e.g. "Use Postgres row-level security for tenant
isolation">
Context: <the forces at play, stated as facts, not argument. Paste mine below>
Decision: <state it in one or two full sentences, active voice, "We will...">
Status: proposed
Consequences: list every consequence, positive, negative, and neutral. Do not
list only the upside.

MY CONTEXT: <paste>
MY DECISION: <paste, in your own words>
ALTERNATIVES I WEIGHED: <paste, or write NONE>

Keep the whole thing to one or two pages. Do not add a consequence I have not
implied from what I gave you.

File these in sequence as you go, rather than trying to fold every decision into one living document that nobody can review in a single sitting.

How do you make requirements testable instead of vibes?

By tagging every requirement line the way RFC 2119 defines, and then treating the tag as a rule rather than decoration.

Here is my requirements list, written as plain sentences:
<paste>

Rewrite it as a table with columns: Requirement, Level, Rationale.

Level must be exactly one of MUST, SHOULD, or MAY, using the RFC 2119
definitions: MUST is non-negotiable, SHOULD allows deviation only for a
documented reason, MAY is genuinely optional.

For each row, Rationale states why that level applies in one sentence. If you
cannot justify a level from what I gave you, mark it NEEDS REVIEW instead of
guessing.

Do not tag more than a third of the list as MUST. If everything is
non-negotiable, nothing is, and that itself should be flagged back to me.
LevelWhat it commits you to
MUSTA test that fails on violation should block a release
SHOULDA deviation needs a written reason someone can review later
MAYAbsence is a legitimate choice, not a bug

That last rule in the prompt, capping how much gets tagged MUST, exists because a requirements list where every line is marked non-negotiable is a list nobody actually enforces. This table is also the direct input to a narrower job: turning a MUST-tagged requirement into an actual test case is a structured-data prompting problem more than a writing one, and it's exactly what feeds prompting for test generation that finds real bugs. A requirement you can't phrase as a testable MUST line is usually a requirement nobody has actually agreed on yet, which is worth discovering before the spec ships, not after.

How do you keep the doc from going stale before the sprint ends?

By separating what changes often from what doesn't, and only re-generating the part that moved. A spec that bundles a stable architecture with fast-changing task lists goes stale in the second half within a week, and then nobody trusts any of it.

Log decisions as you make them, as separate ADRs, instead of editing one master document each time. Prompt versioning applies just as well to the prompts that generate these documents as it does to the documents themselves: keep the context-dump prompt and the base-draft prompt as saved, reusable templates rather than retyping them per project, so the only thing that changes between specs is the facts you paste in.

The facts themselves are also worth saving rather than retyping. Prompt Architects' Global Variables exist for exactly this: reusable blocks, including recurring technical specs, that drop into any prompt instead of getting re-pasted from memory every time a constraint comes up, and the Personal Context Library holds the project-level facts (what the system does today, who owns it, what's already been tried) so a new spec prompt doesn't start from zero. It's a saved library and context system, not a documentation platform: there's no wiki, no Confluence or Notion integration, and nothing here writes the doc to a shared drive for you. The free plan includes 5 prompt enhancements per day, forever, per our FAQ page, which is enough to try the context-dump prompt above without committing to anything.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Where do these prompts reliably go wrong?

Four ways, and all four are cheap to catch once you know to look.

The invented constraint. A model given write a spec for a rate limiter with no context will assume a plausible one: Redis, a sliding window, a round number for the limit. None of that may be true of your system, and it reads exactly as confident as a real constraint would.

The fabricated alternative. Covered above, and the most damaging one, because a reviewer who trusts the alternatives considered section is trusting analysis that never happened.

Decision blur. A spec that reads as settled when the actual decision is still open. This is the cost of skipping the UNDECIDED marker in the context-dump prompt: the model's confident prose makes an open question look closed.

The document nobody updates. A single sprawling spec drifts from reality the moment the first decision changes, and updating prose is more friction than updating a table or filing a new ADR. Splitting by decision, the way ADRs intend, is what keeps this from happening.

A checklist before you share the doc

Six checks, a few minutes each.

  1. Read the open questions section first. Anything the model marked UNKNOWN needs an answer or an explicit UNDECIDED, not a quiet guess.
  2. Check alternatives considered against what you actually discussed. If you wrote NONE and the section isn't empty, the model invented options.
  3. Check every MUST tag is something you'd block a release over. If you can't defend that, it's a SHOULD.
  4. Check the constraints section matches reality, not a plausible-sounding default for a system like yours.
  5. Check the doc's scope against its length. If it covers three decisions, it should probably be three ADRs plus a short design doc linking them, not one document.
  6. Check who inherits this. An ADR or design doc is written for the person who reads it after the decision is old news, not for the reviewer approving it today.

The failure this list guards against isn't a bad spec. It's a fluent one that states something nobody actually decided, filed somewhere a future engineer will trust without checking.

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account