TL;DR: Least-to-most prompting is a two-stage technique: first the model decomposes a hard problem into an ordered list of simpler subproblems, then you solve them one at a time, feeding each confirmed answer into the next prompt. It comes from a 2022 Google Research paper and was built specifically for problems harder than any example you can show the model.
What Is Least-to-Most Prompting?
Least-to-most prompting is one specific, named technique inside the broader discipline our prompt engineering guide covers, not a general synonym for writing a well-structured prompt. It comes from a specific paper, "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models", by Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi, all with Google Research's Brain Team. It was first posted on arXiv on May 21, 2022, revised twice, and published as a conference paper at ICLR 2023.
The paper describes the method as two prompting stages that run in sequence:
- Decomposition. You show the model a small number of worked examples of breaking a problem into an ordered list of simpler subproblems, then hand it your actual question. The paper's own description: "The prompt in this stage contains constant examples that demonstrate the decomposition, followed by the specific question to be decomposed."
- Sequential solving. You then answer the subproblems one at a time. Each solving prompt carries three things forward: worked examples of how subproblems get solved, the subquestions already answered along with their answers, and the next subquestion to solve. In the authors' words, that prompt "consists of three parts: (1) constant examples demonstrating how subproblems are solved; (2) a potentially empty list of previously answered subquestions and generated solutions, and (3) the question to be answered next."
That second stage is the actual mechanism. Each new subproblem gets solved with the confirmed answers to every earlier subproblem already sitting in its context, so a five-step problem never asks the model to hold five unverified steps in its head at once. The abstract sums up why: "The key idea in this strategy is to break down a complex problem into a series of simpler subproblems and then solve them in sequence. Solving each subproblem is facilitated by the answers to previously solved subproblems."
How Is This Different From Chain-of-Thought and From Just Splitting a Task?
Two comparisons matter here, and it's worth being precise about both, because the names get used loosely.
Versus chain-of-thought prompting. Our chain-of-thought guide covers a technique that asks a model to show its reasoning inside one continuous answer to one question: "Let's think step by step," followed by a single unbroken chain of intermediate steps and a final answer. Least-to-most prompting is structurally different. It is two separate prompting stages, and the paper built it to address something specific that chain-of-thought does not solve: generalizing to problems harder than any example in the prompt. The authors state the gap directly, describing chain-of-thought prompting as a technique that "often performs poorly on tasks that require generalization of solving problems harder than the demonstration examples," including what the paper calls compositional generalization. Least-to-most prompting exists specifically to close that gap, by decomposing the hard problem into pieces no harder than what the examples already showed.
Versus generic task decomposition. Our task decomposition guide covers the general skill of finding checkable seams in any multi-part job and splitting it into an ordered chain of prompts, something you do by inspecting your own task and deciding where the seams go. Least-to-most prompting is a specific, named version of that idea with one structural difference: the decomposition itself is a prompted step, driven by worked examples of how to decompose that kind of problem, rather than something you work out by hand for each new case. That only pays off when you actually have (or can build) a small set of decomposition examples for your problem type. If you already know exactly how to split your task, plain decomposition is simpler and skips a step; least-to-most prompting earns its keep when the model needs to find subproblems for cases you did not hand-pick in advance.
What Did the Original Research Actually Measure?
This is the part worth reading past a summary of, because the size of the improvement depends enormously on which benchmark you look at.
The clearest result is on SCAN, a benchmark that maps natural-language commands to sequences of actions and is widely used to test compositional generalization, specifically whether a model can handle test cases that are structurally longer than anything in its examples. On SCAN's length split, where the test sequences are longer than every training sequence, the paper reports:
| Method | SCAN, length split | GSM8K, overall | GSM8K, problems needing ≥5 steps |
|---|---|---|---|
| Standard prompting | 16.7% | 17.06% | not reported |
| Chain-of-thought | 16.2% | 60.87% | 39.07% |
| Least-to-most prompting | 99.7% | 62.39% | 45.23% |
(All figures use the code-davinci-002 model, the strongest one tested in the paper.)
Two things stand out. On SCAN's length split, chain-of-thought barely helps over standard prompting at all (16.2% versus 16.7%), while least-to-most prompting jumps to 99.7%, using only a handful of demonstration examples and no extra training or fine-tuning. That's the paper's central claim: chain-of-thought reasons well, but it does not generalize to structurally harder cases the way this decomposition strategy does.
On GSM8K, a set of grade-school math word problems, the overall gain is far more modest: 60.87% to 62.39%. But the paper breaks results down by how many reasoning steps a problem actually needs, and the pattern matches the SCAN story on a smaller scale: on problems needing four steps or fewer, least-to-most prompting and chain-of-thought are close to even, and on problems needing five or more steps, the gap opens up to 39.07% versus 45.23%. The technique's advantage shows up specifically on the harder end of the distribution, which is exactly the population this post is about.
The paper is also candid that decomposition is doing most of the actual work: "We’ve observed that nearly all problems in GSM8K can be accurately solved if the large language models are provided with the correct decomposition of those challenging problems." In other words, finding the right subproblems is close to solving the problem. That's a real result, but it also tells you where the effort in using this technique actually goes.
How Do You Write a Least-to-Most Prompt for Your Own Hard Problem?
The mechanics translate cleanly to a normal chat interface; you don't need API access or any special tooling, just two prompts run in order, with the first prompt's output pasted into the second.
Stage 1: get an ordered list of subproblems, not a solution yet:
STAGE 1 — DECOMPOSE ONLY, DO NOT SOLVE ANYTHING YET
Here is an example of decomposing a problem of this type:
Problem: [worked example matching your problem's structure]
Decomposition: [ordered list of subproblems for that example, easiest
and most foundational first]
Now decompose this problem the same way. Output only the ordered list
of subproblems. Do not answer any of them yet.
Problem: [your actual hard problem]
Stage 2: solve the subproblems one at a time, carrying forward only the confirmed answers so far:
STAGE 2 — SOLVE ONE SUBPROBLEM AT A TIME
Here is an example of solving subproblems of this type, in order:
[worked example: subquestion 1 -> answer 1 -> subquestion 2 -> answer 2]
Confirmed so far:
Subquestion 1: [text]
Answer 1: [confirmed answer]
Subquestion 2: [text]
Answer 2: [confirmed answer]
Now answer only the next subquestion, using the confirmed answers above.
Subquestion 3: [next item from your Stage 1 list]
You run stage 2 once per subproblem, appending the newly confirmed answer to the "Confirmed so far" block each time before asking for the next one. That repetition is the entire mechanism: every later step gets to see every earlier answer already verified, instead of re-deriving it inside one long, unbroken response.
Where this earns its cost over a single prompt is a problem that is genuinely harder in structure than any single example you can write out: a multi-part policy question where later parts depend on how earlier parts get classified, a pricing calculation with several conditional branches, a multi-hop question where each hop's answer changes what the next hop needs to look up. If you can already write out the full worked solution for your problem type in one example, you don't need two stages, plain few-shot prompting or chain-of-thought will do the job. This is also why the same team of researchers wrote decomposition guidance for tasks you haven't fully scoped yet; see our guide to prompting a task you don't understand if the problem is that you can't yet write a decomposition example at all.
When Does This Break Down?
That limitation has a practical consequence: the effort in using least-to-most prompting is mostly in the decomposition example, not in the mechanics of running two stages. If your decomposition example doesn't actually resemble the structure of the problem you're feeding it, the model will produce a plausible-looking but wrong list of subproblems, and every subsequent step inherits that mistake. Two things follow from that:
- Write your decomposition example from a real, representative case, not a simplified toy version of your problem. A decomposition example that's easier or more regular than your actual hard cases will not transfer to them.
- Check the subproblem list before you solve anything. Stage 1's whole output is disposable if it's wrong; catching a bad decomposition before you've spent five follow-up prompts solving its subproblems is the entire point of splitting these into two stages in the first place.
This is also not a technique for problems that are simply long, rather than structurally hard. A ten-item shopping list summarized into one paragraph doesn't need a decomposition stage; nothing about item six depends on how you handled item three. Save the two-stage approach for problems where a later part genuinely depends on how an earlier part got resolved.
Where This Fits in a Real Workflow
The hard part of this technique is almost never the mechanics of running two stages; it's writing a decomposition example that actually fits your problem, going from a rough, half-formed sense that a problem is hard and multi-part to an actual worked example the model can follow. That's the gap between having a hard problem and having a usable Stage 1 prompt, and it's the part worth getting help with rather than staring at a blank prompt box.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An AccountThe Bottom Line
Least-to-most prompting is a specific, named technique from a real 2022 paper, not a synonym for "break your task into smaller pieces." Its actual value, backed by the paper's own numbers, shows up on problems that are structurally harder than any example you can show the model: a jump from roughly 16% to 99.7% on a benchmark built to test exactly that, and a real but narrower gain on hard multi-step math problems specifically. It costs you one extra prompting stage and a decomposition example that has to genuinely match your problem's structure. Reach for it when a problem's difficulty comes from its structure, not just its length, and when you can write one honest worked example of how to split it.