Few-shot chain of thought examples work best when the demonstrations mirror the reasoning structure your task demands, not just the input-output format. Most published examples are too clean — they solve textbook problems but fail on messy real-world inputs where the model must decide which reasoning pattern applies. Below are six patterns that hold up under production traffic, each with a template you can adapt.
1. Math word problems with explicit unit tracking
The failure mode for math problems isn’t arithmetic — it’s unit confusion and implicit assumptions. Force the model to declare units at every step and verify dimensional consistency before computing.
{
"messages": [
{"role": "system", "content": "Solve step by step. Declare units for every number. Verify dimensions match before each operation."},
{"role": "user", "content": "A factory produces 2,400 widgets per 8-hour shift. Each widget requires 3.2 grams of alloy. The alloy costs $18.50 per kilogram. What is the alloy cost per shift?"},
{"role": "assistant", "content": "Step 1: Widgets per shift = 2,400 widgets/shift (given)\nStep 2: Alloy per widget = 3.2 g/widget (given)\nStep 3: Total alloy per shift = 2,400 widgets/shift × 3.2 g/widget = 7,680 g/shift\nStep 4: Convert to kg: 7,680 g/shift ÷ 1,000 g/kg = 7.68 kg/shift\nStep 5: Alloy cost = $18.50/kg (given)\nStep 6: Cost per shift = 7.68 kg/shift × $18.50/kg = $142.08/shift\nAnswer: $142.08 per shift"},
{"role": "user", "content": "A pump moves 1,200 liters per hour. The fluid density is 0.85 kg/L. The tank holds 15,000 kg. How many hours to fill?"}
]
}
The key is the system instruction: “Declare units for every number. Verify dimensions match before each operation.” This prevents the model from multiplying liters by dollars or dividing kg by hours. In production, prepend a small schema of known units (currency, time, mass, volume) so the model can catch mismatches like “kg per shift” vs “kg per hour” before they propagate.
2. Code generation with property-based test cases
Don’t ask for code. Ask for code plus the properties it must satisfy, then demonstrate the verification loop. This shifts the model from “write something that looks right” to “write something that passes these invariants.”
# Few-shot prompt template
SYSTEM = """Write a Python function that satisfies the given properties.
Include a test block using hypothesis that verifies each property.
If a property fails, revise the implementation and re-test."""
EXAMPLES = [
{
"task": "Implement run-length encoding for strings.",
"properties": [
"decode(encode(s)) == s for all strings s",
"len(encode(s)) <= len(s) * 2 for all strings s",
"encode('') == ''"
],
"solution": '''def encode(s: str) -> str:
if not s:
return ""
result = []
count = 1
for i in range(1, len(s)):
if s[i] == s[i-1]:
count += 1
else:
result.append(f"{count}{s[i-1]}")
count = 1
result.append(f"{count}{s[-1]}")
return "".join(result)
def decode(s: str) -> str:
result = []
i = 0
while i < len(s):
count = 0
while i < len(s) and s[i].isdigit():
count = count * 10 + int(s[i])
i += 1
if i < len(s):
result.append(s[i] * count)
i += 1
return "".join(result)
# Property tests
from hypothesis import given, strategies as st
@given(st.text())
def test_roundtrip(s):
assert decode(encode(s)) == s
@given(st.text())
def test_length_bound(s):
assert len(encode(s)) <= len(s) * 2
def test_empty():
assert encode("") == ""''''
}
]
The pattern: state properties first, then show an implementation that passes them. When the model sees the test block as part of the expected output, it internalizes the verification step. For production, store property schemas separately and inject them dynamically — this lets you reuse the same few-shot structure across different coding tasks while keeping the property definitions maintainable.
3. Logical reasoning with explicit contradiction checks
Models hallucinate intermediate conclusions when the reasoning chain exceeds 3-4 steps. Force a “checkpoint” step where the model restates known facts and verifies no contradiction exists before proceeding.
{
"messages": [
{"role": "system", "content": "Reason step by step. After every 3 inference steps, pause and list all established facts. Explicitly check for contradictions before continuing."},
{"role": "user", "content": "All engineers use Git. Some Git users use Vim. All Vim users hate Emacs. Sarah is an engineer. Does Sarah hate Emacs?"},
{"role": "assistant", "content": "Step 1: Sarah is an engineer (given)\nStep 2: All engineers use Git (given) → Sarah uses Git\nStep 3: Some Git users use Vim (given) → Sarah *might* use Vim, not certain\n--- CHECKPOINT ---\nEstablished facts:\n- Sarah is an engineer\n- Sarah uses Git\n- Some Git users use Vim (existential, not universal)\n- All Vim users hate Emacs\nNo contradictions.\nStep 4: 'Some Git users use Vim' does not imply Sarah uses Vim\nStep 5: Cannot conclude Sarah uses Vim\nStep 6: Cannot conclude Sarah hates Emacs\nAnswer: Unknown / cannot be determined from premises"},
{"role": "user", "content": "All production services have alerts. Some alerts page on-call. No paging alerts are silent. The payment service is a production service. Does the payment service have a paging alert?"}
]
}
The checkpoint forces the model to externalize its working memory, making gaps visible. The phrase “existential, not universal” in the example teaches the model to distinguish quantifier scope — a common failure point. In practice, add a rule: “If a checkpoint reveals an existential quantifier where a universal is needed, state the gap explicitly and stop.” This prevents the model from bridging the gap with an unjustified assumption.
4. Multi-hop QA with evidence citations
Retrieval-augmented generation fails when the model synthesizes an answer without tying each claim to a specific source chunk. Require inline citations at the sentence level, and demonstrate what “insufficient evidence” looks like.
{
"messages": [
{"role": "system", "content": "Answer using only the provided context. Cite sources inline like [doc3]. If context is insufficient, say 'Insufficient evidence' and list what's missing."},
{"role": "user", "content": "Context:\n[doc1] Acme Corp Q3 revenue: $42.1M, up 12% YoY.\n[doc2] Acme Corp Q3 operating margin: 18.3%, down 210 bps.\n[doc3] Industry average Q3 operating margin: 22.1%.\n[doc4] Acme Corp announced layoffs affecting 8% of workforce on Oct 15.\n\nQuestion: Why did Acme's operating margin decline in Q3?"},
{"role": "assistant", "content": "The context does not state the cause of the operating margin decline. Revenue increased 12% YoY [doc1], but operating margin fell 210 bps to 18.3% [doc2], below the industry average of 22.1% [doc3]. Layoffs were announced Oct 15 [doc4] but no causal link to Q3 margin is provided.\nAnswer: Insufficient evidence. Missing: breakdown of operating expenses (R&D, SG&A, COGS), one-time charges, or management commentary linking specific factors to margin compression."},
{"role": "user", "content": "Context:\n[doc1] ...\nQuestion: ..."}
]
The example shows a “negative” case — the model correctly refuses to hallucinate a cause. This is critical: if your few-shot examples only show successful answers, the model learns to confabulate. Include at least one “insufficient evidence” demonstration per 3-4 examples. For production, parse the citations post-generation and verify each cited doc ID exists in the retrieved set; reject or retry if citations are fabricated.
5. SQL generation with schema-aware reasoning
Text-to-SQL breaks on implicit joins, ambiguous column names, and business logic not encoded in the schema. Demonstrate a “schema inspection” step where the model lists relevant tables, columns, and foreign keys before writing the query.
-- Few-shot example
/* Task: Find customers who ordered product 'Widget A' in Q3 2024 but never ordered 'Widget B' */
-- STEP 1: Inspect schema
-- customers(id, name, email, created_at)
-- orders(id, customer_id, order_date, status)
-- order_items(id, order_id, product_id, quantity, unit_price)
-- products(id, name, category, list_price)
-- STEP 2: Identify joins and filters
-- customers -> orders (customer_id)
-- orders -> order_items (order_id)
-- order_items -> products (product_id)
-- Filter: products.name = 'Widget A', orders.order_date BETWEEN '2024-07-01' AND '2024-09-30'
-- Anti-join: NOT EXISTS for 'Widget B' across all time
-- STEP 3: Write query
WITH widget_a_customers AS (
SELECT DISTINCT c.id, c.name
FROM customers c
JOIN orders o ON c.id = o.customer_id
JOIN order_items oi ON o.id = oi.order_id
JOIN products p ON oi.product_id = p.id
WHERE p.name = 'Widget A'
AND o.order_date >= '2024-07-01'
AND o.order_date < '2024-10-01'
AND o.status = 'completed'
)
SELECT wac.id, wac.name
FROM widget_a_customers wac
WHERE NOT EXISTS (
SELECT 1
FROM orders o2
JOIN order_items oi2 ON o2.id = oi2.order_id
JOIN products p2 ON oi2.product_id = p2.id
WHERE o2.customer_id = wac.id
AND p2.name = 'Widget B'
AND o2.status = 'completed'
);
The three-step structure (inspect → plan → write) mirrors how a human engineer approaches the problem. The comments are not decorative — they’re the chain of thought. In production, feed the actual schema (or a sampled subset for wide tables) as a preamble before the few-shot examples. This grounds the model in real column names and types, preventing hallucinated identifiers like customer_name when the column is name.
6. Debugging logs with hypothesis elimination
Incident response requires generating hypotheses, then eliminating them with evidence. Show the model how to structure a differential diagnosis and cross off causes as logs rule them out.
{
"messages": [
{"role": "system", "content": "You are debugging a production incident. List initial hypotheses. For each log entry, update the hypothesis list: CONFIRMED, RULED OUT, or NEEDS MORE EVIDENCE. State the decisive evidence for each ruling."},
{"role": "user", "content": "Service: payment-api\nError: 500 on POST /charge, latency p99 > 30s\n\nHypotheses:\n1. Database connection pool exhaustion\n2. Upstream provider timeout (Stripe)\n3. Memory leak causing GC pauses\n4. Bad deploy (config flag flip)\n5. Deadlock in application code\n\nLogs:\n[1] 2024-01-15 14:22:11 db.pool: active=18/20, idle=2, wait_queue=0\n[2] 2024-01-15 14:22:12 stripe.client: request_id=ch_abc, latency=28.4s, status=200\n[3] 2024-01-15 14:22:13 jvm.gc: pause=12ms, heap_used=42%\n[4] 2024-01-15 14:22:14 deploy: version=v2.3.1, config.charge_timeout=30s (was 10s)\n[5] 2024-01-15 14:22:15 thread.dump: 45 threads, 2 BLOCKED on java.util.HashMap$Node"},
{"role": "assistant", "content": "Hypothesis updates:\n1. Database connection pool exhaustion → RULED OUT. Evidence [1]: pool 18/20 active, wait_queue=0. No contention.\n2. Upstream provider timeout (Stripe) → CONFIRMED. Evidence [2]: Stripe latency 28.4s, status 200 (success but slow). Matches 30s p99.\n3. Memory leak causing GC pauses → RULED OUT. Evidence [3]: GC pause 12ms, heap 42%. Healthy.\n4. Bad deploy (config flag flip) → CONTRIBUTING. Evidence [4]: charge_timeout increased from 10s to 30s. Allows slow Stripe calls to hang instead of failing fast.\n5. Deadlock in application code → NEEDS MORE EVIDENCE. Evidence [5]: 2 threads BLOCKED on HashMap. Could be symptom of Stripe latency (threads holding locks while waiting), not root cause.\n\nRoot cause: Stripe latency spike (external) exacerbated by timeout config increase (deploy). Fix: revert timeout to 10s, add circuit breaker."},
{"role": "user", "content": "Service: auth-service\nError: 401 spikes on GET /validate\nHypotheses:\n...\nLogs:\n..."}
]
The format forces three disciplines: enumerate hypotheses upfront (prevents anchoring), cite specific log lines for each ruling (prevents narrative fallacies), and distinguish “ruled out” from “needs more evidence” (prevents premature closure). The HashMap blocking example also shows a “contributing” category — a config change that didn’t cause the incident but amplified it. In production, pipe structured logs (JSON) into this prompt template automatically; the model only sees the last N relevant entries, not the full firehose.
Summary table
| Pattern | Core discipline | Failure it prevents | Production hook |
|---|---|---|---|
| Math with units | Declare units at every step | Dimensional nonsense | Prepend unit schema |
| Code with properties | State invariants, show tests | Plausible but broken code | Inject property schemas dynamically |
| Logic with checkpoints | Pause every 3 steps, list facts | Quantifier scope errors | Enforce “stop on existential gap” rule |
| Multi-hop with citations | Inline cite or say insufficient | Hallucinated synthesis | Verify citations post-generation |
| SQL with schema inspection | Inspect → plan → write | Hallucinated columns, wrong joins | Feed live schema as preamble |
| Debugging with differential | Hypothesize → cite → rule | Anchoring, narrative fallacy | Auto-pipe structured logs |
These six patterns share a principle: the few-shot examples must demonstrate the cognitive moves you want the model to make, not just the input-output mapping. If the example skips the verification step, the model learns to skip it too. Write the examples the way you want the model to think — painfully explicit, checkpointed, and willing to say “I don’t know.”