Next Talk: KI-Agenten als QA-Engineer: Automatisierte Qualitätssicherung mit Claude Code & GitHub Actions

November 11, 2026 — QS-Tag, Frankfurt am Main

Conference
Skip to content

pstack's Playbooks, Explained as Factorio Factories

Published: at 

If you have played Factorio, you know that every factory has a shape. A smelter line is long and straight. A circuit factory has loops. A train station has buffers. You can tell what a factory makes just by looking at it from above.

I’ve been using pstack, a plugin by poteto that turns a coding agent into something closer to a careful engineer. After a while I noticed the same thing: every pstack playbook has a shape. A bug fix has a loop back to the start. A refactoring has a gate that throws work away. A hillclimb is one big loop.

So I built them as factories.

What pstack is#

pstack is a set of skills. You don’t have to learn most of them, because one skill, /poteto-mode, runs the rest for you. You describe the task and how you’ll know it’s done. poteto-mode matches the task to one of 23 playbooks, copies the playbook’s steps into a todo list, and calls the other skills (how, architect, arena, interrogate, and so on) as each step needs them.

The idea behind it is simple: AI writes a lot of code fast, and most of it is slop. pstack doesn’t try to make the agent faster. It makes the agent prove its work.

pstack was built for Cursor. I use Claude Code, so I ported it: pstack-claude. The skills and playbooks are poteto’s. The port only swaps the Cursor-specific parts.

How to read the factories#

Every playbook below is one factory. Work flows from left to right.

BeltCarries work. Items queue when the machine ahead is busy.
InserterMoves an item from a belt into a machine, and out again.
AssemblerA step that builds: code, a fix, a candidate.
LabA step that checks: repro, verify, measure.
Drafting tableA step that plans: architect, or a brief for a worker.
RadarReads the code (how).
RecyclerDeletes code: subtract before you add.
Chemical plantAdversarial review (interrogate).
Ghost blueprintNever ships: throwaway code, or a PR that won't happen.
SplitterSends items whose dot matches its own onto the side belt.
Red loopRework: the item goes back around for another try.
Rocket siloA merged PR.
ChestOutput that isn't code, like a report.
Dots and numbersA dot is a verdict or a variant: green pass, red fail, orange contested. A number is a count: suspects, ms, PR #.

The caption under each map says what to watch. The list under it is the real playbook: click a step and the factory pauses and lights up the machines that do it. Click it again to resume. The numbers are made up. The shapes are not: each one follows the steps in poteto’s playbook files.

poteto-mode: the router#

This is the part you actually type. You never pick a playbook yourself. The same command with different words lands in a different factory:

What you type after /pstack:poteto-modePlaybook
users get two notifications after a retry. repro first, then fix and verify.Bug fix
investigate why background jobs time out every few hours. don’t change any code yet.Investigation
move parsing into one module, zero behavior change.Refactoring
startup takes 1.8s on this fixture. trace it, show me before and after.Perf

In Cursor the command is /poteto-mode. In Claude Code, with my port, it’s /pstack:poteto-mode.

Playbook · poteto-mode

Watch the coloured dots. Each splitter pulls off the items with its colour. Items with no dot ride to the end and fall through to figure-it-out.

The playbook · click a step to pause on it
Production stats
Tasks routed0
Bug fix
0
Investigation
0
Refactoring
0
Perf
0
Feature
0
figure-it-out
0
Playbooks you picked
0
On the belts
  • bug report → Bug fix
  • question → Investigation
  • cleanup → Refactoring
  • slow path → Perf
  • new behavior → Feature
  • no fit → figure-it-out

Five shapes of task get their own splitter here. The real list has 23 playbooks. When none of them fits, the task doesn’t default to Feature. It falls through to figure-it-out, a skill that designs a bespoke playbook for that one task.

Investigation: a factory without a silo#

What you type:

/pstack:poteto-mode investigate why background jobs time out every few hours. give me what we know and your best hypotheses. don't change any code yet.

A read-only question (“why does X happen”, “don’t change any code”) is what routes it to Investigation.

Playbook · Investigation

Purple-dot questions also stop at the why lab. Orange ones turned out to need a code change, so the splitter sends them out to be re-routed. Everything else ends as a report in the chest. Nothing reaches the No PR blueprint.

The playbook · click a step to pause on it
Production stats
Questions answered1
Also asked why
2
Handed back to re-route
0
Files changed
0
PRs opened
0
On the belts
  • how-question
  • why-question
  • actually needs a change
  • history found
  • report

There is no silo, only a blueprint marked “No PR”. I like that this is a separate playbook, because “explain this to me” and “change this” are different jobs. An agent that mixes them up starts editing files you only asked about. When the answer turns out to need a code change, the playbook hands it back to be re-routed to Bug fix or Feature.

Bug fix: verify on the same machine#

What you type:

/pstack:poteto-mode users get two notifications after a retry. repro first, then fix and verify.

A reported defect, plus “repro first”, routes it to Bug fix.

Playbook · Bug fix

Watch the long green belt: every fix rides back to the red lab that reproduced the bug, and only counts when that lab flashes green. The bisect label halves the suspects each lap.

The playbook · click a step to pause on it
Production stats
Bugs fixed0
Hypotheses ruled out
4
Search passes
1
Repros that needed logging
0
Fixed without a repro
0
On the belts
  • bug report
  • won't reproduce yet
  • reproduced, n suspects
  • one cause left
  • fix, same repro passes
  • red test

Nothing gets past the red lab until the bug actually fires. If it won’t reproduce, the agent adds logging until it does. It doesn’t ask you to reproduce it.

The long green belt is the point. The fix only counts when the original repro passes on the same surface. A unit test that passes somewhere else doesn’t count.

The architect only runs when the fix crosses a function boundary. A one-line fix inside one function skips it.

Last detail: the commit machine drops two items. First a red test, then the fix. In git history the failing test lands before the fix, so anyone can check out the commit before and watch it fail.

Feature: an arena#

What you type:

/pstack:poteto-mode add a --json flag to this command. text output stays byte-identical. verify both.

New behavior is what routes it to Feature.

Playbook · Feature

Watch the splitter after the architect. A plan with one obvious shape (blue dot, 1) goes straight to Build B and skips the judge. The rest go to the arena: three runners build the same brief and each surfaces its own shape.

The playbook · click a step to pause on it
Production stats
Candidates built0
Features shipped
0
Parts grafted
0
One shape, no arena
0
Interrogated
0
On the belts
  • ticket
  • plan
  • runner brief A / B / C
  • one obvious shape
  • contested design
  • verified feature

When there are several valid ways to build something, pstack doesn’t let one agent pick. It runs an arena: the same brief goes to three runners, each surfaces its own shape, and a judge takes one as the base and grafts the best parts of the other two into it. When there’s one obvious shape, the arena is skipped.

Designs that are contested (orange dot) take a detour through interrogate, where reviewers try to break the change. More on both skills below.

Refactoring: a gate that throws work away#

What you type:

/pstack:poteto-mode move parsing into one module, zero behavior change. record the current output first and prove it's unchanged after.

“Zero behavior change” is what routes it to Refactoring.

Playbook · Refactoring

Watch the two splitters after the yellow lab. Red means the output changed: it loops back under reshape for another small step. Orange means same output but no easier to read: it goes up into the reverted chest.

The playbook · click a step to pause on it
Production stats
Lines deleted0
Refactors landed
0
Pin breaks caught
0
Reverted, not simpler
0
On the belts
  • old module
  • reshaped
  • pin broke
  • not simpler
  • same output, simpler

A refactoring must not change what the code does, so the first machine is a pin: a test that records what the old module does today. Two things stand out:

  1. The recycler comes before the assembler. pstack deletes dead code, one-caller wrappers and old validators before it builds the new shape. Subtract before you add.
  2. There are two ways to fail. A broken pin goes back for another small step. A change that keeps the output but doesn’t read easier gets reverted. A refactoring that doesn’t lower reader load has no reason to exist.

Perf: seven mantras on one belt#

A perf fix starts with a baseline, and that number gets vetted before anyone trusts it (benchmark-checklist). Then the slow path tries the seven performance mantras in order, cheapest first:

  1. Don’t do it.
  2. Do it, but don’t do it again.
  3. Do it less.
  4. Do it later.
  5. Do it when they’re not looking.
  6. Do it concurrently.
  7. Do it cheaper.

What you type:

/pstack:poteto-mode startup takes 1.8s on this fixture. trace it, fix the measured cause, show me before and after.

A measured slowness is what routes it to Perf.

Playbook · Perf issue

The number on each item is its time in ms. Watch for the green dot: the moment a mantra gets a path under 300ms, the next splitter drops it onto the bypass and it skips the rest.

The playbook · click a step to pause on it
Production stats
Stopped at mantra (avg)#–
Target
≤ 300ms
Speedup
–
Fixes landed
0
Avg after
–
On the belts
  • slow path, ms on top
  • met the target (≤ 300ms)
  • measured fix

The playbook has a stop rule: when an earlier mantra meets the target, stop. That’s the bypass belt. Nobody reaches for “do it concurrently” when “don’t do it” already solved the problem. Then the green lab measures again, and the before and after go into the PR. No number, no PR.

Hillclimb: one big loop#

What you type:

/pstack:poteto-mode hillclimb p95 search latency. at least 50% better than baseline and at least 10 attempts. keep a decision log.

A metric, a target and a floor on attempts is the shape Hillclimb asks for.

Playbook · Hillclimb

Watch the two gauges under the silo. Lap 1 is a lucky +55%: the metric gauge passes the target, but the laps gauge is at 1/10, so the loop keeps going. Only when both are full does the rocket go.

The playbook · click a step to pause on it
Production stats
Stop predicatenot yet
Metric vs baseline
+0%
Laps
0
Kept
0
Reverted
0
Stop when
≥50% and ≥10 laps
PRs landed
0
On the belts
  • hypothesis #n
  • change under test
  • faster, tests green
  • reverted

Hillclimb is for pushing one metric up over many attempts. The stop rule has two halves on purpose. Without the floor on attempts, a lucky first try ends the run. Every lap writes one row to decision.tsv, so you can read the whole run the next morning. A win only counts if the regression tests stay green too.

After three losses in a row, watch for “plateau → pivot category”. That’s the playbook telling the agent not to stop at the first plateau, but to try a different kind of idea.

Prototype: only blueprints#

What you type:

/pstack:poteto-mode prototype a few options for the new dropdown menu. take screenshots for me to compare.

“Prototype” is one of the words that routes straight to this playbook.

Playbook · Prototype

The variants are blueprints: throwaway code, deleted after. Watch the green item on the exit belt. The only thing that leaves this factory is a decision, and Feature builds the real thing.

The playbook · click a step to pause on it
Production stats
Decisions made0
Variants built
0
Production code
0 lines
Tests written
0
On the belts
  • open question
  • variant brief A / B / C
  • throwaway (deleted after)
  • the decision

A prototype exists to answer one question, like “which layout?”. No question, no prototype. The variants are throwaway code in a scratch folder: no framework, no tests. All three sit behind one switcher, so the agent can screenshot and compare them. The output is a decision, not code.

Shipping: stop at the first red#

What you type:

/pstack:poteto-mode land the stack.

Asking to land a stack routes it to Shipping.

Playbook · Shipping

Watch the rack under the lander. Verdicts come back out of order, and each PR waits in its slot with its light. The lander only takes the next slot from the bottom with a green light. A red light draws the ceiling, and every green PR above it waits.

The playbook · click a step to pause on it
Production stats
PRs landed0
Next to land
#1
Ceiling
none
Green, but above the ceiling
0
On the belts
  • PR #n, no verdict yet
  • PASS
  • FAIL
  • merged

Shipping is for a stack of PRs that depend on each other. Every PR gets its own verifier: an agent that didn’t write the code. A green CI run doesn’t count as a verdict. The ceiling is the rule I find most useful: a verified PR on top of an unverified one isn’t safe to land, however green it looks.

Orchestrate: a rolling window#

What you type:

/pstack:poteto-mode orchestrate the store migration. own it until every package is converted and merged. i'll check in twice a day.

A project you hand over for days (“own it until…”) is what routes it to Orchestrate.

Playbook · Orchestrate

The coordinator writes briefs, never code. Watch the first brief run alone as the pilot. Only when it lands does the window open to four workers, and each one gets a new brief the moment it finishes.

The playbook · click a step to pause on it
Production stats
Units merged0 / 16
In flight
0 / 1
Phase
pilot
Programs done
0
Idle worker-s avoided vs batches of 4
0
On the belts
  • brief for one unit
  • worker output
  • verified
  • merged

The biggest playbook. One coordinator chat runs a project that takes days and many PRs. It starts with a goal you can count, like “16 units merged”. The coordinator never writes code. Its product is the brief, because a worker can’t ask it a question.

The pilot tests the brief and the verify recipe while a mistake costs one agent instead of fifty. After that, a rolling window beats blocking batches: a batch waits for its slowest worker, a window refills each worker the moment it’s done. Landing runs the whole time. It’s never a final phase.

The skills inside the playbooks#

The playbooks call smaller skills as single machines. Here are five of them opened up.

interrogate#

This is what’s inside the chemical plant in the Feature factory. One reviewer per configured model reads the same diff with the same prompt and rubric. The adversarial signal comes from the reviewers being different, not from personas. A finding two reviewers raise on their own is the strongest signal. The lead then sorts every finding into four bins, and nothing gets applied automatically.

One honest caveat: upstream pairs a Claude model with a Grok model. In my Claude Code port, all subagents run on Claude models, so the reviewers differ by separate context, not by model family.

/pstack:interrogate the whole branch, but skeptically. no nitpicks unless it's an actual bug or regression.
Playbook · interrogate

Both reviewers get the same diff and rubric. Watch synthesis: a finding raised by both is merged into one item tagged 2×, the consensus. The lead then sends every finding into one of four chests.

The playbook · click a step to pause on it
Production stats
Act on0
Consensus findings
0
Findings raised
3
Dismissed
0
Changes auto-applied
0
On the belts
  • same brief per reviewer
  • finding #n
  • act on
  • consider
  • noted
  • dismissed

arena#

This is the fan of A/B/C builders in the Feature factory. The prompt is the contract, and a short rubric that only the picker sees decides the base. The base is the candidate a future maintainer can extend most easily. If all candidates converge, that’s agreement, and nothing gets grafted. If they wildly diverge, the frame was too vague, so the task goes back to be re-framed.

/pstack:arena this, 5 candidates. the cache key format is expensive to change later.
Playbook · arena

Three candidates build the same prompt in their own worktrees. The judge scores all three against the rubric. Usually the base takes one idea from each loser. If they all agree, nothing is grafted. If they wildly diverge, the red loop sends the task back to be re-framed.

The playbook · click a step to pause on it
Production stats
Ideas grafted0
Artifacts shipped
0
Converged, no graft
0
Re-framed
0
On the belts
  • same brief, candidate A / B / C
  • candidate
  • diverged → re-frame
  • verified artifact

architect#

This is the drafting table in the Bug fix and Feature factories. It designs before code: at least two structurally different designs, screened against a list of red flags, then one sketch to implement against. When implementation keeps hitting the same workaround, the sketch gets thrown out instead of patched.

/pstack:architect design the import pipeline before writing any code. i care most about how callers use it.
Playbook · architect

Every task gets at least two structurally different designs before any code. The screen picks one sketch. Most sketches skip the white checkpoint; it only runs when you ask. When implementation keeps hitting the same workaround, the red loop throws the sketch out and starts again from grounding.

The playbook · click a step to pause on it
Production stats
Built against a sketch0
Sketches scrapped
0
Human checkpoints
0
On the belts
  • grounded brief per design
  • design sketch
  • code against the sketch
  • scrap → back to ground

swarm#

Swarm fans out workers, one brief each, and returns one report. A result without its evidence doesn’t count: that worker gets respawned once, and after a second miss the slice is a gap. A gap is not a pass.

/pstack:swarm check every package under packages/ against its check.sh. one worker per package. one report.
Playbook · swarm

Four workers, one slice each. Results come back with a dot: green PASS, orange ISSUES, red BLOCKED. A purple dot means the result skipped its evidence: that worker gets one respawn on the red loop. Miss twice and it's a gap.

The playbook · click a step to pause on it
Production stats
Reports0
Respawned once
0
Gaps (not a pass)
0
On the belts
  • standalone brief, slice n
  • PASS
  • ISSUES
  • BLOCKED
  • no SHAs or method

blast-radius#

Listing callers is not the job, the agent can grep those in a second. Blast-radius looks for the one fact a change is safe because of, and then proves it by running real code. The staircase is the skill’s own confidence ladder. A writeup that sounds right is worth nothing until the fact reaches step 4.

/pstack:blast-radius what could this change to the session cache break?
Playbook · blast-radius

Watch the staircase. The one fact climbs while the green lab runs real code. Only step 4 (ran it) or step 5 (reproduced in the app) counts as proven. A fact that stalls lower leaves as unproven.

The playbook · click a step to pause on it
Production stats
Facts proven by running code0
Marked unproven
0
Callers listed
not the job
On the belts
  • the change
  • the one fact
  • proven (step 4+)
  • unproven

What all the factories have in common#

Look at them again. Every factory has a machine whose only job is to check, and every checker can send work back or stop the line:

None of the playbooks make the assembler faster. The coding part was never the slow part. In my Delivery Factory, AI made dev fast and the work piled up in review and QA. pstack puts the review and QA inside the agent’s own factory, so less broken work reaches you.

Try it#

In Cursor:

/add-plugin pstack

In Claude Code, with my port:

/plugin install pstack --marketplace alexanderop/pstack-claude

Then start every non-trivial task with /pstack:poteto-mode (or /poteto-mode in Cursor), and say what done looks like.

Press Esc or click outside to close

Stay Updated!

Subscribe to my newsletter for more TypeScript, Vue, and web dev insights directly in your inbox.

  • Background information about the articles
  • Weekly Summary of all the interesting blog posts that I read
  • Small tips and trick
Subscribe Now