If you have played Factorio, you know that every factory has a shape. A smelter line is long and straight. A circuit factory has loops. A train station has buffers. You can tell what a factory makes just by looking at it from above.
I’ve been using pstack, a plugin by poteto that turns a coding agent into something closer to a careful engineer. After a while I noticed the same thing: every pstack playbook has a shape. A bug fix has a loop back to the start. A refactoring has a gate that throws work away. A hillclimb is one big loop.
So I built them as factories.
What pstack is
pstack is a set of skills. You don’t have to learn most of them, because one skill, /poteto-mode, runs the rest for you. You describe the task and how you’ll know it’s done. poteto-mode matches the task to one of 23 playbooks, copies the playbook’s steps into a todo list, and calls the other skills (how, architect, arena, interrogate, and so on) as each step needs them.
The idea behind it is simple: AI writes a lot of code fast, and most of it is slop. pstack doesn’t try to make the agent faster. It makes the agent prove its work.
pstack was built for Cursor. I use Claude Code, so I ported it: pstack-claude. The skills and playbooks are poteto’s. The port only swaps the Cursor-specific parts.
How to read the factories
Every playbook below is one factory. Work flows from left to right.
The caption under each map says what to watch. The list under it is the real playbook: click a step and the factory pauses and lights up the machines that do it. Click it again to resume. The numbers are made up. The shapes are not: each one follows the steps in poteto’s playbook files.
poteto-mode: the router
This is the part you actually type. You never pick a playbook yourself. The same command with different words lands in a different factory:
What you type after /pstack:poteto-mode | Playbook |
|---|---|
| users get two notifications after a retry. repro first, then fix and verify. | Bug fix |
| investigate why background jobs time out every few hours. don’t change any code yet. | Investigation |
| move parsing into one module, zero behavior change. | Refactoring |
| startup takes 1.8s on this fixture. trace it, show me before and after. | Perf |
In Cursor the command is /poteto-mode. In Claude Code, with my port, it’s /pstack:poteto-mode.
Watch the coloured dots. Each splitter pulls off the items with its colour. Items with no dot ride to the end and fall through to figure-it-out.
- Bug fix
- 0
- Investigation
- 0
- Refactoring
- 0
- Perf
- 0
- Feature
- 0
- figure-it-out
- 0
- Playbooks you picked
- 0
- bug report → Bug fix
- question → Investigation
- cleanup → Refactoring
- slow path → Perf
- new behavior → Feature
- no fit → figure-it-out
Five shapes of task get their own splitter here. The real list has 23 playbooks. When none of them fits, the task doesn’t default to Feature. It falls through to figure-it-out, a skill that designs a bespoke playbook for that one task.
Investigation: a factory without a silo
What you type:
/pstack:poteto-mode investigate why background jobs time out every few hours. give me what we know and your best hypotheses. don't change any code yet.
A read-only question (“why does X happen”, “don’t change any code”) is what routes it to Investigation.
Purple-dot questions also stop at the why lab. Orange ones turned out to need a code change, so the splitter sends them out to be re-routed. Everything else ends as a report in the chest. Nothing reaches the No PR blueprint.
- Also asked why
- 2
- Handed back to re-route
- 0
- Files changed
- 0
- PRs opened
- 0
- how-question
- why-question
- actually needs a change
- history found
- report
There is no silo, only a blueprint marked “No PR”. I like that this is a separate playbook, because “explain this to me” and “change this” are different jobs. An agent that mixes them up starts editing files you only asked about. When the answer turns out to need a code change, the playbook hands it back to be re-routed to Bug fix or Feature.
Bug fix: verify on the same machine
What you type:
/pstack:poteto-mode users get two notifications after a retry. repro first, then fix and verify.
A reported defect, plus “repro first”, routes it to Bug fix.
Watch the long green belt: every fix rides back to the red lab that reproduced the bug, and only counts when that lab flashes green. The bisect label halves the suspects each lap.
- Hypotheses ruled out
- 4
- Search passes
- 1
- Repros that needed logging
- 0
- Fixed without a repro
- 0
- bug report
- won't reproduce yet
- reproduced, n suspects
- one cause left
- fix, same repro passes
- red test
Nothing gets past the red lab until the bug actually fires. If it won’t reproduce, the agent adds logging until it does. It doesn’t ask you to reproduce it.
The long green belt is the point. The fix only counts when the original repro passes on the same surface. A unit test that passes somewhere else doesn’t count.
The architect only runs when the fix crosses a function boundary. A one-line fix inside one function skips it.
Last detail: the commit machine drops two items. First a red test, then the fix. In git history the failing test lands before the fix, so anyone can check out the commit before and watch it fail.
Feature: an arena
What you type:
/pstack:poteto-mode add a --json flag to this command. text output stays byte-identical. verify both.
New behavior is what routes it to Feature.
Watch the splitter after the architect. A plan with one obvious shape (blue dot, 1) goes straight to Build B and skips the judge. The rest go to the arena: three runners build the same brief and each surfaces its own shape.
- Features shipped
- 0
- Parts grafted
- 0
- One shape, no arena
- 0
- Interrogated
- 0
- ticket
- plan
- runner brief A / B / C
- one obvious shape
- contested design
- verified feature
When there are several valid ways to build something, pstack doesn’t let one agent pick. It runs an arena: the same brief goes to three runners, each surfaces its own shape, and a judge takes one as the base and grafts the best parts of the other two into it. When there’s one obvious shape, the arena is skipped.
Designs that are contested (orange dot) take a detour through interrogate, where reviewers try to break the change. More on both skills below.
Refactoring: a gate that throws work away
What you type:
/pstack:poteto-mode move parsing into one module, zero behavior change. record the current output first and prove it's unchanged after.
“Zero behavior change” is what routes it to Refactoring.
Watch the two splitters after the yellow lab. Red means the output changed: it loops back under reshape for another small step. Orange means same output but no easier to read: it goes up into the reverted chest.
- Refactors landed
- 0
- Pin breaks caught
- 0
- Reverted, not simpler
- 0
- old module
- reshaped
- pin broke
- not simpler
- same output, simpler
A refactoring must not change what the code does, so the first machine is a pin: a test that records what the old module does today. Two things stand out:
- The recycler comes before the assembler. pstack deletes dead code, one-caller wrappers and old validators before it builds the new shape. Subtract before you add.
- There are two ways to fail. A broken pin goes back for another small step. A change that keeps the output but doesn’t read easier gets reverted. A refactoring that doesn’t lower reader load has no reason to exist.
Perf: seven mantras on one belt
A perf fix starts with a baseline, and that number gets vetted before anyone trusts it (benchmark-checklist). Then the slow path tries the seven performance mantras in order, cheapest first:
- Don’t do it.
- Do it, but don’t do it again.
- Do it less.
- Do it later.
- Do it when they’re not looking.
- Do it concurrently.
- Do it cheaper.
What you type:
/pstack:poteto-mode startup takes 1.8s on this fixture. trace it, fix the measured cause, show me before and after.
A measured slowness is what routes it to Perf.
The number on each item is its time in ms. Watch for the green dot: the moment a mantra gets a path under 300ms, the next splitter drops it onto the bypass and it skips the rest.
- Target
- ≤ 300ms
- Speedup
- –
- Fixes landed
- 0
- Avg after
- –
- slow path, ms on top
- met the target (≤ 300ms)
- measured fix
The playbook has a stop rule: when an earlier mantra meets the target, stop. That’s the bypass belt. Nobody reaches for “do it concurrently” when “don’t do it” already solved the problem. Then the green lab measures again, and the before and after go into the PR. No number, no PR.
Hillclimb: one big loop
What you type:
/pstack:poteto-mode hillclimb p95 search latency. at least 50% better than baseline and at least 10 attempts. keep a decision log.
A metric, a target and a floor on attempts is the shape Hillclimb asks for.
Watch the two gauges under the silo. Lap 1 is a lucky +55%: the metric gauge passes the target, but the laps gauge is at 1/10, so the loop keeps going. Only when both are full does the rocket go.
- Metric vs baseline
- +0%
- Laps
- 0
- Kept
- 0
- Reverted
- 0
- Stop when
- ≥50% and ≥10 laps
- PRs landed
- 0
- hypothesis #n
- change under test
- faster, tests green
- reverted
Hillclimb is for pushing one metric up over many attempts. The stop rule has two halves on purpose. Without the floor on attempts, a lucky first try ends the run. Every lap writes one row to decision.tsv, so you can read the whole run the next morning. A win only counts if the regression tests stay green too.
After three losses in a row, watch for “plateau → pivot category”. That’s the playbook telling the agent not to stop at the first plateau, but to try a different kind of idea.
Prototype: only blueprints
What you type:
/pstack:poteto-mode prototype a few options for the new dropdown menu. take screenshots for me to compare.
“Prototype” is one of the words that routes straight to this playbook.
The variants are blueprints: throwaway code, deleted after. Watch the green item on the exit belt. The only thing that leaves this factory is a decision, and Feature builds the real thing.
- Variants built
- 0
- Production code
- 0 lines
- Tests written
- 0
- open question
- variant brief A / B / C
- throwaway (deleted after)
- the decision
A prototype exists to answer one question, like “which layout?”. No question, no prototype. The variants are throwaway code in a scratch folder: no framework, no tests. All three sit behind one switcher, so the agent can screenshot and compare them. The output is a decision, not code.
Shipping: stop at the first red
What you type:
/pstack:poteto-mode land the stack.
Asking to land a stack routes it to Shipping.
Watch the rack under the lander. Verdicts come back out of order, and each PR waits in its slot with its light. The lander only takes the next slot from the bottom with a green light. A red light draws the ceiling, and every green PR above it waits.
- Next to land
- #1
- Ceiling
- none
- Green, but above the ceiling
- 0
- PR #n, no verdict yet
- PASS
- FAIL
- merged
Shipping is for a stack of PRs that depend on each other. Every PR gets its own verifier: an agent that didn’t write the code. A green CI run doesn’t count as a verdict. The ceiling is the rule I find most useful: a verified PR on top of an unverified one isn’t safe to land, however green it looks.
Orchestrate: a rolling window
What you type:
/pstack:poteto-mode orchestrate the store migration. own it until every package is converted and merged. i'll check in twice a day.
A project you hand over for days (“own it until…”) is what routes it to Orchestrate.
The coordinator writes briefs, never code. Watch the first brief run alone as the pilot. Only when it lands does the window open to four workers, and each one gets a new brief the moment it finishes.
- In flight
- 0 / 1
- Phase
- pilot
- Programs done
- 0
- Idle worker-s avoided vs batches of 4
- 0
- brief for one unit
- worker output
- verified
- merged
The biggest playbook. One coordinator chat runs a project that takes days and many PRs. It starts with a goal you can count, like “16 units merged”. The coordinator never writes code. Its product is the brief, because a worker can’t ask it a question.
The pilot tests the brief and the verify recipe while a mistake costs one agent instead of fifty. After that, a rolling window beats blocking batches: a batch waits for its slowest worker, a window refills each worker the moment it’s done. Landing runs the whole time. It’s never a final phase.
The skills inside the playbooks
The playbooks call smaller skills as single machines. Here are five of them opened up.
interrogate
This is what’s inside the chemical plant in the Feature factory. One reviewer per configured model reads the same diff with the same prompt and rubric. The adversarial signal comes from the reviewers being different, not from personas. A finding two reviewers raise on their own is the strongest signal. The lead then sorts every finding into four bins, and nothing gets applied automatically.
One honest caveat: upstream pairs a Claude model with a Grok model. In my Claude Code port, all subagents run on Claude models, so the reviewers differ by separate context, not by model family.
/pstack:interrogate the whole branch, but skeptically. no nitpicks unless it's an actual bug or regression.
Both reviewers get the same diff and rubric. Watch synthesis: a finding raised by both is merged into one item tagged 2×, the consensus. The lead then sends every finding into one of four chests.
- Consensus findings
- 0
- Findings raised
- 3
- Dismissed
- 0
- Changes auto-applied
- 0
- same brief per reviewer
- finding #n
- act on
- consider
- noted
- dismissed
arena
This is the fan of A/B/C builders in the Feature factory. The prompt is the contract, and a short rubric that only the picker sees decides the base. The base is the candidate a future maintainer can extend most easily. If all candidates converge, that’s agreement, and nothing gets grafted. If they wildly diverge, the frame was too vague, so the task goes back to be re-framed.
/pstack:arena this, 5 candidates. the cache key format is expensive to change later.
Three candidates build the same prompt in their own worktrees. The judge scores all three against the rubric. Usually the base takes one idea from each loser. If they all agree, nothing is grafted. If they wildly diverge, the red loop sends the task back to be re-framed.
- Artifacts shipped
- 0
- Converged, no graft
- 0
- Re-framed
- 0
- same brief, candidate A / B / C
- candidate
- diverged → re-frame
- verified artifact
architect
This is the drafting table in the Bug fix and Feature factories. It designs before code: at least two structurally different designs, screened against a list of red flags, then one sketch to implement against. When implementation keeps hitting the same workaround, the sketch gets thrown out instead of patched.
/pstack:architect design the import pipeline before writing any code. i care most about how callers use it.
Every task gets at least two structurally different designs before any code. The screen picks one sketch. Most sketches skip the white checkpoint; it only runs when you ask. When implementation keeps hitting the same workaround, the red loop throws the sketch out and starts again from grounding.
- Sketches scrapped
- 0
- Human checkpoints
- 0
- grounded brief per design
- design sketch
- code against the sketch
- scrap → back to ground
swarm
Swarm fans out workers, one brief each, and returns one report. A result without its evidence doesn’t count: that worker gets respawned once, and after a second miss the slice is a gap. A gap is not a pass.
/pstack:swarm check every package under packages/ against its check.sh. one worker per package. one report.
Four workers, one slice each. Results come back with a dot: green PASS, orange ISSUES, red BLOCKED. A purple dot means the result skipped its evidence: that worker gets one respawn on the red loop. Miss twice and it's a gap.
- Respawned once
- 0
- Gaps (not a pass)
- 0
- standalone brief, slice n
- PASS
- ISSUES
- BLOCKED
- no SHAs or method
blast-radius
Listing callers is not the job, the agent can grep those in a second. Blast-radius looks for the one fact a change is safe because of, and then proves it by running real code. The staircase is the skill’s own confidence ladder. A writeup that sounds right is worth nothing until the fact reaches step 4.
/pstack:blast-radius what could this change to the session cache break?
Watch the staircase. The one fact climbs while the green lab runs real code. Only step 4 (ran it) or step 5 (reproduced in the app) counts as proven. A fact that stalls lower leaves as unproven.
- Marked unproven
- 0
- Callers listed
- not the job
- the change
- the one fact
- proven (step 4+)
- unproven
What all the factories have in common
Look at them again. Every factory has a machine whose only job is to check, and every checker can send work back or stop the line:
- Bug fix: the repro lab.
- Feature: the judge and interrogate.
- Refactoring: the pin gate.
- Perf and hillclimb: the measuring labs.
- Shipping: the verifiers and the ceiling.
None of the playbooks make the assembler faster. The coding part was never the slow part. In my Delivery Factory, AI made dev fast and the work piled up in review and QA. pstack puts the review and QA inside the agent’s own factory, so less broken work reaches you.
Try it
In Cursor:
/add-plugin pstack
In Claude Code, with my port:
/plugin install pstack --marketplace alexanderop/pstack-claude
Then start every non-trivial task with /pstack:poteto-mode (or /poteto-mode in Cursor), and say what done looks like.