Software factories — automated systems that plan, write, test, and ship code beyond a single laptop agent — aren’t a “whether” anymore. They’re a “how.” Ran’s take is rooted in running a real human software shop (Automaze) and an internal AI factory (Cloop): you need extreme constraints to trust machines, and you still need humans in the right places.
The spark
Dex Horthy (Human Layer) wrote about a lights-off factory experiment — always-on, minimal human interaction — and got a lot of slop. Ran’s joking reply: you’re holding it wrong. Not because lights-off is evil, but because removing humans from the wrong seats is the failure mode.
Uncle Bob’s extreme-constraint story rhymes: when unit tests, QA, mutation testing, coverage, linters, and performance gates surround the agent hard enough, maybe you stop reading every line. Trust comes from the harness, not vibes.
Workstation → work environment
Next projects outgrow one machine. Agents should run when you’re offline. Steering from a phone, dedicated “local-cloud” boxes (from the previous episode’s setup) — the station becomes a virtual environment. Principle holds wherever the metal lives.
How Cloop runs (lights on)
Setup. Index the whole project (multi-repo aware), build a code knowledge graph / AST map, ask until it can stand up an instance.
Intake. Tickets from humans, clients, agents — or Ladybug (Sentry-like production catch → verify real bug vs third-party outage → issue). Or a mini-PRD dumped from Claude Desktop brainstorming.
Plan. Factory drafts a PRD against that graph. Scope matters (button color vs army-of-agents feature). A second-opinion agent (GPT + Fable rotation) reviews code + PRD for gaps, first principles, reuse. Human reviews artifacts/diagrams — iterate or approve.
Implement. Break into parallel tracks / Kanban sub-issues (Linear-like). Factory Droid as the coding harness; model routing by task (e.g. GLM for frontend) and cost.
Prove. Unit / QA / e2e. Then a simplification pass (strongest models: efficiency, profiling — not “make it pretty”). Then code + security review (GPT/Claude; Fable for security — with workarounds so it doesn’t refuse as if you’re hacking the NSA).
Loop. Fail → implement → test → simplify → review, up to ~five times, then ping the human (“I’m stuck”). Pass → human gets a PR (plus temp deploy for visual once-over). Merge triggers doc/wiki/graph updates after merge — and a parallel loop rebases main hotfixes into the working branch so the PR doesn’t rot.
Why lights-off fails
Without humans at plan and PR, agents optimize for green tests: mocks, TODOs, false positives, or endless retries and a huge token bill with no off-ramp. Temporary deploy + human eyes on the PR is the anti-slop gate.
Ran’s claim: lights-out removed the wrong human. Keep humans for (1) what needs doing, (2) whether the plan matches, (3) whether it actually got done. Don’t bother them mid-loop while the machine still thinks it can finish — but don’t merge unattended either.
Old school, new tech
We’ve always done loops: write, test, rewind. Machines sped them up; the methodology shouldn’t pretend otherwise. Fancy “loop engineering” language doesn’t replace judgment seats. Human time is expensive; machine time isn’t — so stage humans where proof of work (videos, evidence via Automaze’s proof skill) actually matters.
As models get stronger, where the human sits may shift. The factory shape — constraints + humans at plan and merge — is what Ran would bet on sticking.