Stop Babysitting What You Built
Turn the signals your product already emits into one queue, with agents that find, agents that finish, and you reviewing the exceptions.
Ben Sufiani · The article: The Self-Driving Startup
The problem is invisible babysitting
A ticket buyer missed the welcome sequence. Ad tracking looked plausible and was still wrong. The website stayed online while parts of the intended experience quietly failed.
These are not dramatic outages that force attention. They are the small gaps that accumulate when nobody is continuously comparing what should happen with what did.
“Some people who purchased a ticket didn't get a seat assigned. If they didn't get a seat, they didn't get the welcome emails – but how would I know?”
Agents that find turn systems into signals
Logs, session replays, errors, database records, payment events, authentication, and dependency alerts all know something about the product. Ben's daily scout reads those sources and translates the useful findings into one work queue.
Its job is coverage, not completion. It records a stable problem with enough evidence for the next role to judge and act on.
“I don't have the time to look through 39 different signals. So what does an agentic coder do? Let the agent do it.”
Agents that finish take one issue through proof
A second routine runs more often and chooses one worthwhile task. It investigates, changes the product, verifies the outcome, and watches production before marking the work done.
When the action involves a product decision or uncertain tradeoff, it assigns the question to Ben. He reviews the exception instead of supervising every step.
“Whenever it puts something in review, that means a human is required, and it even assigns me the task.”
Keep the two roles and one queue
The finder needs permission to scan widely and leave work behind. The finisher needs permission to narrow down and complete one thing. Mixing those jobs in one run makes the broad sweep compete with the careful finish.
The queue is the handoff between them and the record for the human. Start with one source, run the finder manually, let a finisher attempt one item, and lengthen the leash only after the proofs are trustworthy.
“I differentiate between two agents, two routines – one that finds the problems once a day, then the ones that run every four hours that fix the problem.”
Dive deeper
Read the complete operating model, the proof gates, and the seven-day path to your first source.
Read the insightBring the source your business already pays for and decide what its first agent should watch.
Register freeKeep the full library of operating sessions available with the Pirate Pass.
See the Pirate Pass