---
title: "My AI Team Followed Three Rules. It Kept Two."
date: "2026-10-04"
excerpt: "Watch an AI team judge three rules it followed, keep two, and delete one with no proven benefit as goals and retrospectives help it manage its work."
ogImage: "/images/blog/my-ai-team-followed-three-rules-it-kept-two.jpg"
tags: ["Agents", "Agile", "Retrospective", "Orchestrator", "Goals"]
---

At eleven last night a message arrived in my Slack channel: the schedule was at risk. Below that came what was done, what had come back for rework, what was still waiting, and what the team meant to do next. I had not asked for it.

This is the team from [my last post](/blog/the-day-i-left-the-team): two AI directors who decide what to build, executors who build it, and an orchestrator that moves work between them on a board of markdown files. The goal the message was about had not been reached, and it still hasn't. A week ago, though, this team could not tell it was late until it already was, and when it got stuck it waited for me. That afternoon it had also deleted one of its own rules, one it had been following. A few things had to exist before either could happen.

## Something it could miss

The first was a way to see trouble coming. The team finished one item and picked up the next, and nothing in that said whether the whole was on time. If a result is due at six, I want to hear at five that it will be late, while something can still be done. Goals gave it that: a file per period, with a due time, checkpoints, and the risks that should make it come to me.

The first goal the team wrote was due at seven in the evening and declared met at 9:13 that morning. The small test behind it had failed, and the report said so plainly, but the goal had been written so that handing over a result counted as meeting it. I wrote back that a goal finished by mid-morning was only a checkpoint. From then on a goal names an outcome, worked back from the purpose above it, and the outcome is either something of value or a failure that tells us what to change.

Its estimates looked like a person's. My guess is that models learn to estimate from what people wrote, so a job they call a month's work is done in three days. I asked them to estimate from their own record instead, and a script now writes it every night from the agent logs. The next goal was planned at 30.2 hours, took 27.4, and ended with my choice of a design.

The sizing instruction also asked for goals the team would reach about half the time with real effort, the way a stretch objective is set. The directors read "even odds" as a probability they would have to prove. They could not prove it, so they wrote "not even odds" beside every estimate, and with nothing left to check the numbers, a 27-hour plan ended up resting on a design goal that had taken 27.4 hours, a different kind of work. I took the stretch out. What I need from a due time is the team's best estimate, with no margin, because the gap between plan and finish is what the next retrospective learns from.

## A promise to someone

A goal also had to survive the work of reaching it. The team kept finding a problem halfway and treating it as a reason to start something else. Now the means can be questioned at any time while the goal holds until it is judged, and a problem found on the way is fixed inside it. If the evidence makes the goal pointless, the team abandons it out loud. There is no third way out.

The harder part was the end. On one goal the team checked its build against conditions it could verify itself: labels legible, failures explained, a run that could be stopped. The reviewer rejected it four times and passed it on the fifth, the team marked the goal met, and the result came to me at 3:11 in the morning. I tried it and wrote back that I could not judge the idea from it. The orchestrator turned my reply into a request to redo the work. I asked whether the goal had been reached or not, and whether anyone meant to settle it.

It was settled as missed. Since then a goal carries key results the team and I can both observe, and it is not met until I agree it is.

Yesterday afternoon that was tested. A build reached me well ahead of its due time, and again I could not tell whether the idea held. The look was the best they had shown me, and working it felt like filling in a web form. Nobody marked anything met this time. The orchestrator wrote back that the test had failed and the premise was still untested, kept the goal open, and the team went back to work inside it.

## Rules it can change

I keep the studio's base instructions. The team keeps its own working agreements, one short file per role, and changes them in a retrospective, where it looks back at how the work went. That gives it a place to put what it learns without waiting for me.

It is also a place to put a mistake. One retrospective added a rule telling the directors to reopen any premise my words seemed to question. A few hours later they were halfway through a test that was nearly done, my last message changed one detail about the audience, and under the new rule they reopened the subject and started comparing three new ones. The test was left where it stood.

Every change to the agreements now has to say what effect it should have. New work waits until the changes are in, and the next retrospective checks whether the effect showed up. It starts by deleting agreements whose contribution the evidence does not show, since following a rule does not, by itself, show that it helped. It also follows every goal, including the ones that go well. The old version opened one only after a miss or a very early finish, and I asked the obvious question: if a goal ends one second past its halfway mark, does nobody look back?

Yesterday afternoon, a retrospective had three rules to judge, all added that morning and all followed.

The first was meant to keep the directors from testing only what agents can see. It told them to choose the observations that would decide a hypothesis, and who could supply them, from the claim itself. They did. My answer still confirmed nothing about the hypothesis, because the test had failed for another reason, so there was no sign the rule had changed anything. It was deleted, with a note on when a narrower version should come back.

The second told the orchestrator, when handing me something to judge, to lay out the hypothesis, how the thing in front of me was supposed to produce what it claimed, and what each observation would change. It stayed, and the evidence was me. When I told them the design was not good enough, I had quoted those lines from their packet and argued from them.

The third told the orchestrator to keep a failed test apart from a failed premise when it settled my answer. It stayed too. It is why the goal was still open that afternoon.

The handoff as a whole had failed. Inside it, one rule had shown no clear benefit, one had given me the words to say why, and one had kept the failure from being written up as the wrong kind.

Neither of the two it kept is in the agreements now, and not because they stopped helping. What they asked for has since been written into the goal and the design documents, and the team deleted them rather than say it twice.

## What the practices were for

These are agile practices, and I have worked this way with human teams for a long time. Goals with checkpoints, a review with the person the work is for, a retrospective that changes how the team works: each of them exists to get a result out of uncertain work quickly, and to find out early when that is not happening. I brought them in one at a time, as the team stalled, at whatever size the problem needed.

Handing the building to agents did not remove the problems they answer. The team could build fast and still had to be told what counted as a result. A build could pass every check the team ran and leave me unable to judge the idea. And a team that could rewrite its own agreements still had to find out whether a change had helped.

The repairs that failed here went wrong the way these practices go wrong with people: the form was there and the reason was not. None of them was fixed by dropping the practice. Each was fixed by putting the reason back, and the reason decided which parts stayed.

## What happened without me

Last night an independent review turned back a new screen on five conditions before it reached me. At 22:01, after a piece of repair work came back twice for the same reason, the team opened a retrospective on its own process. And at the 23:00 checkpoint, the orchestrator sent the message I started with.

It said the schedule was at risk, what had been done, and what came next. A week earlier, almost every change of course had started with a message from me. The goal is still open, and the plan to recover it is theirs to work out.
