Week 9: Noj Got Faster Than Its Human Gate
Every piece of process Noj added this week routes through me, and the list of things waiting on me went from 19 to 27 in seven days. Phased tickets, ticket dependencies, and two tickets that ran out of rounds and came back to me as designed. Plus what happened on my toy apps when I stopped reviewing.
Noj is the control room I write about here, which I used to call Mission Control and describe as a software factory. Every morning Noj files a standup issue for me, and one of its sections is “Blocked on you”. On 26 September it listed 19 tickets. On 2 October it listed 27. It went up on all but one day in between, and the standup on 3 October showed 33.
The board tells the same thing. 127 issues were opened this week and 89 closed. Last week it was the other way around, 99 opened and 119 closed.
Most of what landed in Noj this week is more process, and every new piece of it ends at me.
Phased tickets
Since 29 September a ticket can be phased: a plan on Opus, execution on Sonnet, a review on Opus, and QA on Sonnet, or an audit on Opus when the ticket is marked high risk. Nothing gets executed until I reply “Plan approved” to the latest plan. Each phase hands off to a new session on its own model, and the dispatcher now starts that session on the model the ticket asks for. Before this, every session it started ran Sonnet, even on tickets labelled for Opus.
Why borrow a checkpoint from uzi, which I wrote about in last week’s post? I was trying to get kairos-lab to work properly with a shared network. Neither Claude nor Codex, on their own, got it to a final state. There was always another bug. uzi got it done, and I think that is because of its planning. So I want to keep using uzi for work like that. But I cannot wait on its long execution times.
Phased tickets are the compromise I have right now. The plan gets the careful treatment, on Opus, and the execution is the cheap part. That is uzi’s plan checkpoint with the model tiers flipped: uzi runs its coder on Opus.
It is too early to say much about the model mix. The first count, on 29 September, found 10 dispatcher starts in the previous seven days that recorded a model: 1 on Opus and 9 on Sonnet. The other 35 came before the model was recorded at all.
Ticket dependencies
On 2 October tickets got dependencies. A ticket names what it needs in one line, Depends on: #N, and the dispatcher picks up the dependencies first and only executes a ticket once they are closed. If a plan was built on another plan that has since changed, it goes back to a planner.
The feature went through the phased process itself, and it is the clearest picture of the week. I approved the plan at about seven in the evening. I merged the first pull request before its review ran. The review came back with changes, twice. The third execute round, the last one allowed, passed review. Then the QA audit failed it: the code did what it should, but two of the new tests passed no matter what the code did, and one contract had no test at all. With the round cap reached, the ticket stopped and handed the decision to me: merge as it is, or let an agent add the test fixes. I chose the test fixes, and both pull requests merged before ten, with three QA findings still open for a later ticket.
The cap worked the way it was designed to. It also meant Noj stopped on a question about test coverage until I answered it.
Noj plans my week now
A new skill plans my week: a priority list, a calendar from now on, and my hours against a 38-hour net week. It drafts in chat first and files only when I say so. Four decision records describe it, which makes it most of the week’s six. It is useful. It is also one more thing that produces drafts for me to approve.
What broke
A subagent testing a tool against the real board, instead of a fake one, put a label and a comment on the wrong issue. The session was not allowed to remove them, so the cleanup became a ticket for me. It sat there a week.
The circuit breaker from week 7 tripped for real. Just after midnight on 29 September the dispatcher failed three times in a row and paused itself until a person cleared it, which is the point of the breaker. It was resumed by hand that evening, about 21 hours later. The cause was a session that had left the shared checkout parked on a feature branch, so the dispatcher now runs from a clone nobody works in. That fix merged the morning after, and setting up the new clone and its timer was a step for me.
And I asked why every tool Noj has written is in Python. When I asked, there were 99 tools in Noj’s own repository, every one of them Python, and no decision record that ever chose it. The answer from Noj’s agents was that it just happened: Python is what a model reaches for when asked for a script, and nobody stopped to ask. I chose to move the tools to Go, one slice at a time. The first slice was also the first ticket to reach the round cap. It passed review, failed the QA audit on a gap in one test, and on 1 October came back to me. As I write this, it is still waiting for my call.
Toy apps where I stopped reading
For my personal projects, small apps that only I use, I gave Noj full ownership, without me reviewing. On the first app the records still show me writing “Plan approved” and clicking merge, but I did not review the plan or the code first. On the second app Noj merged its own pull requests.
Some things got merged that were not correct. On the night of 2 October a change to the first app passed Noj’s own review and a QA audit, and was merged and deployed. I tried it, and saving a closed date did nothing, with no message on the page. The fix was merged 33 minutes after the change itself. The next day part of that same change came out again, because once I used it, it was more complex than useful.
The next evening Noj built the second app overnight. I had given it permission in chat to plan, review, merge and deploy without waiting for me. It still stopped for me a few times: for the first commit, and for two merges its own permissions refused. Before the night was out, one fix had gone in for quotes that showed up doubled. By about seven the next morning three more had merged: a setting that failed with an error when switched on, an installed copy that kept running the old version after a deploy, and screens that were taller than the phone and scrolled.
Even so, seeing the problem live was much faster for me than checking a plan and reviewing the code would have been. The cost is small because these are toy apps and I am the only user.
This is an experiment, and I run it only on these toy apps. It is not the case for Kairos tickets: the 13 Kairos changes below were all merged by me. Having a space where I can try this is really helpful.
I am not proposing this for anyone else. I don’t have a better solution yet.
What the uzi pilot cost gets its own post.
The numbers
Between 26 September and 2 October Noj merged 13 pull requests to Kairos, 59 to its own repository, and 5 to my personal website and homelab. It wrote 6 architecture decision records and no postmortems. The internal board saw 127 issues opened and 89 closed. The full table is on the notes page.
Noj’s own repository went from 55 to 59. Kairos went from 17 to 13, counting only the agent account’s pull requests as before. Noj’s own repository counts every author: 51 pull requests by the agent account and 8 by mine.
The 13 Kairos changes, for the record:
- kairos-lab#37
- kairos-lab#40
- kairos-lab#41
- kairos-lab#45
- kairos-lab#53
- workshop-kubernetes-intro#12
- workshop-kubernetes-intro#13
- workshop-kubernetes-intro#14
- workshop-kubernetes-intro#15
- workshop-kubernetes-intro#17
- kairos-docs#713
- kairos-docs#717
- kairos#5068
As before: this post was drafted by the system it describes, and I reviewed, edited and merged it myself. If you want to follow along you can subscribe to the notes feed, or find me on LinkedIn or YouTube.