Testing an open source software factory: uzi, by Vlad Mocanu
Our first try at uzi, Vlad Mocanu's dark factory, on real Kairos work: a shared-networking feature that landed upstream after about $340 of reported Claude calls, two runs that lost finished work, and a default model choice that took me two days to notice.
This is our first try at using uzi, a dark factory by Vlad Mocanu. A dark factory plans and runs the work without a person watching it. Noj and I pointed it at real Kairos work at the end of September. I say “our” on purpose: I couldn’t have played with uzi as fast as I did if Noj, my own agent setup, hadn’t done the hard work.
Please don’t read this as a verdict on uzi. It was our first attempt, and we probably didn’t do things the best way. I share our mistakes too, because I find that more useful than a polished story.
Two things first. The developer is approachable and open to working together. When I reported a version-stamp bug on 25 September (uzi#1682), Vlad replied the next morning with the root cause, took the design I proposed over his own, and shipped it in v0.84.0 that same day. When we filed four feature requests on 30 September, he answered all of them on 2 October with a draft plan, and on two of them he offered that we implement the change and he reviews it (uzi#1983, uzi#1984).
And I got a feature that mattered for my kairos-lab project: shared networking, merged upstream on 27 September (kairos-lab#34). Getting there was bumpy.
The setup, and a note on the numbers
I deployed uzi on my homelab Kubernetes cluster and pointed it at mauro-agent/kairos-lab, a fork of kairos-lab. uzi needs write access to the repository it works on, so it works on the fork, and I open the pull request upstream by hand. Every run here used uzi v0.83.1. Where a later release fixes something, I point to the fix.
uzi is open source and costs nothing. Every dollar figure below is the amount each Claude call reported back. I run on a Claude subscription, so I did not pay these amounts. I use them to measure how much work the pilot asked of the model, not what uzi costs.
The shared-networking feature (NAT networking, a MAC address per VM, address discovery) used about $340 worth of calls from planning to merge, failed runs included. A follow-on run used $107.28 and left nothing behind. On the first day, two runs used about $80 over about nine hours, and neither produced something I could build and run.
Where the money goes
The planning phase for that one feature used $12.01: 162k fresh input tokens, 126k output, and 11.2M cache reads. Several agents each read the same issue and the same repository, so the amount follows how many agents read the context, not how much code gets written. uzi’s per-run figure matched Anthropic’s console once I took out unrelated usage.
The bigger lever was the model. On one run I counted the model field on every agent message: 2,537 of 2,580 (98.3%) were Opus. Only the documenter role ran on Sonnet. The coder alone sent 1,120 of those messages. When I counted, that run had reported $185.72, and it finished at $232.71. The six-milestone run the day before, on the same repository, reported $71.60 before it was killed. Smaller scope, more than three times the amount.
uzi lets you override the model per agent, and an admin can edit the built-in roles, so this is about the default, not a missing feature. I just didn’t know to look, and it took two days and about $270 worth of calls to notice. Later releases come with better defaults, though most agents still run on Opus. Whether that is right is up to you. There’s no right answer, and a lot of it depends on your budget.
Where the time goes
On v0.83.1 a run’s budget is time, not money: a wall-clock limit that defaults to 8 hours, whatever the scope. A six-milestone run and a five-milestone subset of it both got 8 hours, and only the iteration count changed (30 against 25). So splitting a big issue into smaller ones multiplied the timeline instead of dividing it. Time also follows the review policy the plan picks for itself: a small diff that the plan classifies as risky gets a reviewer and a tester on every commit, and that is where the hours went.
I’d like a cost budget next to the time budget. But the bigger question is what happens at the limit. A run of half a day or a full day has done real work by the time it reaches it. A system like this should be able to weigh that work against starting over, and decide whether a little more time is worth it or a hard stop is better, within limits the owner sets, if any. v0.84.0 moves in that direction: a run at its limit now parks and asks the owner to extend, stop or cancel, instead of failing (#1504).
Where the work went
A uzi agent is denied git push at the tool layer. It commits locally, and only the worker’s finalization step publishes anything. An agent said so itself, when I asked it to push after each milestone:
I can commit locally after each milestone, and will, but
git pushis denied to me at the tool layer; the worker opens the merge request after I signal done.
So in our setup, a run that died before finalizing left everything on the worker’s disk. That happened twice on the first day.
The first run was OOM-killed with four milestones committed. uzi noticed the orphaned run and opened custody holds for it, but there was no archive to download, and the UI’s only action was to discard. I got the work back by going into the worker pod and pulling a git bundle off the volume by hand. v0.84.0 partly addresses this: a checkpoint now publishes committed work to the remote on a timer, best effort, so at the default settings at most about 25 minutes of committed work is at risk (#1618).
My worker had a 4 GiB memory limit, which I then raised to 8 GiB. We never established what used the memory. It could have been uzi, or it could have been the code under test, since kairos-lab starts virtual machines. Here the limit hurt, but a hard cap like that can be a useful kill switch, depending on the work. Then the factory needs a way to get feedback from the platform about why the run died. I didn’t expect that from uzi, but it would be a good feature for a later release.
The second run picked up a ticket that was already half done, so we gave it 4 hours instead of the default 8. It hit that limit and was killed two seconds over. Its log shows it had finished: a clean tree at commit d407ca5, with the reviewer and auditor in their final passes. The worktree was cleaned up, and that commit no longer exists anywhere on the worker. uzi’s custody record knew the SHA and said bundle_failed. Both parts are addressed in v0.84.0: the run would now park and ask instead of failing (#1504), and the commit now goes into the worker’s own bare repository before the bundle is built, so cleanup no longer deletes the only copy (#1510).
Milestones 1 to 4 of that feature exist because we recovered them by hand. Milestones 5 and 6 were written, lost, and written again by later runs. That is part of the $340.
A later run on the multi-VM follow-up (mauro-agent/kairos-lab#6) timed out after milestone 1 of 5: $107.28, about 8.5 hours, nothing recoverable. Its plan was good and already approved, but uzi had no way to start a new run from another run’s plan. The only route was copying plan_md out by hand and feeding it back through --plan-file, as if I had written it.
What we built around it
Noj now has a small snapshot tool that copies a live run’s unpublished commits off the worker before they can be lost. It only reads from the cluster: it never pushes, deletes or touches the run. Testing it found two bugs in our own tool: a text-mode transfer that corrupted the binary bundle, and a shared marker that broke when the worker held more than one worktree.
That is a workaround, not a fix. When I talked to Vlad I suggested two directions. One is to let the agent push its own working branch on repositories the owner opts in, with upstream and downstream repositories treated differently, so the default stays strict where it should. The other is to make the internal capture reliable, so that “the agent cannot push” doesn’t also mean “the work is gone”.
What worked well
Our spec was written from a clone ten commits behind, and three of its claims about the code were false. During planning, uzi’s fact-checker checked the issue against the real HEAD, flagged those claims, and named the commit that had already landed the work, before any code was written.
The plan was better than the spec I gave it. It caught a data race my spec would have introduced, and it checked external facts against their sources instead of from memory. The review passes produced real security hardening. The defaults are more conservative than I expected: scheduled jobs and the run judge are both opt-in. And the findings page turned up 17 unrelated bugs along the way.
What I would tell another operator
Check which model each role uses before the first run, not after the totals come in. Prefer the gated path, where the plan is fact-checked and you approve it, over seeding a run with your own plan file: it is what caught our stale spec, and its time budget is sized to the plan, where a seeded run gets a fixed default. And until the push question is settled, snapshot live work off the worker on a timer.
What we filed
We filed eight issues on uzi, among them:
- uzi#1682, the version-stamp bug
- uzi#1983, a fork and pull request flow
- uzi#1984, an Anthropic token per repository
- uzi#1985, linking a handoff back to the run that failed
- uzi#1986, starting a run from another run’s approved plan
The number I can’t settle
A documentation pull request (mauro-agent/kairos-lab#7) used $26.61 worth of calls over about an hour and 25 minutes. I tried to compare that with what a person would cost, and couldn’t do it fairly. A person’s rate hides the employer’s overhead, and an agent’s per-run price hides the review time I still spend before I trust the output. I can see why companies will read the raw number as proof they need fewer people. I don’t think we understand yet what that shift does.
I’m still trying uzi. I see a good fit in running it next to Noj, and whether I keep it depends on how well it fits my workflow as the issues above play out.
This post was drafted by Noj. I reviewed, edited and merged it myself.