finnborge16 hours ago
This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.
That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.
Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.
Thank you for sharing and the care you put into writing this!
ajstorm15 hours ago
Thanks for the thoughtful response. We too were surprised by how deeply we could take the model of the teaching hospital to software engineering. Whenever we think to expand the system in one dimension or another, the teaching hospital model seems to have a nearby analogy.
I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.
zavec3 hours ago
I'll be very interested to read any followups about the part you described at the end about getting it to do more "teaching," both from the perspective of how we can use this to help newer devs out but also from the perspective of helping ourselves understand the software being built.
One thing I've been worrying about recently is the notion of comprehension debt and making sure the humans can still understand the system (so as to be able to do incident response or something if the LLMs are down). The slower, more methodical workflow described here should definitely help by allowing checks like proper docs whenever a piece of the architecture changes for instance, but I'm curious what else could be done to help the humans understand as much of the implications of what the AI is doing as possible.
Edit: having the code author and reviewer be completely separate like you described below is also probably pretty helpful for making sure changes are comprehensible from the outside.
jacquesm14 hours ago
Now imagine what it is like to run an actual hospital with real people whose lives (or the lives of their loved ones) are often at stake, with underpaid and overworked staff and with messy biological creatures as the subjects instead of bits and bytes. If anything this whole exercise should also give you a much deeper appreciation of the people that feel themselves called to help others.
jaggederest14 hours ago
I had a similarly useful analogy of the legal system emerge, in a similar way. I think there's a great analogy between policy in the legal system and policy in software engineering, and they have some awfully good (and very, very historical) ways to think about e.g. amending, repealing, and adjudicating things based on those policies.
baddash13 hours ago
i think role-play and utilizing the full power of language, stories, character, and narrative will unlock very sophisticated use-cases and in general a new dimension to agentic systems much in the way you're describing.
probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and/or texts in general? not really sure.
exe3410 hours ago
I feel like this is something we do with humans as well. Things like "scrum", "sprint", "the clean coder", etc.
slopinthebag13 hours ago
yeah this is really cool. i've been thinking about sort of a mirrored idea, where you end up with wizards, mages, clerics etc and fantasy terminology used. idk if it would be as effective as this is, but maybe it would be more creative somehow?
i kinda want to try to build this off of github. it's essentially just an event bus / message queue that workers (agents) tap into.
anonymous90821312 hours ago
I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.
zshrdlu10 hours ago
Seems to me it's just research and experimentation.
gausswho3 hours ago
The next role: Insurance Rep.
Ensures all other agents are operating efficiently and within reasonable levels of token usage given the expected level of effort to 'resolve' the patient. In moderate to severe cases, may lead to patient defenestration or agent revolt.
Or perhaps, the author considered this role but found it typically costs more than it saves.
nickpeterson2 hours ago
“patient defenestration” certainly conjures an image.
james_marks17 hours ago
A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
devin16 hours ago
I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I've seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.
Zanfa10 hours ago
I just finished ripping out one of these “dark factory” setups. Removed about 70k lines of code and 750k words of generated documentation. For what is effectively a 5-screen app.
Majromax2 hours ago
Funny that we've automated meeting hell. Definitely a truism that we reimplement the org chart in software.
supermdguy14 hours ago
I've gone through this cycle recently. I think static lint/type checks are super useful, but agentic code review loops can easily go off the rails.
zx808012 hours ago
> not helping
It does! It helps getting promotion with tokenmaxxing.
devin11 hours ago
lol, no doubt.
gchamonlive16 hours ago
This is Microsoft scale, I'd be surprised if it made any difference these labs. It's more likely it's plain managerial mishandling of the infra in chasing new profit heights.
ajstorm2 days ago
Rafi and I, who authored this post, will be hanging out here for any questions people may have.
bethekidyouwant3 minutes ago
How much of this is a snake eating its own tail against vs the control group?
Eridrus16 hours ago
Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
ajstorm15 hours ago
No, we haven't performed any ablation studies yet - it's a good suggestion.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.
Majromaxan hour ago
> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.
Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.
It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.
devghost_pro14 hours ago
[flagged]
ArtRichards7 hours ago
Yes agreed, using a system very similar without the hospital vibe :)
dirtbag__dad3 hours ago
This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.
1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?
2. Sorry if I missed this in the post, but will you open source this system?
losteric12 hours ago
are any of these artifacts public and available for inspection/use?
ajstorm12 hours ago
Not currently, but open sourcing this has been discussed. Stay tuned.
what13 hours ago
The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?
contingencies17 hours ago
Which inherent limitations did you recognize in the metaphor before commencing this research?
ajstorm15 hours ago
One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
boparaji307 hours ago
[flagged]
reachableceo6 hours ago
I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?
I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.
The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.
virgilpan hour ago
Why are github issues not a proper ticket system? E.g. rust uses it, seems to work fine for them...
davidmurdoch5 hours ago
Who defines what proper is?
ajstorm3 hours ago
The simple answer is convenience. It was just easiest to build it this way when we started.
K0balt10 hours ago
I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?
kinduff16 hours ago
Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.
ajstorm15 hours ago
I'd say that the first "mistake" we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn't human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.
Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.
When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.
There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.
dkubb9 hours ago
> When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped.
Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?
This way if a single requirement is dropped the system alerts you rather than it being silently dropped.
It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).
aetherspawn14 hours ago
Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.
ajstorm12 hours ago
It often does that too, but we keep the ledger and the code comment as well, to ensure that if it doesn't get resolved by the sibling, that it's not lost. If the sibling does resolve it, it's removed from the ledger and the code.
keyofthedooran hour ago
[dead]
cs7025 hours ago
Great post. Thank you for sharing it on HN.
Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.
In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.
It's still incredible. We sure live in interesting times!
rzzzt4 hours ago
You could also run them like an army, like an artist commune, like Valve (add virtual tables on wheels which they can pull wherever they want), like a lawyer's office, etc.
jebarker2 hours ago
If the correct analogy for a team of SW agents isn’t a team of SW engineers why is that?
virgilpan hour ago
For example, we don't have a clear analogue for "triage nurse". What would be it? Oncall engineer for production issues; "product owner" maybe for the regular tickets?
Thing is, software engineering is still (relatively) young. And more of a craft than true engineering. So the "good practices" are _somewhat established_, but at the same time, not quite. Like, we have glimpses into what works and what doesn't, but not a universally-established workflow to follow. And AI is upending what we already knew; it's really not surprising that people are looking elsewhere for workable metaphors.
jebarkeran hour ago
> For example, we don't have a clear analogue for "triage nurse". What would be it?
Triage nurse seems well defined in software engineering, perhaps just without a single title. There are certainly engineers that focus on triaging GitHub issues and then fixing immediate showstopper bugs and/or prioritizing the issue for further attention from other engineers.
That role can be communicated to agents by writing down that description. The risk with saying that the role is “triage nurse” is that you can’t be sure what other roles and allowances the agent will infer it has.
mimischi9 hours ago
If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?
dingaling91116 hours ago
Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.
ajstorm15 hours ago
It's true that we don't have long-term maintainability data just yet. We've just shipped the first product of this model to customers and likely won't have any detailed maintainability data for several months (and for good data, several years). We hope to author more blog posts on this experiment in the future.
Kinrany13 hours ago
Perhaps doing a random sample of the steps by hand will be a good way to notice maintainability issues?
arbor-group2 hours ago
[flagged]
[deleted]13 hours agocollapsed
mncharity13 hours ago
One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.
ajstorm13 hours ago
We actually have a /sinai-advocate skill, where a human can advocate on behalf of a stuck patient. We use it every once in a while when the labels get screwed up, or a workflow fails for some reason.
fathermarz11 hours ago
I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.
Fable feels like overkill for this also.
tujux6 hours ago
Humans: $160k / 9 months = $600/day
AI Software Factory: $4172 / 2 days = $2086/day
This seems unsustainable, unless you're also generating 3x the revenue.
paid_dot_expert6 hours ago
I don't get your maths? [1] 160/9 months != 600 for any given number of days a week e.g. 5,6,7 (something between 6 and 7), but where did the 9 come from, is it some kind of adjustment for weekends? Anyway there is also a concept of "fully loaded employee cost".
In any case the way to think of this is not "/day".
The reason is simple. If you buy 1000 barrels of oil, you buy 1000 barrels of oil not 17.4 days of oil. There is now a disconnect between work done and time. Infact you would be sane if you said "that result I can get in 2 days for $4172, if you can get me that same result in 1 hour, I'd pay $8344". See where this is going?
Yes a lot of thought work is now a commodity, and if you want the commodity faster (last minute booking, uber to come quicker etc.) you pay more not less. Value being $/hour is over.
[1] The ? acts as both a question and a regex.
thomascountz4 hours ago
I have qualms and disagreements with the whole "software factory" concept, but your math doesn't help the argument. The cost of human labor is more than salary—or even hourly rate, if you're thinking of paying a contractor.
epolanski6 hours ago
Only one of the two trends is downwards.
abuani4 hours ago
... Token cost is still heavily subsidized by the frontier labs while they're private companies and raise billions whenever they decide. The downward cost in tokens is unsustainable long term without revolutions training or inference of models.
Veelox16 hours ago
You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
ajstorm15 hours ago
All that credit goes to Rafi. I believe that he either noticed that they were getting long winded in a code review, or suspected that they needed trimming after reading this blog post from Anthropic: https://claude.dev/blog/the-new-rules-of-context-engineering....
It was definitely human driven, but I believe that the agents did the actual trimming.
Veelox15 hours ago
That is helpful. Thank you :)
jordanlewis2 days ago
Great post!
One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?
ajstorm2 days ago
That's a very interesting idea. Sounds more like a "routine follow-up" in the medical model.
amirkhanianan hour ago
[flagged]
sglim2 days ago
[dead]
Spooky2317 hours ago
Reminds me of the “surgical team” development model in the Mythical Man Month.
sroerick15 hours ago
I thought this too. It's fun that they were working with DB2 as well.
zmj15 hours ago
Nice writeup. Structured handoffs and external plan reviews are good takeaways.
ajstorm15 hours ago
Thanks! And thanks for reading.
[deleted]16 hours agocollapsed
gafferongames3 hours ago
This is fantastic work. I've been exploring my own work system here: https://github.com/mas-bandwidth/nova-sprint and I'm adopting your ideas. Thanks!
gafferongames3 hours ago
[flagged]
drc500free12 hours ago
I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.
git_rancher16 hours ago
The patient “leaves” when the bug is fixed?
tough16 hours ago
What would be the analogy if the patient dies?
unrented797716 hours ago
Bug becomes a feature
singularity200111 hours ago
congenially my agents started calling bugs gaps
actionfromafar2 hours ago
That's hilarious, tell us more!
pwdisswordfishq4 hours ago
"Just lost another bug."
"It never gets any easier, huh?"
pards4 hours ago
It brings a whole new definition to some medical terms like the "crash cart" used when a patient "codes" [0]
Trusteando8 minutes ago
[dead]
orbitaldesk2 hours ago
[flagged]
shledery2 hours ago
[flagged]
alyssassan6 hours ago
[dead]
alyssassan6 hours ago
[flagged]
khotem16 hours ago
[flagged]
ContinuityLab15 hours ago
[flagged]