
Every engineering team feels the pull between two kinds of work. One is the steady stream of fixes and quality improvements that keep a product polished and resilient. The other is the large, ambitious features that open up new business opportunities.
With this tension in mind, Ambrook’s Product Pod set out to automate as much of the former as possible so we can focus more engineering bandwidth on the latter. Our goal was simple: increase the throughput of small and medium sized improvements without spending more time on them. We started this project during Momentum Month, where the entire company spent six weeks building agents for every function in the company. In the end, we built a “software factory” that allows operators (i.e., not engineers) to go from product problem definition to a merged PR in less than an hour, and yielded a 5x faster time-to-fix for bugs.
Our automation pipeline consists of four major components:
Linear was the natural choice for our state, because it already holds the context of the feature requests and bugs that the agents will act on. We built a simple set of labels to identify each ticket’s state, and used label-driven webhook events to trigger agents via Zapier[0] which update labels in Linear after completing their work.

[0]: There are other ways of reacting to webhooks events. We introduced Zapier as an operator-friendly if-X-then-Y interface but will soon move these production loops to something more easily configured via code.
The state machine begins when a ticket receives an Agent Triage label. This will trigger a Linear webhook that proxies through our application to Zapier with the ticket information as well as an idempotency key to deduplicate requests. The Zap will then split paths, separating tickets with a “Bug” label from feature request tickets which do not have that label.
Each path sends a request to our homegrown cloud agent, Amos. Anyone at Ambrook can interact with Amos in Slack or use it programmatically for automation. Amos has a shared set of skills that help it triage our tickets, root causing them, sizing the fix, flagging scope creep, and suggesting priority.
Amos will investigate the ticket, post a Linear comment with its findings, and label the ticket with its classification and reasoning.
Tickets are then sized to qualify or disqualify them for agentic implementation. Features are given a t-shirt size (small, medium, and large), with large features being set aside for an engineer to own.
Bugs are designated into three categories:
Agent Fixable when the fix is well-defined and an agent wouldn’t reasonably find any alternative. Examples might include UI issues like a button not opening the correct dialog, a cache not busting and leading to UI inaccuracies, or an off-by-one invoicing date issue caused by timezones.Needs Engineer when the fix is trickier, often requiring in-depth knowledge of our accounting or payments systems. They could also be just not easily traceable by the triaging agent, or with a fix complex enough that we don’t have confidence the agent will get it right on its own. This is a dynamic set as the set of things models can do on their own is always increasing.Not a Bug when the behavior was intended, or the request is really a new feature.Tickets labeled Not a Bug, labeled Needs Engineer, or sized as a large feature request exit the loop. Humans take over to either re-triage, rewrite the ticket, or close it.
Any Agent Fixable tickets will trigger Niteshift, our model-independent cloud coding agent. Once a PR is up, Niteshift will iterate with one of several automated review agents to get the PR into a healthy shape. When this process is complete, Amos leaves a comment on the original ticket letting the ticket owner know a PR is ready and requests them to preview the fix on a staging deployment.
If required, Amos can also request input from the ticket assignee or creator and will include instructions on how to send the ticket back to the triage phase with updated context.
We have a team-wide policy that the creator of a ticket shepherds it through automation. An operator who files a bug is responsible for answering Amos when it comes back with “Needs Info,” and for testing the fix on the preview link once a PR is up. Niteshift includes a short manual-testing path in every PR description to make that quick. If the fix is wrong, the operator iterates with Niteshift or pulls in an engineer.
Wherever possible we lean on automated code review to shorten review times and raise our confidence in a fix. We use Greptile, Claude code review, and run a home-grown agentic risk assessment that flags large diffs, changes to sensitive systems like accounting or authentication, and anything with undefined behavior. Today, though, nothing ships through the loop without a human looking at it.
Which human a PR goes to depends on how the ticket was classified:
Agent Fixable bug PRs go to an engineering review rotation shared across the whole team.Needs Engineer bug tickets go to a separate, dedicated set of engineers.The goal of this split is to give engineers confidence that we are shipping high-quality code without cluttering the standard review request lane with auto-generated PRs.
We use two sources of data to improve the loop over time: OTEL sessions from the individual agent runs and GitHub PRs. OTEL session data gives us details on where agents are inefficient: excessive tool calls, errors, duration. GitHub PRs (in particular, the reviews left by both automations and humans) give us insight into what the agent actually produced. Three times a week we run another automation that reviews OTEL session data and reviews to suggest automated improvements to agent instructions, tools, and so forth.
We also maintain several views of the loop in Hex and Linear in order to identify places where the tickets are getting stuck. We surface metrics like time-to-fix, time-to-merge, merges per week, human touches in the pipeline, and how long tickets are remaining stuck when humans are intended to intervene.

Today, if a customer success operator is on a customer call and notices that some timeline events are off by a day (somewhere in the app, there’s a timezone bug), they can file a ticket, add the agent triage label, and move on to their next call. Meanwhile Amos triages the ticket, Niteshift opens a PR, and automated systems review that PR. The rep tests the fix on the preview link, an engineer merges, and the rep emails the customer back the same afternoon.
That speed of turnaround back to the customer is paramount. We’ve seen the median time to completion for bugs drop 80% since the beginning of the year, with a time to close under 24 hours and 1 in 4 bugs being resolved in under three hours.
We’ve also seen our teammates in sales, CX, and marketing make tons of small improvements to the app, often shipping improvements that could have otherwise been easily lost in the backlog. Just as importantly, the Ambrook’s engineering team is not underwater trying to fight our backlog. Instead, we’ve carefully engineered a system that ships the backlog.
In the end, we didn’t build this to ship more code. We built it so the small fixes and improvements don’t need to compete with our more ambitious work, because both are important in building the best product possible for our customers.
I’ve always been more interested in the application of software than software itself.

Build an exact model of an inexact world.
