
Every large migration eventually forces a choice between two bad options.
The first is the mega-PR. Models like Fable or Astra can now produce tens of thousands of lines of code in under an hour, which makes merging massive projects in one big swing more tempting than ever. But massive PRs are near-impossible to review closely, can be a mess to rollback properly, and are not a good fit for a product like Ambrook that is a critical finance system for thousands of businesses across the country.
The second is the more classical, conservative approach: build new components or classes alongside the old ones and move call sites over gradually. Each PR is safe to review and merge, but the codebase lives in two systems for months at a time. While mid-migration, engineers and agents need to learn both APIs, and new code keeps being written against the old API, leading to delays, confusion, and headache.
| Short migration time | Long migration time | |
|---|---|---|
| Big diffs | The mega-PR: fast, unreviewable, with an all-or-nothing rollback | The stalled "V2" branch that never merges |
| Small diffs | Our goal | Incremental migration: safe per PR, slow overall |
This summer we migrated Ambrook’s design system from Material UI (MUI) to Base UI. After years of being well served by MUI, we had reached a breaking point, having overridden nearly all of MUI’s style defaults in favor of our own styles. Base UI’s headless, agent-friendly approach gave us the web accessibility best practices both libraries include without the drawbacks.
The migration was massive, spanning 54k line changes across 794 files, in a cross-platform design system that targets React web and React Native, and we completed it in just over two weeks without any major regressions or customer disruption.
To get this outcome, we wanted to break the false choice between mega-PR and drawn out migration, in favor a few project goals:
Rather than building as a mega-PR or one massive stack, we ran the migration as a graph. We combined long-lived planner agents, groups of cloud worker agents, thoughtful testing, human review, and rapid bug bash iteration in a process we coordinated to get the best of both worlds. Here’s how we broke it down:

Stacking our migration into a hundred small pieces wasn’t tenable because stacks are serial, with each change affecting the next downstream. Instead, we worked to simplify to create isolated, verifiable, parallelizable tasks that allowed us to break down the project into four phases. Each phase fanned out parallel worker agents, then brought their output back through human review, following the same basic steps:
We routed tasks to the appropriate coding model based on the kind of judgment they needed. Mechanical, well-specified batches (test writing, simple bug-fixes) went to Sonnet. Planning, dispatching, and the core migration, where one agent had to hold the whole system in context, went to Fable. Visual polish and complex component migrations went to Opus with a long context window. For review, we used multiple adversarial code review agents from different providers, so the reviewers didn’t share the authors‘ blind spots.
The first phase was prototyping, and most of it was meant to be thrown away. Agents produced exploratory mega-PRs that migrated large parts of the app at once, so we could see how it felt and where it broke. Alongside them, a handful of careful single-component PRs showed what “done right” looked like for one component.
From those, the plan made a decision for every component:
AmSpinner, AmSkeleton, and AmLink components were easier to write from scratch than use a library for.CalendarDate matches how we store dates.We also leaned on standard best practices for planning large migrations. In our design system, component props live in a .types.ts file shared with native, and those interfaces stayed frozen (with one planned exception, covered under Codemods), so that we didn’t run into conflicts between PRs while updating consumers. We also worked to avoid bundling any forward fixes that could clutter or complicate reviews.
Before we migrated the component implementation, we first filled in the gaps in our component tests. We worked to ensure that both component behavior and appearance was tested, so that our reviewers could more easily spot regressions in our migration PRs.
First, we expanded our Storybook harness. Prop galleries and auto-docs render each component in isolation, one prop at a time, but the regressions we were worried about would appear in a specific combination of props: a chip whose padding collapses inside a dense ledger row, a destructive menu item that loses focus styling, a form that misbehaves inside a dialog’s focus trap. We added dozens of stories, modeling the combinations of component props for each story based on the real-world use cases across our product, and used sanitized fixture data to ensure the cases were realistic.
Second, we ensured that every component had a suite of component tests for the behavior stories can’t see: keyboard navigation, focus, value parsing, and accessibility roles, all of which were provided by Material UI. A planner enumerated test cases per component and split them into eight batches, each a ticket referencing the shared plan. The resulting nine PRs added about 11,000 lines of tests (roughly 1,270 cases), giving future migration workers a spec to build against.
Because these test PRs were independent of each other, we didn’t need to stack them all up, and could merge them as they were ready, without any coordination overhead.
Most of the migration left component interfaces untouched and required the flexibility of an LLM to decide how to best adapt Material to Base UI. Icons, we realized, were the only exception. AmIcon, the component every icon in the app renders through, wasn’t easily compatible with Base UI without API change, which meant editing the hundreds of call sites across our app. That kind of change is formulaic, so a codemod fit it better than hand-written or agent-written edits, and we ran it before swapping any implementations to Base UI. Getting the one interface change out of the way first lets every later change keep true to our “don’t change the component API” rule.
The change touched about over 600 files, with eight batches of roughly 100 files ran in parallel. Once each PR passed CI, we cherry-picked all eight onto a throwaway branch and tested the combined result for visual regressions. The batches were then reviewed separately, stacked, and merged together as a single squashed commit.
What was left after merging tests and codemods was the actual migration of our design system components from MUI to Base UI. Our goal was to merge a single stack of PRs in one go.
We forked this phase out like the others. A Fable session planned the migration, using example PRs we had carefully reviewed beforehand as a template, fanning out ten Niteshift agents by component family (menus, dialogs, selects, popovers, text fields, and so on), each building its part of the stack with one commit per component.
With the stack assembled, we deployed it to a staging link and bug bashed the whole app on the new foundation. While bashing and fixing and bashing and fixing could have taken weeks, with Supercut, we cut this down to just two days end-to-end with a far simpler process.
In parallel with the visual bashing by humans, we put the stack through its paces through a number of other checks and reviews:
Then, just three days after opening this stack, we merged the rebuilt components into main.
Removing MUI was our first milestone, but after merging, we set to work shipping enhancements that had been too awkward or complex to add before we migrated to Base UI. In the weeks since launch, we’ve shipped dozens of improvements to our design system:
AmMenu into components like AmMenu.Root, AmMenu.Trigger, AmMenu.Item, and so on.The cleanup mattered just as much for engineers. We removed style overrides written against .Mui* class names, aligned on a consistent dataTestId contract on web and native, deleted unused components, and added lint rules to prevent backsliding.
Two weeks after the first commit, our new design system launched into production to very little external fanfare. Across the project, we merged 60 PRs across over 120 agent sessions, modified 794 files and changed 54,000 lines of code, a total of 4.5B total tokens. When it comes to migrations like these, no news is good news – we had replaced MUI without any major regressions.
| Phase | Sessions | Input | Output | Cache writes | Cache reads | Total Tokens |
|---|---|---|---|---|---|---|
| Planning | 15 | 0.6M | 2.1M | 21.0M | 370.6M | 394.3M |
| Tests | 21 | 1.9M | 3.1M | 29.2M | 575.3M | 609.5M |
| Codemods | 13 | 0.8M | 1.8M | 28.8M | 672.5M | 703.9M |
| Component Migrations | 48 | 6.8M | 8.3M | 81.1M | 2,040.0M | 2,136.2M |
| Fast Follows | 24 | 2.0M | 2.3M | 28.4M | 661.0M | 693.7M |
| Total | 121 | 12.1M | 17.7M | 188.5M | 4,319.4M | 4,537.6M |
We’re using agent graphs to ship large features, tech debt cleanup, and other larger migrations across Ambrook; we no longer have to choose between a mega-PR and a slow, prolonged migration. Using graphs and parallel PRs that factor out dependencies, we can merge code faster and more confidently.
Human review is a critical part of moving quickly and confidently, especially when building accounting and payments software. By designing our graphs to optimize for reviewability, we built the confidence we needed to make sweeping, ambitious changes.
Ambrook models complex financial worlds.

Build an exact model of an inexact world.
