Art byAdam Dixon
Artwork #01 / 01
Ambrook LogoEngineering
  1. AI

    Migrating fast with agent graphs

    By Dan Schlosser, Adam Markon

    Every large migration eventually forces a choice between two bad options.

    The first is the mega-PR. Models like Fable or Astra can now produce tens of thousands of lines of code in under an hour, which makes merging massive projects in one big swing more tempting than ever. But massive PRs are near-impossible to review closely, can be a mess to rollback properly, and are not a good fit for a product like Ambrook that is a critical finance system for thousands of businesses across the country.

    The second is the more classical, conservative approach: build new components or classes alongside the old ones and move call sites over gradually. Each PR is safe to review and merge, but the codebase lives in two systems for months at a time. While mid-migration, engineers and agents need to learn both APIs, and new code keeps being written against the old API, leading to delays, confusion, and headache.

    Short migration timeLong migration time
    Big diffsThe mega-PR: fast, unreviewable, with an all-or-nothing rollbackThe stalled "V2" branch that never merges
    Small diffsOur goalIncremental migration: safe per PR, slow overall

    This summer we migrated Ambrook’s design system from Material UI (MUI) to Base UI. After years of being well served by MUI, we had reached a breaking point, having overridden nearly all of MUI’s style defaults in favor of our own styles. Base UI’s headless, agent-friendly approach gave us the web accessibility best practices both libraries include without the drawbacks.

    The migration was massive, spanning 54k line changes across 794 files, in a cross-platform design system that targets React web and React Native, and we completed it in just over two weeks without any major regressions or customer disruption.

    To get this outcome, we wanted to break the false choice between mega-PR and drawn out migration, in favor a few project goals:

    • All changes should be packaged in small or simple diffs a human can actually review
    • No “overlap” period where the codebase has a confusing mix of old and new patterns
    • No user disruption or component regressions
    • No API churn for web or native consumers (the change should remain within the design system wherever possible)
    • Leave the system better tested at the end.

    Rather than building as a mega-PR or one massive stack, we ran the migration as a graph. We combined long-lived planner agents, groups of cloud worker agents, thoughtful testing, human review, and rapid bug bash iteration in a process we coordinated to get the best of both worlds. Here’s how we broke it down:

    Our Graph Design

    Stacking our migration into a hundred small pieces wasn’t tenable because stacks are serial, with each change affecting the next downstream. Instead, we worked to simplify to create isolated, verifiable, parallelizable tasks that allowed us to break down the project into four phases. Each phase fanned out parallel worker agents, then brought their output back through human review, following the same basic steps:

    1. A long-lived Fable planner agent running in Claude Code held the context for the migration, writing and proposing a plan.
    2. A human reviews the plan, giving feedback on prototype PRs and the proposed batching to minimize conflicts.
    3. The planner writes a brief per batch and files it as a Linear ticket assigned to Niteshift, one of our cloud-based coding agents, running either Opus or Sonnet, depending on the ticket complexity. Niteshift allows these tickets to be addressed by an independent cloud agent with its own environment, branch, and context window, not as a sub-agent inside the planner’s session, so one crashed or rate-limited session can’t take the batch down with it.
    4. Each Niteshift agent returns a separately mergeable PR, incorporating agentic code review feedback.
    5. A human reviews and merges, either independently or as a stack (depending on the phase).

    We routed tasks to the appropriate coding model based on the kind of judgment they needed. Mechanical, well-specified batches (test writing, simple bug-fixes) went to Sonnet. Planning, dispatching, and the core migration, where one agent had to hold the whole system in context, went to Fable. Visual polish and complex component migrations went to Opus with a long context window. For review, we used multiple adversarial code review agents from different providers, so the reviewers didn’t share the authors‘ blind spots.

    1. Planning & Prototyping

    The first phase was prototyping, and most of it was meant to be thrown away. Agents produced exploratory mega-PRs that migrated large parts of the app at once, so we could see how it felt and where it broke. Alongside them, a handful of careful single-component PRs showed what “done right” looked like for one component.

    From those, the plan made a decision for every component:

    • Base UI, where it had the right primitive. Tooltips, menus, dialogs, selects, and popovers all moved onto Base UI’s headless components.
    • Plain JSX and CSS, where MUI had been doing little. Our AmSpinner, AmSkeleton, and AmLink components were easier to write from scratch than use a library for.
    • A dedicated library, where Base UI had nothing to offer. Calendars and date pickers specifically aren’t provided by Base UI, so we used React Aria, whose timezone-less CalendarDate matches how we store dates.

    We also leaned on standard best practices for planning large migrations. In our design system, component props live in a .types.ts file shared with native, and those interfaces stayed frozen (with one planned exception, covered under Codemods), so that we didn’t run into conflicts between PRs while updating consumers. We also worked to avoid bundling any forward fixes that could clutter or complicate reviews.

    2. Tests

    Before we migrated the component implementation, we first filled in the gaps in our component tests. We worked to ensure that both component behavior and appearance was tested, so that our reviewers could more easily spot regressions in our migration PRs.

    First, we expanded our Storybook harness. Prop galleries and auto-docs render each component in isolation, one prop at a time, but the regressions we were worried about would appear in a specific combination of props: a chip whose padding collapses inside a dense ledger row, a destructive menu item that loses focus styling, a form that misbehaves inside a dialog’s focus trap. We added dozens of stories, modeling the combinations of component props for each story based on the real-world use cases across our product, and used sanitized fixture data to ensure the cases were realistic.

    Second, we ensured that every component had a suite of component tests for the behavior stories can’t see: keyboard navigation, focus, value parsing, and accessibility roles, all of which were provided by Material UI. A planner enumerated test cases per component and split them into eight batches, each a ticket referencing the shared plan. The resulting nine PRs added about 11,000 lines of tests (roughly 1,270 cases), giving future migration workers a spec to build against.

    Because these test PRs were independent of each other, we didn’t need to stack them all up, and could merge them as they were ready, without any coordination overhead.

    3. Codemods

    Most of the migration left component interfaces untouched and required the flexibility of an LLM to decide how to best adapt Material to Base UI. Icons, we realized, were the only exception. AmIcon, the component every icon in the app renders through, wasn’t easily compatible with Base UI without API change, which meant editing the hundreds of call sites across our app. That kind of change is formulaic, so a codemod fit it better than hand-written or agent-written edits, and we ran it before swapping any implementations to Base UI. Getting the one interface change out of the way first lets every later change keep true to our “don’t change the component API” rule.

    The change touched about over 600 files, with eight batches of roughly 100 files ran in parallel. Once each PR passed CI, we cherry-picked all eight onto a throwaway branch and tested the combined result for visual regressions. The batches were then reviewed separately, stacked, and merged together as a single squashed commit.

    4. Component Migrations

    What was left after merging tests and codemods was the actual migration of our design system components from MUI to Base UI. Our goal was to merge a single stack of PRs in one go.

    We forked this phase out like the others. A Fable session planned the migration, using example PRs we had carefully reviewed beforehand as a template, fanning out ten Niteshift agents by component family (menus, dialogs, selects, popovers, text fields, and so on), each building its part of the stack with one commit per component.

    With the stack assembled, we deployed it to a staging link and bug bashed the whole app on the new foundation. While bashing and fixing and bashing and fixing could have taken weeks, with Supercut, we cut this down to just two days end-to-end with a far simpler process.

    1. Designers, engineers, and other Ambrook team members all recorded themselves bashing the app on real devices, narrating issues they saw along the way. Instead of having to type or capture perfect bug reproductions, Supercut allowed them to just record one hour-long session, vocalizing bugs naturally: “this chip’s padding collapses,” “make this button one size bigger,” etc.
    2. We passed all the Supercut URLs to a single Fable session, which used the MCP to read the transcript, identify issues, merge duplicates identified by multiple people, and filing Linear tickets for each. The Supercut MCP allows agents to get a screengrab at a particular timestamp, which meant that the tickets were rich with context, ready to be picked up.
    3. After a quick review, we sent all these tickets to their own Niteshift agents in parallel, stacking fixes onto our main migration branch. A quick review on our stack’s staging link confirmed the fixes.

    In parallel with the visual bashing by humans, we put the stack through its paces through a number of other checks and reviews:

    • Our adversarial code review agents checked for accessibility, keyboard, focus, contract, styling, effect leaks and worked with Niteshift to automatically fix detected issues.
    • We sent Fable to smoke test our most critical pages, using browser-use to visually diff production and the staging link.
    • The tests introduced earlier (storybook and component tests) ran and flagged issues to Niteshift until they were all green.
    • Humans reviewed every PR in the stack.

    Then, just three days after opening this stack, we merged the rebuilt components into main.

    5. Fast Follows

    Removing MUI was our first milestone, but after merging, we set to work shipping enhancements that had been too awkward or complex to add before we migrated to Base UI. In the weeks since launch, we’ve shipped dozens of improvements to our design system:

    • Accessibility. We enhanced the accessibility of dozens of components, from accordion triggers and tabs to selects and autocompletes. Nested menus now work entirely from the keyboard, and every control shares one consistent set of visible focus ring styles.
    • Motion. A motion token system replaced MUI’s JavaScript transitions, with popups scaling and fading from their origin, optimized drawers’ sliding animations, and reduced-motion preferences respected everywhere.
    • More powerful date pickers. We shipped a brand new set of calendar and date picker components built on React Aria, enabling us to understand dates entered in shorthand formats.
    • Design token coverage. We moved all hardcoded colors to tokens, menus, selects, and autocompletes share one compact treatment, and popups no longer open behind dialogs.
    • Performance. We reduced our bundle size by removing MUI and Emotion, moved transitions to run on compositor-friendly properties, and removed unnecessary re-renders.
    • Component API Changes. We intentionally avoided changing our component APIs to move towards Base UI’s preference for composition over configuration (the opposite of Material UI’s approach) during the migration itself. We are now decomposing large components like AmMenu into components like AmMenu.Root, AmMenu.Trigger, AmMenu.Item, and so on.

    The cleanup mattered just as much for engineers. We removed style overrides written against .Mui* class names, aligned on a consistent dataTestId contract on web and native, deleted unused components, and added lint rules to prevent backsliding.

    Silence is Golden

    Two weeks after the first commit, our new design system launched into production to very little external fanfare. Across the project, we merged 60 PRs across over 120 agent sessions, modified 794 files and changed 54,000 lines of code, a total of 4.5B total tokens. When it comes to migrations like these, no news is good news – we had replaced MUI without any major regressions.

    PhaseSessionsInputOutputCache writesCache readsTotal Tokens
    Planning150.6M2.1M21.0M370.6M394.3M
    Tests211.9M3.1M29.2M575.3M609.5M
    Codemods130.8M1.8M28.8M672.5M703.9M
    Component Migrations486.8M8.3M81.1M2,040.0M2,136.2M
    Fast Follows242.0M2.3M28.4M661.0M693.7M
    Total12112.1M17.7M188.5M4,319.4M4,537.6M

    We’re using agent graphs to ship large features, tech debt cleanup, and other larger migrations across Ambrook; we no longer have to choose between a mega-PR and a slow, prolonged migration. Using graphs and parallel PRs that factor out dependencies, we can merge code faster and more confidently.

    Human review is a critical part of moving quickly and confidently, especially when building accounting and payments software. By designing our graphs to optimize for reviewability, we built the confidence we needed to make sweeping, ambitious changes.

  2. Next
    Infrastructure

    Optimized Real-time Firestore Changelogs with BigQuery

    Ambrook models complex financial worlds.

Build an exact model of an inexact world.

Art byAdam Dixon
Artwork #01 / 01