Skip to content
All case files

Button, 2021 to 2026

Commerce infrastructure at scale

I owned the routing, decisioning, attribution and order pipeline that every partner and every dollar moved through, at a startup where that pipeline was the product. Then my manager left, I inherited four functions with no transition plan, and the job became rebuilding the organization around the platform at the same time as running it.

Commerce a month
$3B+
Orders a day
600K+
Functions run at once
Four
Engineers
13, through leads as well as directly

15x

peak traffic scaled, no downtime

400 to 15

instances, on one service

30%

faster partner onboarding

15 min

to acknowledge a P0

Button sat between large publishers and large merchants and moved the commerce between them. My teams owned the part in the middle: the routing and decisioning that chose where a transaction went, the attribution that decided who got paid for it, and the order pipeline that carried it. When any of that was wrong, it was wrong about money, for partners including Amazon, Walmart, Uber and a long list of others.

That is a specific kind of engineering pressure. There is no degraded mode where the numbers are approximately right. So the reliability mandate and the tech debt mandate were the same mandate, and both sat with me.

Scale and reliability

  1. 01

    Prime Day readiness: 80 to 1,200 requests a second

    I led readiness for the year's largest retail traffic event, repeatedly, year over year. The work is unglamorous and it is all in the preparation: modeling what the traffic will actually look like rather than what the partner says it will, load testing to the model, finding what breaks first, fixing it, and then running a formal go or no-go rather than hoping. The platform held through its biggest days with nobody pulling an all-nighter to make it happen, which is the outcome I care about. A held peak that cost the team a weekend is a system that is still broken.

  2. 02

    Right-sizing: 400 instances to 15, and the cost

    Two services were provisioned for a peak they no longer had. I took one from 400 instances to 15 and another from 400 to 130, which cut peak event cost from roughly $20,000 a month to about $7,500. The reason this is worth telling is that it went the opposite direction to the instinct in the room. The conversation about a slow service defaults to adding capacity, and it takes measurement rather than opinion to establish that most of the fleet was doing nothing.

  3. 03

    The order pipeline at 600,000 orders a day

    Queueing through SQS, autoscaling in front of it, migrations onto Aurora with serverless database scaling underneath, and a release-confidence program so that shipping into the pipeline was a normal Tuesday rather than an event. The release work is the part that compounds: a team that is afraid to deploy accumulates risk in large batches, and then every deploy really is dangerous, which confirms the fear.

Getting the interrupts off the engineers

When I arrived, a partner with a problem sent a message to whoever they happened to know, and it landed on a platform engineer. There was no routing, no ownership, and no answer to how long anything was supposed to take. The visible cost was partner frustration. The invisible and larger cost was that the engineers who owned the most critical systems in the company were being interrupted all day.

So I built the function that did not exist. I made the case to leadership, secured the budget and the tooling, hired into it, and ran the migration from Salesforce to a Zendesk and Jira spine end to end: business case, vendor selection, procurement, and delivery. Zendesk became the source of truth for routing, prioritization and service levels. I defined the tiered agreements, the incident priority matrix and the escalation decision tree, with a P0 acknowledged inside fifteen minutes. Then I formed the Support Engineering team that owns it, and stood up an offshore arm through a vendor to take tier one and tier two load, so the function could scale without the headcount scaling with it.

Where a partner problem goes

most of itPartner reports aproblemAmazon, Uber, Best BuySam's Club, FetchTiered intakeRouting, priority matrix,SLAsAnswered outrightSelf-serve tooling, runbooksImpact analysisAbsorbed by the data teamPlatform engineeringOnly what needs a codechangeWeekly remediationreview
Before this existed, a partner with a problem sent a direct message to whoever they knew, and it landed on a platform engineer. Tiered intake now routes and prioritizes it, self-serve tooling answers the common questions outright, and the data team absorbs impact analysis. Engineering sees what genuinely needs a code change.

Partners, product and revenue

  1. 01

    Standardizing integration: 30% faster onboarding

    Every new enterprise partner used to mean bespoke engineering, which meant the cost of a partner never came down and the team that could land one was always the same team. I standardized the enterprise integration path and removed the per-partner custom build, so a new partner lands on a repeatable process. Onboarding timelines came down 30% and revenue per integration went up 20%. The second number follows from the first: work that is repeatable is work you can price.

  2. 02

    The partner integrations, and the security work

    Amazon, Walmart, Best Buy, Target, Sam's Club, Nike, Lululemon, Expedia, Marriott, Uber One, Lyft, Disney+ and StubHub, including a full platform migration for Sam's Club. Enterprise delivery is also a security job: I worked infosec questionnaires through partner procurement and drove penetration test findings to remediation, which is unglamorous and is frequently the thing standing between a signed contract and a live integration.

  3. 03

    Retail media: forming a new product line

    I led the formation and initial delivery of the retail media product: set the roadmap and release plan with product and revenue, and shipped the first releases with NY Post, Forbes and BuzzFeed. Being early on a product line is mostly about deciding what not to build in the first six months.

Inheriting four functions

When the Senior Director above me left in 2024 I inherited the core platform engineering team, Infrastructure and Marketplace on top of Solutions Engineering. No transition plan, no defined scope. I ran all four for a year until the reporting lines caught up.

I split the work into owned lanes rather than routing all four through myself: architectural debt on one side, scaling on the other, each with a lead accountable for running their team against it. I reorganized the core platform engineering team into clear domain ownership across the order pipeline, the commerce data services and the shared platform foundations, and stood up a Leads forum so context moved sideways between teams rather than up through me and back down.

Alongside all of it I carried the delivery cycle for my org end to end: roadmap, backlog, capacity planning, quarterly planning and estimation, every sprint ceremony run by me, and a feasibility gate between revenue, product and engineering so the roadmap met reality before it met a customer. I also owned escalations and incident management company-wide, and took it from a backlog nobody worked to a routed process with a weekly remediation review.

Growing the people who took it over

  1. 01

    Three of four functions run by leads I developed

    That is the number I care about most on this page. I did not hand those functions to hires, I handed them to leads who were already there and were not yet running anything. Splitting the work into owned lanes was the mechanism: architectural debt on one side, scaling on the other, each with a lead accountable for running and operating their team against it, with the accountability genuinely attached rather than reserved for me. You cannot develop somebody into an owner by letting them advise.

  2. 02

    Writing the growth ladder with the lead it was for

    I wrote the Solutions Engineering growth ladder and the role definitions for Solutions Engineering and Support, and I built them with the lead I was moving into the role rather than handing them down finished. That was the point. Writing a ladder forces you to answer what actually separates one level from the next, and doing that work is how somebody learns to evaluate their own people. The artifact was the vehicle. The development was the deliverable. I also contributed to the software engineering ladder.

  3. 03

    Defining Solutions Engineering before hiring into it

    I established Solutions Engineering as a defined function: titles, role definitions, an intake and partner tiering model, and a hiring plan that doubled the team off a two-person breaking point. Then I handed operational ownership to the lead I had grown into it. Defining the function first is what makes the growth plan honest, because otherwise you are asking somebody to grow toward a target nobody has written down.

  4. 04

    A Leads forum, so context moved between teams

    I stood up a Leads forum so the people running teams talked to each other rather than routing every cross-team question up to me and back down. When the manager is the only connection between their leads, nothing moves sideways and every lead's development depends on one person's attention.

Stack

Platform
AWS (ECS, EC2, RDS and Aurora, S3, SQS, SES, Route 53), Go, Python, Node.js, Postgres and Aurora, MongoDB, BigQuery
Operations
Terraform, Docker, GitHub Actions, PagerDuty and on-call, Observability and traffic modeling, Blameless postmortems
The support spine
Zendesk, Jira, Salesforce, Tiered SLAs, Incident priority matrix