Button, 2021 to 2026
Commerce infrastructure at scale
I owned the routing, decisioning, attribution and order pipeline that every partner and every dollar moved through, at a startup where that pipeline was the product. Then my manager left, I inherited four functions with no transition plan, and the job became rebuilding the organization around the platform at the same time as running it.
- Commerce a month
- $3B+
- Orders a day
- 600K+
- Functions run at once
- Four
- Engineers
- 13, through leads as well as directly
15x
peak traffic scaled, no downtime
400 to 15
instances, on one service
30%
faster partner onboarding
15 min
to acknowledge a P0
Button sat between large publishers and large merchants and moved the commerce between them. My teams owned the part in the middle: the routing and decisioning that chose where a transaction went, the attribution that decided who got paid for it, and the order pipeline that carried it. When any of that was wrong, it was wrong about money, for partners including Amazon, Walmart, Uber and a long list of others.
That is a specific kind of engineering pressure. There is no degraded mode where the numbers are approximately right. So the reliability mandate and the tech debt mandate were the same mandate, and both sat with me.
Scale and reliability
01
Prime Day readiness: 80 to 1,200 requests a second
I led readiness for the year's largest retail traffic event, repeatedly, year over year. The work is unglamorous and it is all in the preparation: modeling what the traffic will actually look like rather than what the partner says it will, load testing to the model, finding what breaks first, fixing it, and then running a formal go or no-go rather than hoping. The platform held through its biggest days with nobody pulling an all-nighter to make it happen, which is the outcome I care about. A held peak that cost the team a weekend is a system that is still broken.
02
Right-sizing: 400 instances to 15, and the cost
Two services were provisioned for a peak they no longer had. I took one from 400 instances to 15 and another from 400 to 130, which cut peak event cost from roughly $20,000 a month to about $7,500. The reason this is worth telling is that it went the opposite direction to the instinct in the room. The conversation about a slow service defaults to adding capacity, and it takes measurement rather than opinion to establish that most of the fleet was doing nothing.
03
The order pipeline at 600,000 orders a day
Queueing through SQS, autoscaling in front of it, migrations onto Aurora with serverless database scaling underneath, and a release-confidence program so that shipping into the pipeline was a normal Tuesday rather than an event. The release work is the part that compounds: a team that is afraid to deploy accumulates risk in large batches, and then every deploy really is dangerous, which confirms the fear.
Getting the interrupts off the engineers
When I arrived, a partner with a problem sent a message to whoever they happened to know, and it landed on a platform engineer. There was no routing, no ownership, and no answer to how long anything was supposed to take. The visible cost was partner frustration. The invisible and larger cost was that the engineers who owned the most critical systems in the company were being interrupted all day.
So I built the function that did not exist. I made the case to leadership, secured the budget and the tooling, hired into it, and ran the migration from Salesforce to a Zendesk and Jira spine end to end: business case, vendor selection, procurement, and delivery. Zendesk became the source of truth for routing, prioritization and service levels. I defined the tiered agreements, the incident priority matrix and the escalation decision tree, with a P0 acknowledged inside fifteen minutes. Then I formed the Support Engineering team that owns it, and stood up an offshore arm through a vendor to take tier one and tier two load, so the function could scale without the headcount scaling with it.
Where a partner problem goes
Partners, product and revenue
01
Standardizing integration: 30% faster onboarding
Every new enterprise partner used to mean bespoke engineering, which meant the cost of a partner never came down and the team that could land one was always the same team. I standardized the enterprise integration path and removed the per-partner custom build, so a new partner lands on a repeatable process. Onboarding timelines came down 30% and revenue per integration went up 20%. The second number follows from the first: work that is repeatable is work you can price.
02
The partner integrations, and the security work
Amazon, Walmart, Best Buy, Target, Sam's Club, Nike, Lululemon, Expedia, Marriott, Uber One, Lyft, Disney+ and StubHub, including a full platform migration for Sam's Club. Enterprise delivery is also a security job: I worked infosec questionnaires through partner procurement and drove penetration test findings to remediation, which is unglamorous and is frequently the thing standing between a signed contract and a live integration.
03
Retail media: forming a new product line
I led the formation and initial delivery of the retail media product: set the roadmap and release plan with product and revenue, and shipped the first releases with NY Post, Forbes and BuzzFeed. Being early on a product line is mostly about deciding what not to build in the first six months.
Inheriting four functions
When the Senior Director above me left in 2024 I inherited the core platform engineering team, Infrastructure and Marketplace on top of Solutions Engineering. No transition plan, no defined scope. I ran all four for a year until the reporting lines caught up.
I split the work into owned lanes rather than routing all four through myself: architectural debt on one side, scaling on the other, each with a lead accountable for running their team against it. I reorganized the core platform engineering team into clear domain ownership across the order pipeline, the commerce data services and the shared platform foundations, and stood up a Leads forum so context moved sideways between teams rather than up through me and back down.
Alongside all of it I carried the delivery cycle for my org end to end: roadmap, backlog, capacity planning, quarterly planning and estimation, every sprint ceremony run by me, and a feasibility gate between revenue, product and engineering so the roadmap met reality before it met a customer. I also owned escalations and incident management company-wide, and took it from a backlog nobody worked to a routed process with a weekly remediation review.
Growing the people who took it over
01
Three of four functions run by leads I developed
That is the number I care about most on this page. I did not hand those functions to hires, I handed them to leads who were already there and were not yet running anything. Splitting the work into owned lanes was the mechanism: architectural debt on one side, scaling on the other, each with a lead accountable for running and operating their team against it, with the accountability genuinely attached rather than reserved for me. You cannot develop somebody into an owner by letting them advise.
02
Writing the growth ladder with the lead it was for
I wrote the Solutions Engineering growth ladder and the role definitions for Solutions Engineering and Support, and I built them with the lead I was moving into the role rather than handing them down finished. That was the point. Writing a ladder forces you to answer what actually separates one level from the next, and doing that work is how somebody learns to evaluate their own people. The artifact was the vehicle. The development was the deliverable. I also contributed to the software engineering ladder.
03
Defining Solutions Engineering before hiring into it
I established Solutions Engineering as a defined function: titles, role definitions, an intake and partner tiering model, and a hiring plan that doubled the team off a two-person breaking point. Then I handed operational ownership to the lead I had grown into it. Defining the function first is what makes the growth plan honest, because otherwise you are asking somebody to grow toward a target nobody has written down.
04
A Leads forum, so context moved between teams
I stood up a Leads forum so the people running teams talked to each other rather than routing every cross-team question up to me and back down. When the manager is the only connection between their leads, nothing moves sideways and every lead's development depends on one person's attention.
Stack
- Platform
- AWS (ECS, EC2, RDS and Aurora, S3, SQS, SES, Route 53), Go, Python, Node.js, Postgres and Aurora, MongoDB, BigQuery
- Operations
- Terraform, Docker, GitHub Actions, PagerDuty and on-call, Observability and traffic modeling, Blameless postmortems
- The support spine
- Zendesk, Jira, Salesforce, Tiered SLAs, Incident priority matrix