Matthew Liste, Executive Vice President and Global Head of Infrastructure at American Express, believes that a successful technology platform is not measured by the number of its components, but by its ability to hide complexity and provide developers with a stable, clear experience. In his presentation at QCon San Francisco, Liste drew on more than 20 years spent building platforms and infrastructure for critical systems at Goldman Sachs, JPMorgan Chase, and American Express.
The platforms Liste discussed serve large numbers of internal developers: approximately 20,000 users at American Express and approximately 60,000 at JPMorgan Chase. Although their scale differs from that of cloud service providers, he believes the same fundamental principles apply to any team building a layer on which others depend.
A good platform hides complexity without hiding what is happening
Liste begins by defining a platform as an integrated set of technologies that forms a foundation for building applications on top of it. He compares it to water and sewage services: users do not think about the infrastructure behind them as long as they work, but they notice it immediately when it fails.
Therefore, the platform experience should be intuitive and use shared, interchangeable components, such as unified foundations for monitoring, identity, and namespaces. This composability makes the platform more like Lego blocks rather than a collection of separate services that work together only with additional effort.
Stability, security, and scalability are not optional features
Liste identifies what he calls the “three pillars”: stability, security, and scalability. The required level of availability varies according to the value of the system; according to his presentation, credit card authorization systems at American Express operate at levels reaching six nines, while other systems can tolerate a greater amount of downtime.
He also emphasizes that initial success can conceal scalability problems. A platform that works well at the beginning may encounter bottlenecks as its usage grows, directly affecting stability. For this reason, service-level objectives (SLOs) must be defined with consumers and honored over time, not only on launch day.
Continuous updating and reducing undifferentiated work
Liste considers keeping a platform up to date one of the most difficult and most postponable tasks. The accumulation of deferred upgrades can make moving to a newer version costly and affect customers. He proposes an internal standard called 0114: zero people required for routine maintenance through automation, the ability to upgrade every component in the fleet, completing the upgrade in less than a day, and running an update cycle every 14 days or less.
He also advises against rebuilding what is already available at sufficient quality. Instead of writing a new PostgreSQL engine, he gives the example of building a control layer that performs daily backups to object storage, which is the part required to meet regulatory requirements. The idea is to focus on the value the organization needs, not on the task that is most technically interesting.
Clear decisions and a contractual relationship with consumers
Platform owners should listen to customers, but they should not implement every request. Resources are limited, and keeping old features without retiring them accumulates technical debt and prevents the development of more important capabilities. Therefore, the platform should have a clear “point of view”: it should define what it will support, what it will stop supporting, and what serves the majority of users.
This includes formally defining responsibility boundaries through APIs, service-level agreements, and processes. It is not enough for the team to know what it provides; the consumer must also know what falls under its responsibility and which boundaries the platform does not cross. Liste compares platforms to components with fixed shapes: a small team cannot build a custom product for every customer.
Practical experience: experiment early, and use abstraction without hiding details
Liste advises failing quickly and repeatedly after deciding to build, while protecting customers during the transition. He cites early experiments with Linux containers and different orchestration tools that ultimately led to a move to Kubernetes; early experimentation allowed the team to learn before waiting for a single solution to mature.
At the same time, abstraction layers should not obscure what is happening underneath. A user interface, APIs, and Infrastructure as Code, such as Terraform, can be provided, but enough visibility and detail must remain available to understand failures and adjust behavior when necessary. Liste concludes by emphasizing building on open-source software and open standards, so that the platform team can focus on integration and added value rather than recreating every layer from scratch.
Why do these principles matter?
The core value of Liste’s argument is that it shifts platform building from being a project to launch a collection of tools into a long-term operational commitment. The platform affects many teams after they adopt it, making updatability, clarity of responsibilities, and retirement management more important than quickly adding a new feature. These principles remain practical-experience guidelines, not a unified standard that guarantees success; availability levels, the scope of automation, and abstraction choices remain tied to the nature of the system and the needs of its users.