Event-driven architecture is evolving from a limited model involving a few services into a broad network of producers, consumers, and messaging channels. At this point, the core problem is no longer sending the message, but knowing what is being sent, who owns it, who depends on it, and how its schema or underlying infrastructure can be changed or recreated without causing production failures. This is the focus of Ian Cooper’s session on managing asynchronous APIs at scale, based on real-world engineering practices at Just Eat Takeaway.
Why Does It Become More Complicated as Systems Scale?
In small organizations, team members may know which services publish and consume events, and topics or channels can be created through the messaging framework or by submitting a direct request to the platform team. But this approach loses its effectiveness as the number of teams and services grows. It may not be clear who owns a particular endpoint, what contract describes the message, or whether there are unknown consumers in data and analytics teams.
The risk emerges when a team changes a message schema after consulting only known consumers, while other teams depend on the message without being included in the communication loop. Deleting a topic believed to be unused may result in the loss of a critical path, especially if there is no reliable way to determine its owner or recreate the required infrastructure other than redeploying large parts of the system.
Three Axes for Managing Event APIs
Cooper divides the problem into discovery, governance, and provisioning. This begins with understanding what must be described for any endpoint through what he calls the “ABCs”: address, binding, and contract. The address identifies where messages flow, the binding describes the protocol, transport, and encoding, while the contract defines the headers and data carried by the message.
In this context, AsyncAPI provides an equivalent to HTTP API documentation mechanisms such as OpenAPI, modeling servers, channels, messages, operations, and protocol-specific bindings. It supports describing messages using JSON Schema, Avro, and Protobuf, and allows definitions to be reused instead of repeated across multiple files. But Cooper emphasizes that placing hundreds of AsyncAPI files in a single repository does not solve discovery by itself; text search remains limited as the event inventory grows.
Organizations may therefore need catalog tools such as EventCatalog or Marmot, or an open registry such as xRegistry, to display message flows and their relationships and link schemas to producers and consumers. These tools differ in their level of visualization and ready-made interfaces, but their common function is to move discovery from a question asked in Slack or searched for across scattered repositories to a queryable service.
The Role of CloudEvents in Standardizing Metadata
CloudEvents addresses a different aspect of the problem by providing a standardized set of metadata such as the identifier, source, version, and type, with optional fields for the data content type, subject, time, and schema URL. The event type can be used to route the message or select its deserialization mechanism when multiple types share a single channel.
The session explains that the binary and structured modes in CloudEvents relate to the transport protocol’s ability to carry headers, not to whether the message is “binary” or “structured” in the conventional sense. This detail is important in systems such as SNS, where a limited number of message attributes can be consumed quickly if all CloudEvents data is placed in headers, making encapsulating it inside the body a practical option in some scenarios.
Governance Is Not Limited to Documentation
Documenting the contract tells consumers what the message is, but it does not prevent the producer from publishing a change that breaks their dependencies. The approach presented therefore recommends using a schema registry that applies compatibility rules when contracts are updated and rejects messages or changes that do not comply with the approved rules.
Safe practices typically include adding fields as optional, postponing the deletion of fields until all consumers have stopped using them, and treating a rename as a change consisting of a deletion and an addition. Changing a field’s type is generally considered a breaking change and can be implemented by adding a new field with the required type, migrating consumers to it, and then removing the old field.
What Changes in Practice?
The core message for the technical reader is that managing events at scale requires an interconnected chain, not a single tool: AsyncAPI to describe interfaces, CloudEvents to standardize metadata, a catalog or registry to facilitate discovery, a schema registry to enforce compatibility, and automation to provision infrastructure and monitor its drift. Without these layers, event-driven architecture can become a network of invisible dependencies.
Practical constraints remain: standards alone do not provide complete knowledge of consumers, and tools differ in their support for interfaces, visualization, and broker integration. Applying this approach therefore requires clearly assigning ownership of contracts and channels, linking changes to CI/CD processes, and maintaining a recovery path capable of recreating resources rather than relying on individual knowledge or outdated documentation.