Before you startDDD and Microservices: The Atlas#
A very short book for developers and architects who have heard every one of these words and would like to know what each one is for.
The promise
Read this in two evenings and you will be able to draw the boundaries of a system, say which of them deserve to be separate services and which do not, explain why each pattern in the field exists and what it reacted against, name the trap someone is about to fall into before they fall into it, and hold your own in a design review against anyone who says "microservices" with confidence.
It reads as one story. Bramble is an online shop. It starts as a monolith that works, splits itself badly, gets rescued by a set of ideas that are older than most of the people using them, and ends with the judgement to know when not to split. Every chapter starts with something breaking or someone asking for something the current design cannot do. The fix is the chapter's idea. By the end of the chapter, a new crack has appeared.
Every idea arrives four ways: a place in one running metaphor (a world atlas of countries, languages and treaties), a diagram of how it actually works, a one-sentence Remember it as line you can repeat out loud, and a dated account of who thought of it, when, and why. Every chapter then does the work a mentor would do at a whiteboard: it names the traps and their tell-tale signs, gives you a small design with one planted flaw to find, and hands you three questions to ask whoever proposes the pattern next, plus the one point on which experts still genuinely disagree, argued fairly from both sides.
Who this is for
You have built or maintained a system with more than one team on it. You have heard "bounded context", "aggregate", "hexagonal", "saga" and "outbox" used in meetings and nodded. You may have seen a microservices migration go badly and not been able to say precisely why. You want the twenty percent of the field that lets you design well, review sharply and argue clearly, without reading four thousand pages.
If you have already led two decompositions, run Event Storming workshops and can recite Vernon's aggregate rules, this book will be too slow for you; skim the At the whiteboard boxes and the epilogue.
How to read it
- Straight through the first time. The plot carries the ideas, and each chapter's failure is the reason the next idea exists.
- The boxes are where the book becomes a tool. The traps is a checklist. Spot it trains your eye; cover the answer. At the whiteboard is what to say in the meeting.
- The Epilogue is the reference: the full metaphor table, the complete review checklist, the timeline of the field, a reading list by year, and two Java package layouts.
The cast
Bramble sells stationery online, then adds marketplace sellers and, late on, a subscription box. It takes orders, takes payment, picks and ships, handles returns and runs promotions. Ana is the senior developer who becomes its architect and makes every early mistake herself. Kofi is the staff engineer who joins after the first disaster; he has done this twice before and carries the history. Marcus is the consultant who proposes microservices; his arguments are the real ones, and when he is wrong it is because a premise fails at Bramble, not because he is foolish. Lena founded the company and keeps asking for things.
Code, where it appears, is Java 21 from Bramble's Ordering and Billing contexts. It illustrates shape; it is not a buildable project.
Dates and sources
Every "who and when" in this book was checked against the primary source named in research/milestones.md in the book's repository, and the epilogue's reading list gives them by year. Where an idea has a famous date and a publication date that differ, the text uses the first public one and says so.
Contents
Act I: The monolith that worked (then didn't)
- One Database, One Dream (the monolith, the Big Ball of Mud, language drift)
- Let's Do Microservices (the 2014 definition, entity services, the shared database)
- The Night Everything Deployed Together (the distributed monolith, call chains, the fallacies)
Act II: Finding the seams (strategic DDD)
- The Workshop With the Orange Stickies (Event Storming)
- Problem Space, Solution Space (core, supporting and generic subdomains)
- Drawing the Borders (bounded contexts; context ≠ subdomain ≠ service)
- The Map of Treaties (context mapping; Conway's law)
- Nobody Owns the Product Table (data ownership; database per service)
Act III: Inside a border (architecture styles and tactical DDD)
- Which Way Do the Arrows Point? (layered, hexagonal, onion, clean, vertical slices, modulith)
- The Order That Ate the Database (aggregates, entities, value objects, invariants)
- Something Happened (domain events, integration events, the anti-corruption layer in code)
- Two Models Are Cheaper Than One (CQRS)
- The Ledger Never Lies (event sourcing)
Act IV: Talking across borders (integration)
- The Payment Taken Twice (sync versus async, timeouts, retries, idempotency, circuit breakers)
- The Message That Never Left (dual write, outbox, CDC, inbox, contracts, schema evolution)
- Who Says the Order Is Done? (sagas, compensation, eventual consistency as a business decision)
- The Report Nobody Could Run (analytics across contexts, API composition, BFF, data mesh)
Act V: People, migration and knowing when to stop
- You Ship Your Org Chart (Conway's law, Team Topologies, cognitive load)
- Strangling the Monolith (strangler fig, branch by abstraction, parallel run)
- Can You See It? (contract tests, tracing, independent pipelines, versioning)
- The Consultant Who Said Microservices (the premium, MonolithFirst, merging back, the decision)
Epilogue
- The Atlas: the metaphor table, the review checklist, the timeline, the reading list, the package layouts
Chapter 1 · Act IOne Database, One Dream#
The problem
Bramble began in a spare room with Lena, a laptop and four hundred fountain pens. Ana was the third hire. The first system took her six weeks: one Spring Boot application, one PostgreSQL database, one git push to deploy. Orders, stock, payments, returns and the newsletter all lived in the same codebase and talked through method calls. It was, by any honest measure, the right design.
Three years later Bramble ships two thousand parcels a day and the application has forty thousand lines, eleven developers and a customer table with sixty-one columns. Marketing uses it for people who browse. Finance uses it for people who owe money. Support uses it for people who ring up. A column called status means "has verified their email" to one team and "is allowed credit" to another, and nobody remembers who added status2.
This morning a two-line change to the returns policy broke the promotions engine. Both read the same Order.total() method, and each had, over the years, quietly adjusted what it returned. The fix took nine hours because nobody could say who else depended on it. Lena has stopped asking "how long will it take" and started asking "what will it break".
"We need to split this up," a developer says in the retrospective. Ana is not sure yet what "this" is.
The idea
A monolith is a system deployed as one unit. That is all the word means. It says nothing about how well it is organised inside, and for a new business with a small team and an unknown domain it is the right first choice: one build, one deploy, one place to debug, and refactoring is a rename in an IDE instead of a negotiation between teams. Bramble's first three years were fast because of the monolith, not despite it.
Inside, Ana's application has the shape most applications have: a layered architecture. Controllers on top, services in the middle, repositories at the bottom, the database underneath. Each layer may call the one below it. It is a sensible way to separate "how we talk to the web" from "what the business does" from "how we store things".
The trouble is that layers separate technical concerns and say nothing about business ones. Every feature touches every layer, so every layer grows in every direction. Returns and promotions are different parts of the business, but in a layered monolith they are the same three folders. Over time the dependencies between the pieces stop following any rule, and the system becomes what Brian Foote and Joseph Yoder named a Big Ball of Mud: haphazardly structured, held together by duct tape, and understood by nobody in full. It is not a design failure so much as the absence of a design decision, repeated daily.
The deeper symptom is in the words. When Bramble was small, everyone meant the same thing by "customer". Now there are at least three meanings in one table: the shopper marketing wants to convert, the account holder who can log in, and the payer finance sends invoices to. Domain-Driven Design calls the shared, precise vocabulary that a team and its code agree on the ubiquitous language. When one word carries three meanings, the language has split and the code has not noticed. That gap, between the words the business uses and the words the code uses, is where most of the nine hours went.
In the atlas metaphor this book uses, a business is a territory. Bramble's territory started as one village where everyone spoke the same tongue. It has grown into a region with several dialects and no borders, where every dialect is shouted at once. The ball of mud is not a lack of buildings. It is a lack of a map.
Remember it as: a monolith is one deployable, and that is fine; a ball of mud is one deployable with no borders inside, and that is the problem.
Notice what Ana's colleague did not say. Nobody said "the business has three meanings of customer and the code has one". They said "split it up", which is a deployment answer to a vocabulary problem. Holding on to that distinction is the whole of Act I.
What Ana learned
- Keep the monolith while the domain is still being learned; it is the cheapest place to move a boundary.
- Treat a word with two meanings as a defect, and find where in the code the meanings diverge.
- When someone says "split it", ask which business boundary they mean before asking which deployment.
…but
Lena has already booked a consultant. His deck has the word "microservices" on the first slide, and a diagram of Netflix on the second.
Chapter 2 · Act ILet's Do Microservices#
The problem
Marcus arrives on a Tuesday with a deck and a good reputation. He has done this at a bank and a travel company, and he is not selling snake oil. His pitch is careful: Bramble's problem is not size, it is coupling. Every team deploys the same artefact, so every team waits for every other team. A bad change anywhere is an outage everywhere. Hiring is slower because a new developer must understand forty thousand lines before touching any of them.
His remedy is the one everyone in the room has read about: small services, each owning its data, each deployed on its own, each owned by a team small enough to feed with two pizzas. He quotes the definition from James Lewis and Martin Fowler's 2014 article and cites Stefan Tilkov's argument that a company which already understands its domain should not start with a monolith. "You are not a startup any more," he says. "You know your business. Draw the services around it."
Lena asks how long. Marcus says a quarter for the first cut. Ana asks what the services should be. Marcus says that is a workshop, and books one for Thursday.
On Thursday the whiteboard fills quickly, because the team already has a list of nouns.
The idea
The 2014 article described microservices as a style in which a system is built from a suite of small services, each running in its own process, communicating through lightweight mechanisms, built around business capabilities, independently deployable, with decentralised governance and decentralised data management. Read that list slowly, because Marcus is right about all of it and the room is about to ignore most of it.
The test that matters is the one Sam Newman would make central a year later: a microservice is independently deployable. If you cannot ship a change to one service without coordinating a change to another, you do not have two services. You have one service in two repositories.
What the room drew is a service per noun. Customer, Product, Order, Payment: each becomes a service that stores that table and offers create, read, update and delete over HTTP. It feels natural because it mirrors the database, and that is exactly the problem. A noun is not a business capability. "Customer" is the word chapter 1 found to have three meanings, and the new CustomerService has all three inside it, now behind a network call. Nothing has been decomposed; a table has been given a URL. These are entity services, and every real operation, such as placing an order, now needs three or four of them in a row.
A second developer proposes the other classic cut: split by technical layer, with a web service, a business-logic service and a data service. This is worse. Every feature now spans three deployables owned by three teams, so no change is ever local.
Then comes the sentence that costs the most. Splitting the database is hard, the tables reference each other, and the quarter is short. "We'll keep the shared database for now." Everyone nods, because it is obviously temporary.
In the atlas, a country is a place with its own language, its own law and its own records office. What the room drew is four parliament buildings sharing one records office, and the clerks still walk between them. That is not four countries. It is one country with more corridors.
Remember it as: a microservice is something you can deploy alone; a service per noun sharing a database is a table with a URL.
The confusion to clear up now, because it will recur in every later chapter, is the difference between the mechanics of separation and the fact of independence:
| Looks like independence | Is independence |
|---|---|
| A separate repository | Can be released without releasing anything else |
| A separate container or process | Can change its schema without another team's migration |
| Its own REST API | Keeps working, or degrades gracefully, when a neighbour is down |
| Its own team name | Owns every decision needed to ship a change to it |
Bramble's four services will pass every test in the left column and fail every one in the right.
What Ana learned
- Test every proposed service against one question: can it be deployed alone, schema and all?
- A noun is not a capability; a service per table decomposes nothing.
- A "temporary" shared database is a permanent public contract between every service that touches it.
…but
The quarter ends. Four services are live, the shared database is still shared, and the first big promotion of the year starts on Friday at midnight.
Chapter 3 · Act IThe Night Everything Deployed Together#
The problem
The promotion goes live at midnight and by ten past the checkout is timing out. Ana is on the call with three other developers, and the dashboards all say the same unhelpful thing: everything is slow.
The chain is simple to draw and impossible to fix live. Checkout calls OrderService. OrderService calls CustomerService for the address, ProductService for prices and stock, and PaymentService to take the money, each over HTTP, each waiting for the previous. ProductService is slow tonight because the promotion put a new query on its hot path. Its response time went from forty milliseconds to four seconds. OrderService has a pool of two hundred threads, each now parked for four seconds waiting on products, so it stops answering the checkout. The checkout retries. So does the mobile app. So does OrderService itself, because somebody enabled three automatic retries "for resilience". Product is now receiving four times the traffic that made it slow in the first place.
At one in the morning they find a fix: a new index on the products table. Deploying it means a schema migration on the shared database, which means CustomerService and OrderService must be redeployed too, because they map the same table. Four services, one release, all at once.
At two, the checkout is back. In the morning Ana writes the incident report and, at the bottom, a sentence she will repeat for years: "We have all the costs of microservices and none of the benefits."
The idea
What Bramble built is a distributed monolith: several deployables that must be released together, share a schema, and fail as a chain. It has the operational cost of many services (networks, deployments, dashboards, on-call for each) and the coupling of one. It is the most common outcome of a first decomposition, and it is worth understanding precisely why it fails, because each reason becomes a chapter later.
First, synchronous call chains multiply latency and divide availability. If each of four services answers in time 99 percent of the time, the chain succeeds 0.99⁴ of the time, and every hop you add makes it worse:
| Services in the chain | Each 99% available | Each 99.9% available |
|---|---|---|
| 1 | 99.0% | 99.90% |
| 4 | 96.1% | 99.60% |
| 8 | 92.3% | 99.20% |
| 12 | 88.6% | 98.81% |
The monolith made these calls in-process, where the network cannot fail and a slow method does not hold a socket open. Putting a network between them without changing anything else made the system strictly worse.
Second, slowness travels. A slow dependency exhausts the caller's thread pool, so the caller becomes slow to everyone, including callers that never needed the slow dependency. This is a cascading failure, and retries are its accelerant: every retry is new load on the thing that is already failing. Chapter 14 gives the tools (timeouts, circuit breakers, bulkheads, idempotent retries) that stop the cascade; tonight nobody had them, and the shared thread pool became the fuse box for the whole house.
Third, the shared database means that no service is independently deployable. The schema is a contract between all four, and any migration is a joint release. The room in chapter 2 traded method calls for HTTP calls and kept everything that actually coupled them.
Peter Deutsch's fallacies of distributed computing list the assumptions programmers make when they first put a network inside a system: the network is reliable, latency is zero, bandwidth is infinite, the network is secure, topology does not change, there is one administrator, transport cost is zero, and the network is homogeneous. Bramble's decomposition assumed the first three by default, because in-process code had never needed to think about them.
In the atlas, Bramble built a house whose rooms are in different cities but share one fuse box. Walking from the kitchen to the bedroom now takes a train, and when the bedroom trips the breaker the kitchen goes dark too.
Remember it as: a distributed monolith is rooms in different cities sharing one fuse box: all the travel, none of the independence.
The lesson is not "microservices are bad". It is that the decomposition skipped the only step that matters, which is finding boundaries along which the business, the data and the failures naturally separate. That step has a name, a workshop and a fifty-year history, and it is where Act II begins.
What Ana learned
- Count the services a request needs right now; the product of their availabilities is your availability.
- Never add a retry without a timeout, and never add either without asking what the user sees when it fails.
- A shared schema means a shared release; the database, not the network, decides whether services are independent.
…but
Lena hires a staff engineer called Kofi, who has done this twice. On his first morning he reads the incident report and asks the room a single question: "What does the word customer mean here?"
Chapter 4 · Act IIThe Workshop With the Orange Stickies#
The problem
Kofi's question sits in the room for a while. Marketing says a customer is anyone who has given an email address. Finance says a customer is someone with an invoice. Support says a customer is whoever is on the phone, invoice or not. Ana says a customer is a row in the customer table, and hears how it sounds.
"That's the problem," Kofi says. "Not the services. Not the database. Four teams, four meanings, one word, and the code picked one at random." He asks Lena for a day, a wall, and everyone who touches an order: two developers, the warehouse lead, someone from finance, someone from support, and the person who runs promotions. Lena asks whether this is a technical workshop. Kofi says it is the opposite, which is why it will work.
Ana is sceptical. She has sat through modelling workshops that produced a class diagram nobody used. This one starts with a roll of brown paper eight metres long, a stack of orange sticky notes, and a rule: write down things that happen, in the past tense, one per note, and stick them on the wall in the order they happen.
The idea
Event Storming is a workshop for discovering a business by listing its events. An event is something that happened and that the business cares about: Order Placed, Payment Taken, Parcel Dispatched, Return Requested. Past tense is not decoration; it forces people to describe facts rather than screens or tables, and facts are what every department can agree on even when they disagree about everything else.
The big picture session covers the whole flow from a shopper's first visit to the last refund. Everyone writes at once and sticks notes anywhere; then the group sorts them into a rough timeline. At Bramble the wall fills in twenty minutes with sixty orange notes, and three things happen that no class diagram could have produced.
First, the timeline has gaps and arguments. Finance puts Invoice Issued before Parcel Dispatched; the warehouse puts it after. They are both right, for different kinds of customer, and that difference becomes a pink hotspot note: a place where the business is unclear or disagrees. Hotspots are the most valuable output of the day, because each one is a bug the code currently resolves by accident.
Second, some events change everything after them. Order Placed turns a browsing session into a commitment; Payment Taken turns a commitment into money; Parcel Dispatched moves the order out of the shop's hands. Kofi marks these as pivotal events with a vertical strip of tape. The wall now falls into natural chapters, and the chapters are not "customer" and "product". They are shopping, ordering, paying, fulfilling, supporting.
Third, the word splits in public. Between Basket Checked Out and Order Placed the warehouse lead asks whose order it is. "The customer's," says marketing. "Which one?" says finance. "The one paying, or the one it's going to? Half our corporate orders are bought by one person for another." Support adds that a customer who rings up about a delivery is neither. On the wall, within a metre of brown paper, "customer" has become shopper in the shopping chapter, buyer and recipient in ordering, payer in paying, and caller in support. Nobody proposed this. The events forced it.
The afternoon is a process level pass on the ordering chapter. Now blue notes appear for the commands that cause events (Place Order causes Order Placed), yellow notes for the people or systems that issue them, and green notes for the information they need to decide. The group walks through the order that broke the promotions engine in chapter 1, step by step, and finds that the returns policy and the promotions rule both read "the order total" but mean different totals: one before discounts and shipping, one after. That, too, was a word with two meanings.
The wall is not a design. It is a map of the territory as the people who live in it describe it, and the borders on it were drawn by the places where the language changed. In the atlas, a border is exactly that: the line where a word changes meaning. The next two chapters turn those lines into a solution.
Remember it as: list what happens, in the past tense, in order; where the words change meaning, you have found a border.
A related technique, domain storytelling, tells the same stories as pictographic sentences (actor, activity, object) and is a good follow-up when a single flow needs precision. Event Storming's advantage is breadth and speed: a whole business on a wall in a day, with its disagreements visible.
What Ana learned
- Start with events in the past tense and the people who witness them, not with tables or screens.
- Mark every disagreement as a hotspot and treat it as a defect the current code hides.
- The places where a word changes meaning are the first draft of the system's borders.
…but
The wall has five chapters and thirty hotspots, and Lena, looking at it, asks the question that decides where the money goes: "Which of these is actually us, and which of these could anyone do?"
Chapter 5 · Act IIProblem Space, Solution Space#
The problem
Lena's question is a budget question wearing a modelling costume. Bramble has eleven developers and five chapters on the wall. Marcus's plan treated every service as equally worth building; the shopping experience, the payment integration and the returns desk would each get a team and a roadmap. Kofi says that is how companies end up with a beautifully engineered invoicing module and a checkout that loses to a competitor's.
Ana pushes back, gently. Payments are critical; if they fail, nothing works. Surely critical means core? Kofi draws a line on the whiteboard: critical is about what breaks if it stops, core is about what makes customers choose you. Payments are critical and not core. Bramble does not win because its card processing is better than anyone else's. It wins, Lena says without hesitating, because of the curation: which pens, which papers, which bundles, and the promotions that make people feel clever for buying them.
That sentence reorganises the wall. The question is no longer "what services do we need" but "which parts of this problem deserve our best people, which deserve good-enough, and which should we not build at all".
The idea
Domain-Driven Design draws a line between the problem space, which is the business as it is, and the solution space, which is the software you build to serve it. The problem space is divided into subdomains: areas of the business with their own rules and vocabulary. Bramble's wall already showed five. The trick is that not all subdomains are equal, and treating them as if they were is how effort gets misallocated.
A core subdomain is where the business wins. It is what competitors cannot easily copy, it changes often because the business is always sharpening it, and it is where the most complex and most valuable rules live. At Bramble that is curation and promotions: the bundle logic, the loyalty rules, the "customers who bought this ink also needed this converter" judgement that Lena has been doing by hand. It is not the biggest part of the system. It is the part that must be built in-house, by the strongest people, with the richest model.
A supporting subdomain is necessary and specific to the business but not a differentiator. Bramble's returns handling has its own rules (the fourteen-day window, the restocking exceptions for personalised items) that no off-the-shelf product knows, but nobody chooses Bramble for its returns desk. Build it, keep it simple, do not over-engineer it.
A generic subdomain is a problem everyone has and somebody has already solved. Payments, email delivery, authentication, tax calculation. The right move is to buy or adopt: a payment provider, an identity service, a tax API. Building your own is not ambition; it is spending core-domain talent on a solved problem.
In the atlas, the territory has a capital, the farmland that feeds it, and the utilities it buys in. You defend the capital with your best people. You farm the land well enough. You do not dig your own wells when there is a water company.
A useful tool for the conversation is the core domain chart, which Scott Millett and Nick Tune drew as two axes: business differentiation against model complexity. Things that are high on both are core. High complexity and low differentiation is where teams waste years: a bespoke payments engine, a home-grown search index. Low complexity and high differentiation is a gift, usually a small rule that makes a big difference, and should be protected from being buried under generic plumbing.
Two refinements stop this from becoming a labelling exercise. First, classification changes over time. When Bramble adds marketplace sellers, seller onboarding starts as supporting and may become core if the marketplace becomes the business. When a good open-source promotions engine appears, promotions could slide toward generic, though at Bramble that seems unlikely. Re-ask the question every year. Second, the type of a subdomain should drive how it is built, which chapter 10 makes concrete: a core subdomain earns a rich model and a senior team; a supporting one is well served by plain, simple code; a generic one is an integration.
Remember it as: the capital, the farmland, the utilities: defend the first with your best people, farm the second plainly, buy the third.
The question Lena asked, "which of these is us", is the most important architectural question in the book, and it is not a technical one. Ana notes that Marcus never asked it.
What Ana learned
- Ask "what do customers choose us for" before asking "what services do we need".
- Buy the generic, build the supporting plainly, and spend the strongest people on the core.
- Re-classify every year; subdomains move as the market and the vendors move.
…but
The wall says shopper, buyer, recipient, payer and caller. The database still says customer. Somebody has to decide how many models of a person Bramble is going to have, and where each one stops.
Chapter 6 · Act IIDrawing the Borders#
The problem
Ana's first instinct is the one every developer has: build a single Person model rich enough to satisfy everyone. Shopper, buyer, recipient, payer, caller: five roles on one class. It would be big, but it would be one.
Kofi lets her draw it. Forty minutes later the whiteboard has a Person with twenty-eight fields, six of which are nullable depending on which role is active, three "type" flags, and a comment reading "recipient may not have an email". Marketing's shopper needs consent flags and browsing history and no name. Finance's payer needs a legal name, a VAT number and a credit limit and does not care what the person browsed. Support's caller needs an order history and a phone number and may be neither shopper nor payer.
"You've built the customer table again," Kofi says, "with better intentions." Ana knows he is right. Every field on the class belongs to a different conversation, and any rule about one role has to tiptoe around the nulls of the others.
The alternative feels wrong at first: several models of a person, each small, each complete for its own purpose, each with its own name. Ana asks how that is not duplication. Kofi says it is the opposite, and draws a border.
The idea
A bounded context is a boundary inside which a model and its language are consistent. Within the border, one word has one meaning and one model serves it. Across the border, the same real-world thing may have a different name, a different shape and different rules, and that is correct, not a mistake to be reconciled.
In the atlas, a bounded context is a country: one language, one law. The border is exactly where a word changes meaning. The same human crosses three borders and is a different noun in each, with different papers, and no country needs the others' paperwork. The Shopping context keeps consent and browsing history; Ordering keeps the delivery address; Billing keeps the legal name and the VAT number. Each model is small enough to hold in your head and complete for its own rules. The Person with twenty-eight nullable fields was one passport trying to satisfy three border guards.
Three confusions now need clearing, because they cause more bad architecture than any other misunderstanding in the field.
| Subdomain | Bounded context | Microservice | |
|---|---|---|---|
| Lives in | Problem space: the business | Solution space: the model | Solution space: the deployment |
| Answers | What does the business do? | Where does one language stop? | What ships and runs on its own? |
| Discovered or decided | Discovered from the business | Decided by the design team | Decided by the team and the platform |
| Ideal relationship | One subdomain… | …served by one context (often several) | …deployed as one or several services, or one module |
A bounded context is not a subdomain. Subdomains are found in the business; contexts are drawn by the people building the solution. The ideal is one context per subdomain, and a legacy system frequently has one context sprawled across several subdomains (the ball of mud) or several contexts fighting over one.
A bounded context is not a microservice. A context is a boundary of language and model; a service is a boundary of deployment. A context can be deployed as one service, as several, or as a module inside a single deployable. What must not happen is one service spanning two contexts, because it will then carry two meanings of the same word and become the customer table with a URL. Bramble's three person-models could be three services, or three modules in one application, and the border would be equally real either way. The modular monolith, one deployable with enforced borders between its modules, is a legitimate and often better home for a context, and chapter 9 shows how to enforce the borders in code.
Which raises the question everyone asks: how big should a context, or a service, be? Sam Newman's honest answer in 2015 was that "small" is the wrong axis. The right size is the one a team can own and change independently, and that is bounded by language, not line count. Uber, with thousands of services by 2020, regrouped them into domains and stopped treating the service as the unit of design. The test is whether one team can change it without asking permission, and whether one word means one thing inside it.
Remember it as: a bounded context is a country: one language, one law; the same person is a different noun on each side of the border.
The duplication Ana feared is not duplication. Three small models that each own their rules are cheaper to change than one large model whose every field is a negotiation.
What Ana learned
- Draw the border where a word changes meaning, and give each side its own small, complete model.
- Never let one deployable span two contexts; whether one context is one deployable is a separate decision.
- Judge a context by whether one team can change it alone, not by how many lines it has.
…but
Three countries now exist on paper, and an order crosses all three. Someone has to decide who tells whom what, in whose language, and what happens when Billing changes its mind about the shape of an invoice.
Chapter 7 · Act IIThe Map of Treaties#
The problem
The three countries last a week before the first border incident. Billing changes the shape of an invoice: a new field for the payer's purchase-order number, mandatory for corporate customers. It is a sensible change, made inside Billing's own model, by Billing's own team. On Thursday the Ordering module's build goes red, because Ordering had been reading Billing's Invoice class directly to show an "invoice pending" badge, and the new mandatory field has no value at the moment the badge is drawn.
Nobody did anything wrong by the rules they knew. Billing changed its own model. Ordering used what was available. The failure is that the relationship between the two was never written down: who depends on whom, in whose language, and with what promise of stability. In the old monolith every module could reach every other and the compiler resolved the argument. Across borders there is no compiler.
Kofi puts a fresh sheet on the wall, draws three boxes, and asks a question for each pair: when these two disagree, who wins?
The idea
A context map is the picture of how bounded contexts relate: which one depends on which, in what direction influence flows, and what each side has agreed to do about the other's model. It is the atlas page of treaties between countries. Drawing it is a political act as much as a technical one, because every arrow records who has to change when the other side moves.
The first thing to fix on each relationship is direction. The upstream context is the one whose model flows to the other; the downstream context receives it and must cope with changes. Upstream decides what flows down the river. Billing is upstream of Ordering for invoice status; Ordering is upstream of Billing for what was bought. The second thing to fix is the treaty pattern, and Domain-Driven Design names the ones that recur:
- Partnership: two contexts that succeed or fail together, and plan changes jointly. Allied nations. Ordering and Fulfilment at Bramble: neither ships a change that breaks the other, and they release in step by agreement.
- Shared kernel: a small piece of model that two contexts literally share, such as the identifier scheme for an order, and that neither may change alone. A jointly governed border town. Keep it tiny; every street in it needs two signatures.
- Customer–supplier: upstream supplies, downstream is a customer whose needs are heard and planned for. A trade agreement in which the downstream has a say. Billing supplies invoice status to Ordering and puts Ordering's needs on its backlog.
- Conformist: downstream simply adopts upstream's model, because upstream will not or cannot accommodate it. Adopting the neighbour's language wholesale. Reasonable when upstream is a good, stable external product; a trap when it is the legacy ball of mud.
- Anti-corruption layer: downstream translates upstream's model into its own at the border, so that upstream's language never leaks in. The customs house with translators. Bramble's Ordering should never hold a Billing
Invoice; it should translate what it needs into its ownPaymentStatus. - Open host service and published language: upstream offers a well-defined interface for anyone, in a stable, documented language separate from its internal model. The port open to all ships, and the trade lingua franca. Billing publishes
InvoiceIssuedevents in a public schema, not its internal classes. - Separate ways: the two contexts do not integrate at all, because the cost exceeds the value. No treaty. Bramble's newsletter and its returns desk share nothing and should not pretend to.
- Big ball of mud: a region on the map with no internal borders, which you wall off with an anti-corruption layer and do not try to reform from outside.
How to read the notation, which the rest of the book uses unchanged: arrows point from upstream to downstream; the label names the upstream's pattern, then the downstream's (OHS open host service, PL published language, ACL anti-corruption layer, CF conformist, CS customer–supplier); a two-headed arrow is a partnership; a small purple node joined to two contexts is a shared kernel; separate ways is stated in the caption and never drawn. Bramble's returns desk and newsletter: separate ways.
Thursday's incident now has a name. Ordering was an accidental conformist to Billing's internal model, with no customs house. The fix is two-sided: Billing offers a published language (an event or an interface that is not its internal class), and Ordering installs an anti-corruption layer that translates it into Ordering's own terms. Billing can then add all the fields it likes.
There is a second overlay on this map that most teams discover late. In 1968 Melvin Conway observed that a system's structure mirrors the communication structure of the organisation that builds it. The treaties on the context map will end up matching the treaties between teams whether you draw them or not. If Billing and Ordering are one team, they will drift into a shared kernel. If they are two teams that barely talk, they will drift toward conformist or separate ways. Conway's law is not a warning to fight; it is a tool. Chapter 18 uses it deliberately.
Remember it as: the context map is the atlas of treaties: for every border, who is upstream, and does the downstream translate, conform or co-plan?
The map also gives Bramble its first honest picture of coupling. Vlad Khononov's 2024 framing is useful here: coupling is a product of how strong the connection is, how far apart the two sides are, and how often the upstream changes. A shared kernel between two modules in one team is fine; the same shared kernel between two services in two departments is a standing emergency.
What Ana learned
- Write every border down as a treaty: direction, pattern, and what crosses.
- Translate at the border; never import another context's internal model, especially a vendor's or a legacy system's.
- Expect the map to match the org chart, and redraw it whenever teams change.
…but
Every treaty on the wall assumes each country keeps its own records. Bramble's three contexts still read and write one customer table and one product table in one database, and the first team to run a migration will find out what that costs.
Chapter 8 · Act IINobody Owns the Product Table#
The problem
The Shopping team wants to add a hero_image_url and a marketing description to products. The Fulfilment team wants to add weight and dimensions for the courier. Finance wants to add a tax category. All three write a migration against the product table in the same week, and two of the migrations conflict on a renamed column. The third goes through and breaks the promotions job, which had been selecting * and mapping positionally.
Ana calls a meeting to decide who owns the product table. Everyone claims it. Everyone is right, for their meaning of "product": to Shopping it is a thing to be shown and desired; to Fulfilment it is a thing with a weight and a shelf location; to Billing it is a thing with a price and a tax code. One table, three contexts, no owner. It is the customer table all over again, one noun to the left.
Then someone asks the question that shows the real fear: "If each context has its own product data, won't we have the same product in three places?" The room goes quiet, because that sounds like the thing every developer was taught never to do.
The idea
A context's model is only as independent as its data. If three contexts write one table, the schema is a treaty between all three, and no context can change without a meeting. That was the shared-database lesson of chapter 3, and it applies inside a monolith just as much as between services: data ownership follows the border. Every table, every document, every stream has exactly one context that may write it, and the others get what they need through that context's published interface.
Taken to the deployment level this becomes database per service: each service has a schema, a database or a cluster that no other service touches. The unit of isolation matters less than the rule. Three schemas in one PostgreSQL instance with separate credentials give most of the benefit; three modules sharing one product table give none of it.
Now the fear. Yes, the product's name will exist in three schemas. This is reference data duplication, and it is the correct design, for three reasons. First, each context holds only the fields it needs, under its own name (ProductListing, StockItem, PricedItem), so the "duplicate" is small and shaped for one purpose. Second, the copies are not edited independently; a Catalogue context owns the identity and the name, publishes a ProductAdded or ProductRenamed event, and the others update their copy. Third, and decisively, the alternative is not "no duplication" but "one table nobody can change". Storage is cheap. Coordination is not.
In the atlas, each country keeps its own records office. A citizen who moves is recorded again in the new country, in the new country's forms. Nobody proposes a single world registry, because everyone can see it would be governed by whoever changed it last.
The instinct that resists this is the relational one: normalise, keep one copy, join when you need the picture. Inside one context that instinct is right. Across contexts it is the trap that undoes the whole design. A cross-context join, whether in SQL against another schema or in code by calling three services and stitching the results in a loop, couples the caller to the internal shape of every table it touches. When Shopping's nightly job joins orders to stock to invoices, Shopping breaks whenever any of the three changes, and nobody knows to tell it.
Two honest costs come with data ownership, and the book will pay both later rather than pretend they do not exist. Reads that used to be one query now cross a border, and Act IV shows how to move data across it safely (events, the outbox) and how to build read models that make cross-context reads fast again. And reports, which by nature join everything, need a home of their own; chapter 17 gives them one.
Remember it as: one writer per table, and every other reader gets a copy in its own words; duplication is cheap, a table nobody can change is not.
Ana writes the rule on the wall and the meeting ends in ten minutes. Shopping adds its hero image to its own schema on Friday, and nobody else is in the room.
What Ana learned
- Give every table exactly one writing context, enforced by credentials, not by convention.
- Let each context hold its own small copy of what it needs, updated from the owner's events.
- Never join across a border; if you need the picture, build a read model, which is Act IV's work.
…but
The borders are drawn and the records offices are separate. Then a new developer opens the Ordering module to add a field, finds forty classes in one package that reach straight from the web controller to the database row, and asks Ana what a well-built country is supposed to look like inside.
Chapter 9 · Act IIIWhich Way Do the Arrows Point?#
The problem
The new developer is called Priya, and her question is fair. The Ordering module now has a clean border, its own schema and a treaty with everyone around it. Inside, it is the same forty classes it always was: a controller that builds SQL by hand, a service that knows about HTTP status codes, an Order class with a JPA annotation on every field and a method called toJson(). To add a purchase-order number she has to touch nine files, and the unit tests need a running database.
Ana asks Kofi what a well-built context looks like inside, and braces herself, because she has heard the words before. Hexagonal. Onion. Clean. Ports and adapters. Every conference has a talk on one of them, every talk has a different diagram, and the diagrams look like they disagree. She has never been able to tell whether they are four architectures or one architecture with four logos.
Kofi draws a single arrow on the whiteboard and says they are one idea, told four times, by four people who were each annoyed about the same thing.
The idea
The idea is a rule about which way the arrows point. Source-code dependencies must point toward the business logic, never away from it. The domain, the code that knows what an order is and when it may be placed, depends on nothing: not the database, not the web framework, not the message broker. Everything else depends on it. Robert Martin later named this the dependency rule, but it is the shared core of all four styles.
Ana's layered architecture from chapter 1 breaks the rule in one place that matters. In a classic layered design the domain layer sits on top of the data layer and depends on it, so the Order class knows about the table it is stored in, and nothing about the order can be tested without the table. Hexagonal architecture, also called ports and adapters, fixes this by putting the application in the middle and describing every interaction with the outside as a port, an interface the application owns, with adapters outside that implement or call it. Onion architecture draws the same thing as rings, domain innermost, and states the rule as "outer rings depend on inner rings". Clean architecture adds a ring for use cases between the entities and the adapters and makes the dependency rule explicit. Four names, one arrow:
In the atlas, a context built this way is a walled city: the keep holds the domain, the gates are the ports, the gatekeepers are the adapters, and every road leads inward. The keep does not know which gate a visitor came through.
Here is what the arrow looks like in Java. The port is an interface the domain owns; the adapter, in a package that depends on the domain and not the other way round, implements it.
// ordering/domain/PaymentGateway.java — a port. The domain OWNS this interface.
package bramble.ordering.domain;
public interface PaymentGateway {
PaymentResult take(OrderId orderId, Money amount);
}
// ordering/adapters/out/StripePaymentGateway.java — an adapter. Depends on domain;
// the domain never imports anything from adapters.*
package bramble.ordering.adapters.out;
import bramble.ordering.domain.*;
public final class StripePaymentGateway implements PaymentGateway {
private final StripeClient stripe;
public StripePaymentGateway(StripeClient stripe) { this.stripe = stripe; }
@Override
public PaymentResult take(OrderId orderId, Money amount) {
var charge = stripe.charge(amount.minorUnits(), amount.currency(), orderId.value());
return charge.succeeded() ? PaymentResult.taken(charge.id()) : PaymentResult.failed(charge.reason());
}
}
The PlaceOrder use case can now be tested with a fake gateway and no network, and swapping payment providers is a new adapter, not a change to the domain.
There is a cost, and it produced a reaction. In a small context with simple rules, four layers of interfaces between an HTTP request and a row is ceremony: a controller calling a service calling a use case calling a port calling a repository interface calling a JPA repository, to update one field. Vertical slice architecture, Jimmy Bogard's 2018 answer, organises code by feature rather than by layer: one package per use case, containing its request, its handler and whatever data access it needs, sharing only the domain model. Small features stay small; complex ones may still use the full hexagon internally. This is package by feature rather than by layer, and it is not a rejection of the arrow. The domain still depends on nothing. It is a rejection of building every gate the same size.
// ordering/features/placeorder/PlaceOrderHandler.java — one vertical slice.
package bramble.ordering.features.placeorder;
import bramble.ordering.domain.*;
public final class PlaceOrderHandler {
private final OrderRepository orders; // still a port owned by the domain
private final PaymentGateway payments; // still a port owned by the domain
public PlaceOrderHandler(OrderRepository orders, PaymentGateway payments) {
this.orders = orders; this.payments = payments;
}
public OrderId handle(PlaceOrder cmd) {
Order order = Order.place(cmd.buyerId(), cmd.lines()); // the rules live in Order
payments.take(order.id(), order.total());
orders.save(order);
return order.id();
}
}
One more piece closes the loop with chapter 6. If Bramble keeps several contexts in one deployable, the borders between them must be enforced, or the modular monolith decays into the ball of mud within a year. Spring Modulith (2022) does this for Spring Boot applications: each top-level package is a module, an @ApplicationModule may expose only its declared API, and a test fails the build if ordering imports billing.internal. Similar tools exist for other stacks; the point is that a border you cannot fail a build on is a border in a slide deck.
Remember it as: the keep depends on nothing; every road leads inward; build the gates as big as the traffic needs and no bigger.
And the honest caveat: a generic or simple supporting context, a CRUD screen over a table, deserves none of this. A controller and a repository are fine. The architecture styles earn their cost where the rules are rich, which chapter 5 already told you where to find.
What Ana learned
- Make the domain depend on nothing; every arrow points inward, and a framework import in the domain is a defect.
- Own the ports in the domain, implement them in adapters, and test the use cases with fakes.
- Size the ceremony to the rules: full hexagon for the core, a slice for the simple, a controller and a table for CRUD.
…but
Priya moves Order into a clean domain package with no annotations, and then asks where the rule "an order's total can never be negative" should live, because right now it is checked in the controller, the service and a database constraint, and last week all three disagreed.
Chapter 10 · Act IIIThe Order That Ate the Database#
The problem
The answer to Priya's question is "in the Order class", and the team agrees in about a minute. The argument that follows takes a week, because it is about what the Order class is.
The first draft is generous. Order owns its lines, its buyer's details, its payments, its shipments, its returns, its promotion history and its support notes, because all of those are "part of the order". Loading one order for the badge on the account page pulls three hundred rows. Saving it locks seven tables. Two support agents editing notes on the same order overwrite each other, and a shipment update fails because a promotion rule on the same order is mid-flight. The class has become the database with a constructor.
The second draft is a reaction and no better. Order becomes a bag of fields with getters and setters, and every rule moves to OrderService, OrderValidator, OrderCalculator and OrderStateMachineHelper. The total-cannot-be-negative rule is now checked in two of them, differently. Priya asks, reasonably, why the class exists at all.
Kofi says both drafts made the same mistake: they decided what belongs to the order by what is about the order, instead of by what must be true at the same instant.
The idea
Domain-Driven Design's tactical patterns are a small vocabulary for building the inside of a context. Two kinds of object carry the model. An entity has an identity that persists as its attributes change: an order is the same order after a line is added. A value object has no identity and is defined entirely by its attributes: two amounts of fifty pounds are interchangeable, so Money is a value, immutable, and compared by value. In the atlas, an entity is a person with a passport number; a value object is an amount of money.
An aggregate is a cluster of entities and values that must be consistent together, with one entity, the aggregate root, as the only way in. The rule that decides membership is the invariant: a condition that must hold at every instant, not eventually. "An order's total is never negative" is an invariant of the order, so the lines belong inside. "A dispatched order has a shipment" is true within a few minutes, not within the same transaction, so shipments do not. The aggregate is a household: one head signs for everyone, and a house rule is never broken even for a second. Anything that can be true a minute later belongs in another household.
Around the aggregate sit four helpers. A repository loads and saves whole aggregates by identity, the land registry that hands you the whole household or nothing; it is not a query API for the UI. A factory builds an aggregate when construction has rules of its own. A domain service holds a rule that genuinely involves several aggregates and belongs to none, the notary. An application service is the front desk: it receives a command, loads the aggregate, calls one method, saves, and publishes what happened. It contains no rules.
public final class Order {
private final OrderId id;
private final BuyerId buyerId; // another aggregate: referenced by id
private final List<OrderLine> lines = new ArrayList<>();
private OrderStatus status = OrderStatus.DRAFT;
private Money total = Money.ZERO;
public static Order place(BuyerId buyer, List<OrderLine> lines) {
var order = new Order(OrderId.next(), buyer);
lines.forEach(order::addLine);
if (order.lines.isEmpty()) throw new DomainException("an order needs at least one line");
order.status = OrderStatus.PLACED;
return order;
}
public void addLine(OrderLine line) {
if (status != OrderStatus.DRAFT) throw new DomainException("order already placed");
Money next = total.plus(line.subtotal());
if (next.isNegative()) throw new DomainException("total cannot be negative");
lines.add(line); total = next; // the invariant lives here, once
}
private Order(OrderId id, BuyerId buyerId) { this.id = id; this.buyerId = buyerId; }
}
Money is a value object: immutable, compared by value, and the natural home for arithmetic on amounts.
public record Money(long minorUnits, Currency currency) {
public static final Money ZERO = new Money(0, Currency.getInstance("GBP"));
public Money plus(Money o) { return new Money(minorUnits + o.minorUnits, currency); }
public boolean isNegative() { return minorUnits < 0; }
}
Vaughn Vernon distilled aggregate design into four rules, and each comes with the situation in which it is legitimately broken. Protect true invariants in consistency boundaries: only what must hold this instant goes inside; break it only when the business refuses eventual consistency for a rule and you accept a larger lock. Design small aggregates: the first draft's three hundred rows are the cost of ignoring this; the exception is a genuinely atomic cluster, which is rarer than it feels. Reference other aggregates by identity: BuyerId, not Buyer, so the households can be loaded and later deployed apart; the exception is a read-only navigation inside one process, and it still costs you at extraction time. Update one aggregate per transaction, and use eventual consistency across them: the exception is a single-database system where a cross-aggregate rule cannot wait, and the price is deadlocks now and an impossible split later.
The second draft has a name too. Martin Fowler called it the anaemic domain model: objects that are data and services that are logic, an object-oriented costume on procedural code, with every rule findable only by grep. A rich model puts the rule next to the data it protects.
But rich is not always right. Fowler's own catalogue lists transaction script (a procedure per use case) and active record (an object per row that saves itself) as honest choices, and Vlad Khononov's heuristic ties the choice to chapter 5's classification: a generic or simple supporting subdomain is well served by a transaction script or active record; a supporting subdomain with real rules earns a domain model; a core subdomain earns a rich, sometimes event-sourced model. Applying aggregates to a lookup table is as much a mistake as applying a transaction script to promotions. Skipping strategic design and applying the tactical patterns everywhere is what Vernon called DDD-lite, and it produces beautifully modelled contexts that nobody needed.
Remember it as: an aggregate is a household: one head signs, no house rule is ever broken, and neighbours are known by address, not kept indoors.
What Ana learned
- Put each invariant in exactly one aggregate, and let membership be decided by what must be true at the same instant.
- Keep aggregates small and reference neighbours by identity so households can be loaded, saved and one day deployed apart.
- Match the modelling to the subdomain: a rich model for the core, a transaction script for the plain.
…but
The credit check now belongs to Billing, and Billing has to find out that an order was placed. Priya writes billingService.checkCredit(order) inside the Ordering handler, and Kofi, reading the pull request, puts down his coffee.
Chapter 11 · Act IIISomething Happened#
The problem
Priya's line is billingService.checkCredit(order), called from inside Ordering's PlaceOrderHandler, and Kofi's objection is not that it is wrong today. In one deployable it works. The objection is what it says about the border. Ordering now decides when Billing runs a credit check, knows that Billing has one, and passes Billing its own Order object, which Billing will start reading fields from. Within a month Billing will ask for a field, Ordering will add it, and the two countries will have merged through a side door.
There is a second, quieter problem. Fulfilment also needs to know an order was placed, and so does the promotions engine, and so will the loyalty scheme Lena mentioned yesterday. If Ordering calls each of them, placing an order gets slower and more fragile with every new listener, and Ordering's handler becomes a list of everyone else's business.
"Ordering doesn't need to tell Billing to check credit," Kofi says. "Ordering needs to say what happened. Billing decides what that means for Billing."
The idea
A domain event is a record that something happened in the business, named in the past tense in the context's own language: OrderPlaced, OrderLineAdded, OrderCancelled. The aggregate raises it as a fact; it does not know or care who listens. In the atlas it is a headline in the local paper: printed once, read by whoever subscribes. Contrast it with a command, PlaceOrder, which is an instruction to one named recipient that may be refused. Priya's checkCredit was a command in disguise, issued across a border.
| Command | Domain event | Integration event | |
|---|---|---|---|
| Tense and voice | Imperative: PlaceOrder | Past: OrderPlaced | Past: OrderPlaced (public form) |
| Addressed to | One handler, which may refuse | Anyone inside the context | Anyone outside the context |
| Language | The context's own | The context's own, internal detail allowed | The published language: stable, versioned, minimal |
| Changes when | The use case changes | The model changes | Only with consumers' notice |
The third column is the one teams miss. The OrderPlaced that Ordering's aggregate raises carries whatever the internal model finds convenient, and it changes whenever the model does. Publishing that to Billing would make Billing a conformist to Ordering's internals: chapter 7's accidental treaty. So a context keeps two kinds of event. Internal domain events stay inside; at the border a translator turns the ones other contexts care about into an integration event in the published language, the foreign-language edition of the paper, whose shape is a contract.
// Inside Ordering: the domain event, internal shape, free to change with the model.
public record OrderPlaced(OrderId id, BuyerId buyer, List<OrderLine> lines, Money total, Instant at) {}
// At the border: the integration event, published language v1. Stable. Minimal.
public record OrderPlacedV1(String orderId, String buyerId, long totalMinorUnits, String currency, String placedAt) {}
// The anti-corruption layer, outbound direction: translate, never expose.
public final class OrderingPublisher {
private final EventBus bus;
public OrderingPublisher(EventBus bus) { this.bus = bus; }
public void on(OrderPlaced e) {
bus.publish("ordering.order-placed.v1", new OrderPlacedV1(
e.id().value(), e.buyer().value(),
e.total().minorUnits(), e.total().currency().getCurrencyCode(),
e.at().toString()));
}
}
Billing has the mirror image: an inbound anti-corruption layer that reads OrderPlacedV1 and turns it into Billing's own CreditCheckRequested or a call on Billing's own aggregate. The customs house works in both directions, and neither context ever holds the other's classes.
How much should the integration event carry? Martin Fowler's distinction is the useful one. Event notification carries almost nothing, an identifier and a type, and consumers call back for details: small, always fresh, but every consumer now depends on the producer being up, which reintroduces the call chain. Event-carried state transfer carries the state consumers need, so they can keep their own copy and never call back: decoupled and resilient, at the cost of larger messages and copies that lag. Bramble's OrderPlacedV1 carries the total and the buyer because Billing and Fulfilment need exactly those; it does not carry the whole order, because nobody asked for it.
The published language is a model in its own right. It should be designed for consumers, versioned in its name, and changed only additively until every consumer has moved. That discipline is chapter 15's subject; the point here is that it exists at all.
Remember it as: inside, say what happened in your own words; at the border, print the foreign edition, and let each reader translate it into theirs.
Priya deletes billingService.checkCredit(order). Ordering raises OrderPlaced, the publisher prints the foreign edition, and Billing decides for itself what a placed order means for credit. Adding the loyalty scheme next month will not touch Ordering at all.
What Ana learned
- Raise events for what happened; never command another context across a border.
- Keep internal events internal, and publish a separate, versioned, minimal language at the border.
- Carry the state consumers need, and treat any consumer that calls back as a call chain you chose.
…but
The account page shows an order's status, its payment, its parcel and its loyalty points on one screen, and now that they live in four contexts the page makes four calls and takes two seconds. Lena, refreshing it, asks why a page that shows what already happened has to be slow.
Chapter 12 · Act IIITwo Models Are Cheaper Than One#
The problem
The account page is the most visited screen at Bramble after the checkout, and it is now the slowest. It shows, for each recent order, the status, the amount paid, the parcel's progress and the loyalty points earned. Those four facts live in four contexts, so the page controller calls four interfaces and stitches the answers together, and on a bad afternoon one of the four is slow and the page is slow with it.
Priya's first fix is a cache. It helps until an order is cancelled and the page shows it as dispatched for ten minutes. Her second fix is to ask Ordering's aggregate to expose a method that gathers everything, which Kofi vetoes: the aggregate is shaped for enforcing rules when an order changes, not for painting a screen, and every field added for the screen makes the aggregate heavier for the rules.
"You are asking one model to do two jobs," he says. "Take a payment correctly, and show a list quickly. They want opposite shapes."
The idea
CQRS, Command Query Responsibility Segregation, is the observation that the model that is best for changing state is rarely the model that is best for reading it, so you can keep two. The write side is the aggregate from chapter 10: small, rule-enforcing, loaded by identity, one per transaction. The read side is a read model: a flat, denormalised table or document shaped exactly like the screen that reads it, with no rules and no joins, updated whenever the write side announces a change.
The atlas needs a smaller metaphor here, and it lives inside one country. Think of a restaurant: order slips go one way, to the kitchen, where the rules are enforced; the menu board is what customers read, and it is rewritten when the kitchen changes something. Nobody reads the kitchen's tickets to find out what is for lunch.
The updater is called a projection: it listens to events and projects them onto the read model's shape. Here are the two sides, side by side, for one event:
// Write side: enforce the rule, record the fact. Knows nothing about screens.
public final class CancelOrderHandler {
private final OrderRepository orders;
public CancelOrderHandler(OrderRepository orders) { this.orders = orders; }
public void handle(CancelOrder cmd) {
Order order = orders.load(cmd.orderId());
order.cancel(cmd.reason()); // may refuse: dispatched orders cannot be cancelled
orders.save(order); // raises OrderCancelled
}
}
// Read side: keep the screen's row up to date. Knows nothing about rules.
public final class AccountPageProjection {
private final AccountPageRows rows; // one flat row per order, indexed by buyer
public AccountPageProjection(AccountPageRows rows) { this.rows = rows; }
public void on(OrderPlaced e) { rows.insert(e.id(), e.buyer(), "Placed", e.total()); }
public void on(OrderCancelled e) { rows.setStatus(e.id(), "Cancelled"); }
public void on(PaymentTakenV1 e) { rows.setPaid(e.orderId(), e.amountMinorUnits()); }
public void on(ParcelDispatchedV1 e) { rows.setParcel(e.orderId(), e.trackingRef()); }
}
Notice what the projection is allowed to do that the aggregate never could: consume events from other contexts. Payment and parcel status arrive as published-language events from Billing and Fulfilment, and the projection folds them into the same row. The account page becomes one indexed query against one table that Bramble's own context owns. The four calls are gone, and so is the fear of a slow neighbour, because the neighbours are no longer on the page's path.
The price is honesty about time. The read model lags the write model by however long the event takes to arrive, usually milliseconds, occasionally longer. A user who cancels and immediately refreshes may see "Placed" for a moment. Chapter 16 will show that the business almost always already tolerates this; the design must just not lie about it, which means the read model is never used to make a decision the aggregate should make. Checking "is this order still cancellable" against the account page row is how you cancel a dispatched parcel.
Two clarifications save teams a year each. CQRS is not event sourcing. CQRS says keep two models; event sourcing (next chapter) says store events instead of state. You can have either without the other, and most systems that need CQRS do not need event sourcing:
| CQRS | Event sourcing | |
|---|---|---|
| Decides | How many models to keep (two: write and read) | How to store the write model (as events, not state) |
| Needs the other? | No: the read model can be fed from plain domain events or from table changes | No: an event-sourced aggregate can be read directly, though it usually gets read models |
| Typical reason | A screen that joins many things, or reads that dwarf writes | An audit trail, temporal queries, or replaying history to build new views |
And CQRS is not for everywhere. A context whose screens mirror its aggregates gains nothing but a second table to keep in step. Bramble's returns desk reads the return it just wrote; it does not need a projection. The pattern pays where reads and writes want different shapes, or where a read spans contexts, or where reads outnumber writes by orders of magnitude.
Remember it as: the kitchen enforces the rules and the menu board is what you read; never decide what to cook by reading the board.
The account page loads in forty milliseconds. Priya deletes the cache.
What Ana learned
- Keep the write model for rules and a separate read model for each screen that needs a different shape.
- Feed read models from events, including other contexts', and never let a handler decide from one.
- Use CQRS where reads and writes want different shapes; elsewhere one model is cheaper than two.
…but
Finance asks a question the read model cannot answer and the aggregate cannot either: "What was this order's total on the third, before the discount was applied, and who changed it?" The database holds the order as it is now. Nobody kept the story of how it got there.
Chapter 13 · Act IIIThe Ledger Never Lies#
The problem
Finance's question is not idle. A corporate customer disputes an invoice: the order total they were quoted on the third differs from the total charged on the fifth, and they want to know what changed and who changed it. Bramble's orders table has one row per order, holding the current total. The audit log, added in year two, records that "order 88213 was updated" on the fourth by a support agent. It does not say from what to what.
Priya proposes an order_history table written by a trigger. Kofi asks how it would handle the promotion that was applied and then withdrawn, or the line that was added, priced, and removed within one minute, or the question finance will ask next month, which nobody can predict. A history table records the columns you thought to record. It cannot answer a question you did not anticipate.
"Accountants solved this three hundred years ago," he says. "They don't store a balance. They store every transaction, and the balance is arithmetic."
The idea
Event sourcing stores an aggregate not as its current state but as the sequence of domain events that produced it. The orders row is replaced by a stream: OrderPlaced, LineAdded, PromotionApplied, PromotionWithdrawn, LineRemoved, OrderPaid. To load the order you replay the stream from the beginning, applying each event to an empty aggregate, and the current state falls out. Nothing is ever updated or deleted; a correction is a new event. In the atlas it is the accountant's ledger: never edited, and any summary card, the balance or the account page row, is written from it.
Finance's question becomes trivial: replay to the third, read the total; replay to the fifth, read it again; the events in between say what changed and carry who did it. Better, questions nobody has asked yet can be answered later, because the facts were kept rather than a summary of them. New read models can be built from the whole history, and a corrupted projection is rebuilt by replaying. That is also the honest limit of the pattern: it is a storage decision for the write model of one aggregate inside one context. It is not a message bus, not a way to integrate contexts, and not a substitute for chapter 11's published language.
public final class Order {
private OrderId id; private OrderStatus status; private Money total = Money.ZERO;
private final List<Object> pending = new ArrayList<>();
public static Order replay(List<Object> history) {
var o = new Order();
history.forEach(o::apply); // rebuild state; no rules run during replay
return o;
}
public void cancel(String reason) { // decide: rules run here, on current state
if (status == OrderStatus.DISPATCHED) throw new DomainException("already dispatched");
raise(new OrderCancelled(id, reason, Instant.now()));
}
private void raise(Object e) { apply(e); pending.add(e); }
private void apply(Object e) {
switch (e) {
case OrderPlaced p -> { id = p.id(); status = OrderStatus.PLACED; total = p.total(); }
case OrderCancelled c -> status = OrderStatus.CANCELLED;
case LineAddedV1 old -> apply(Upcasters.toV2(old)); // old shape, translated on read
case LineAddedV2 l -> total = total.plus(l.subtotal());
default -> {}
}
}
public List<Object> pendingEvents() { return List.copyOf(pending); }
}
The costs are real and specific, and each has a standard answer.
Long streams. An order has thirty events; a loyalty account may have thirty thousand. A snapshot stores the state at event N so a load replays only what came after. Add snapshots when a measured load is slow, not before; premature snapshots complicate every schema change for a problem you do not have.
Events change shape. LineAdded v1 stored a price as a float; v2 stores Money. You cannot rewrite history, so you translate on read: an upcaster converts old events to the current shape as they are replayed, which is the LineAddedV1 branch above. Every event type must be versioned from day one, because the first schema change is the first time you discover whether it was.
Personal data. A ledger that is never edited collides with a regulation that says a person may demand erasure. Storing a buyer's name and address in an immutable event is a design error. Keep personal data in a separate, deletable store keyed by identifier and reference it from events, or encrypt each person's fields with a key you can destroy (crypto-shredding). Decide this before the first event is written; retrofitting it means rewriting the ledger you promised never to rewrite.
Reading. You cannot SELECT across event streams. Event sourcing almost always comes with the read models of chapter 12, and every screen is a projection. That is where CQRS and event sourcing became confused: they arrive together, but the first is about two models and the second about storing one.
The pattern earns its cost where history is the product: finance, legal, anything audited, anything where "how did we get here" is a business question. Bramble's Billing context qualifies. Its returns desk does not, and event-sourcing a supporting subdomain because it seemed elegant is one of the more expensive ways to spend a year.
Remember it as: the ledger is never edited; the balance is arithmetic, the summary card is a projection, and the ledger stays inside one country.
Finance gets its answer in an afternoon, from Billing's new invoice stream. Ordering keeps its plain rows, because nobody has ever asked how an order's delivery address got to be what it is.
What Ana learned
- Event-source only where history is the product, and keep it inside one context.
- Version every event from the first one, translate old shapes on read, and keep personal data out of the ledger.
- Add snapshots when a measured load is slow, not because the pattern has them.
…but
The invoice stream is beautiful and the integration event is published, and then a corporate customer rings to say they have been charged twice for one order, because Billing's payment call was retried after a timeout and the provider took the money both times.
Chapter 14 · Act IVThe Payment Taken Twice#
The problem
The corporate customer was charged twice for order 91406, and the trace tells the story in four lines. Billing called the payment provider. The provider took the money and, forty seconds later, returned success. Billing's HTTP client had given up at thirty seconds, thrown a timeout, and been retried by the resilience library somebody added after the night in chapter 3. The retry succeeded too. Two charges, one order, one very reasonable complaint.
Ana's first reaction is that the retry was wrong. Kofi's is that the retry was right and the request was wrong: it asked for "a payment of forty pounds" twice, and there was no way for the provider to know both requests meant the same payment. The timeout was also right. The mistake was to believe that a timeout means the call did not happen. A timeout means you do not know.
Then the second trace arrives, from the same afternoon. The loyalty service was slow, Billing's calls to it queued up, the thread pool filled, and for six minutes Billing could not take payments at all, because the same pool served both. Two different failures, one lesson: the network is a phone line, not a method call, and Bramble has been dialling without a plan for the line going dead.
The idea
Every call across a border is one of two things. A synchronous call is a phone call: both parties must be awake, the caller waits, and if the line drops mid-sentence the caller does not know what the other side heard. An asynchronous message is a letter: the sender posts it and carries on, the receiver reads it when convenient, and neither needs the other awake at the same moment. Act IV is about the cost of each and when to use which; this chapter is about making the phone calls Bramble cannot avoid survive a bad line.
Five tools, each answering one failure:
A timeout on every outbound call, always, because a call with no timeout holds a thread until the other side deigns to answer. The value should be shorter than the caller's own timeout, or the caller times out first and retries into a call that is still running. Pair it with a decision: what does the user see when it fires?
Retries, but only for calls that are safe to repeat, and with backoff and a cap, because chapter 3 showed what unbounded retries do to a struggling service. The retry is not the dangerous part. Repeating a request whose first attempt may have succeeded is.
An idempotency key makes repetition safe. The caller generates a unique key per intended operation (not per attempt), sends it with every attempt, and the receiver remembers keys it has processed and returns the original result for a repeat. It is the reference number on the letter: two letters with the same number are one instruction. Payment providers accept one; so should every handler Bramble exposes.
public final class TakePaymentHandler {
private final ProcessedKeys processed; // durable: same store as the payment, same transaction
private final PaymentGateway gateway;
public TakePaymentHandler(ProcessedKeys processed, PaymentGateway gateway) {
this.processed = processed; this.gateway = gateway;
}
public PaymentResult handle(TakePayment cmd) {
var prior = processed.find(cmd.idempotencyKey());
if (prior.isPresent()) return prior.get(); // a repeat: same answer, no new charge
PaymentResult result = gateway.take(cmd.orderId(), cmd.amount(), cmd.idempotencyKey());
processed.record(cmd.idempotencyKey(), result); // remember before acknowledging
return result;
}
}
A circuit breaker stops calling a dependency that is failing. After a run of errors or timeouts the breaker opens and calls fail fast for a while, which gives the dependency room to recover and the caller a chance to degrade gracefully instead of queueing. It is the fuse. A bulkhead isolates dependencies from each other: a separate thread pool or connection limit per dependency, so a slow loyalty service cannot starve payments. Watertight compartments, so one flooded room does not sink the ship.
Now the honest question: which calls should be phone calls at all? Taking a payment is genuinely synchronous; the customer is waiting for a yes or no. Telling Ordering that a payment was taken is not; a letter will do, and a letter cannot time out. Most of Bramble's remaining synchronous calls turn out to be notifications that someone wrote as calls because calls were familiar. Each one moved to a message removes a dependency from the critical path and a row from chapter 3's availability table.
One disguise to watch for. Putting a request on a queue and blocking until a reply arrives on another queue is request/response with extra steps: the caller still waits, still needs the other side awake, still needs a timeout, and now also needs a broker. It is a phone call routed through the post office. If you need an answer now, make the call and protect it; if you do not, send the letter and stop waiting.
Remember it as: a timeout means you don't know; give every letter a reference number, put a fuse on every line; stop phoning people you could write to.
The provider is called with an idempotency key from that afternoon on. The second charge is refunded, and the loyalty service gets a pool of its own.
What Ana learned
- Treat a timeout as "I don't know", and make every repeatable request carry a key so repeating it is safe.
- Give every dependency its own timeout, breaker and pool, and decide what the user sees on failure.
- Ask of every synchronous call whether it needs an answer now; if not, make it a message.
…but
Ordering now hears about payments by letter, and one afternoon a letter never arrives: an order is saved, the transaction commits, and the process is killed before the OrderPlaced event is published. Billing never hears of the order. Nobody is paid.
Chapter 15 · Act IVThe Message That Never Left#
The problem
Order 92110 exists in Ordering's database and nowhere else. The handler saved the aggregate, the transaction committed, and then, in the two milliseconds before bus.publish(orderPlacedV1) ran, the container was recycled by a routine deployment. The order is real. The event announcing it was never sent. Billing never took payment, Fulfilment never picked, and the customer's confirmation email describes a parcel nobody will pack.
Priya finds the code and sees nothing wrong with it: save, then publish. Kofi points out that there are two writes to two systems, the database and the broker, and no transaction spans both. Whatever order you put them in, the process can die between them. Publish first and the database may then refuse the save, so Billing charges for an order that does not exist. Save first and the publish may never happen, which is 92110.
"This is the dual write," he says, "and every team writes it once."
The idea
The dual write problem is that a single business action needs to update two systems that cannot share a transaction, and any sequence of two separate writes has a gap in which one has happened and the other has not.
The standard cure is the transactional outbox. Instead of publishing to the broker, the handler writes the event into an outbox table in the same database transaction as the aggregate. Either both rows exist or neither does. A separate relay reads the outbox and publishes to the broker, marking rows as sent. In the atlas it is the letter tray by the door, filled in the same sitting as the diary entry; the postman collects later, and a letter in the tray is as certain as the diary.
public final class PlaceOrderHandler {
private final OrderRepository orders;
private final Outbox outbox; // a table in the SAME database as orders
private final OrderingPublisher translate; // domain event -> published language
public PlaceOrderHandler(OrderRepository orders, Outbox outbox, OrderingPublisher translate) {
this.orders = orders; this.outbox = outbox; this.translate = translate;
}
@Transactional // one commit covers both rows
public OrderId handle(PlaceOrder cmd) {
Order order = Order.place(cmd.buyerId(), cmd.lines());
orders.save(order);
for (Object e : order.pendingEvents()) {
outbox.append(translate.toIntegrationEvent(e)); // OrderPlacedV1 as a row, not a publish
}
return order.id();
}
}
// A relay, in another thread or process, polls outbox WHERE sent_at IS NULL,
// publishes each row to the broker, then marks it sent. It may publish twice; see below.
The relay can be a poller, or it can be change data capture: a tool that reads the database's own transaction log and publishes each committed outbox row, so nothing polls and nothing is missed. That is reading the diary to write the letters, and Debezium is the usual tool.
Notice the relay's weakness. It may publish a row and die before marking it sent, then publish it again on restart. Delivery is therefore at-least-once, and every consumer must be an idempotent consumer: it keeps an inbox, a record of event identifiers it has already handled, and ignores repeats. The mailroom log of letters already opened. This is not optional and it is not the broker's job; exactly-once delivery across systems is a marketing phrase, and de-duplication belongs at the receiver.
| Outbox | Change data capture | Inbox | |
|---|---|---|---|
| Solves | Publish atomically with the write | Move committed rows out without polling | Handle each event once despite redelivery |
| Lives in | Producer's database and a relay | Producer's database log and a connector | Consumer's database |
| Guarantee | Event exists iff the write committed | Every committed row is published | Repeats are ignored |
Ordering also matters, and only within limits. A broker can promise that messages with the same key (say, the order id) arrive in the order they were sent; it cannot promise a global order across all keys, and consumers that assume OrderPlaced for order A precedes OrderCancelled for order B will be wrong. Partition by the aggregate id, and design consumers to tolerate an event for the same key arriving out of order or late.
Finally, the letters themselves change shape over time, and the published language is a contract with people you cannot redeploy. Three disciplines keep it honest. Consumers practise tolerant reader: ignore fields you do not know, do not fail on extras. Producers change additively: add fields, never remove or rename in place. When a breaking change is unavoidable, use expand and contract: publish the new version alongside the old, migrate consumers one at a time, retire the old only when its last reader is gone. And make the contract testable with consumer-driven contracts: each consumer records the fields and shapes it relies on, and the producer's build runs those expectations before shipping, so Billing's assumptions about OrderPlacedV1 are checked in Ordering's pipeline, not in production.
Remember it as: write the letter in the same sitting as the diary; expect every letter to arrive twice; never change the form without telling every reader first.
92110 is replayed from the outbox after the fix, Billing takes the payment, and a query for "orders with no outbox row" finds two more from the previous month.
What Ana learned
- Never write to a database and a broker in one handler without an outbox between them.
- Give every consumer an inbox and design every side effect to survive a repeat.
- Change published events additively, expand and contract for anything breaking, and let consumers' contracts run in the producer's build.
…but
The events arrive, exactly once in effect, and Bramble discovers a new kind of failure: an order is placed, stock is reserved, and then the payment fails. The order is half-done across three contexts, and no single transaction can undo it.
Chapter 16 · Act IVWho Says the Order Is Done?#
The problem
Order 93877 is placed in Ordering, its two items are reserved in Fulfilment, and then Billing's payment fails: the card is declined. In the monolith this was one transaction that rolled back. Now it is three contexts, three databases and three facts, one of which is now wrong. The stock stays reserved. Two other customers see the pens as out of stock. Nobody's code is responsible for undoing anything, because nobody's code knows the whole story.
Priya's proposal is a distributed transaction: lock all three, commit or roll back together. Kofi says two things. Technically, two-phase commit across services holds locks in three databases for the duration of a network conversation, and any one of them being slow freezes all three; it is the call chain of chapter 3 with locks. And there is a bigger objection: it does not match how the business already works.
He walks Ana over to the warehouse lead and asks what happens today when a customer's card is declined after the pens are already in the tote. "We put them back," she says. "And if someone else bought the last one in the meantime?" "Then we email the first customer and offer the blue ones, or a refund." Nobody in the warehouse has heard of a lock.
The idea
A saga is a business process that spans several aggregates or contexts, implemented as a sequence of local transactions, each of which publishes an event that triggers the next, with a compensating action for each step that must be undone if a later step fails. There is no global rollback. There is reserve stock, and if payment then fails, release stock. In the atlas it is the travel agent: flight, hotel and car are booked one at a time, each with a cancellation rule, and if the car falls through the agent cancels the hotel rather than pretending the trip never began.
There are two ways to coordinate the steps. In choreography each context reacts to the previous context's event and publishes its own: dancers watching each other, no one in charge. It is simple and loosely coupled for three steps and becomes impossible to follow at eight, when nobody can say what state an order is in without reading five inboxes. In orchestration a process manager (the saga's conductor) sends each step a command, waits for the reply event, and decides what comes next, holding the state of the process explicitly. It is easier to read and to see, and it risks growing into a god service that knows every context's business.
public final class OrderSaga { // an orchestrating process manager
private SagaState state = SagaState.STARTED;
private final Commands send;
public OrderSaga(Commands send) { this.send = send; }
public void on(OrderPlaced e) { state = SagaState.RESERVING; send.to("fulfilment", new ReserveStock(e.id(), e.lines())); }
public void on(StockReservedV1 e) { state = SagaState.PAYING; send.to("billing", new TakePayment(e.orderId(), e.total(), e.orderId())); }
public void on(PaymentTakenV1 e) { state = SagaState.COMPLETED; send.to("ordering", new ConfirmOrder(e.orderId())); }
public void on(PaymentFailedV1 e) { // compensate the step that already happened
state = SagaState.COMPENSATING;
send.to("fulfilment", new ReleaseStock(e.orderId()));
}
public void on(StockReleasedV1 e) { state = SagaState.FAILED; send.to("ordering", new RejectOrder(e.orderId(), "payment declined")); }
}
Two consequences follow, and both are business conversations before they are technical ones.
First, between reserve and release the world sees an intermediate state: the pens look sold. A saga has no isolation. If that is unacceptable, the standard tool is a semantic lock: the reservation is marked pending rather than sold, and other readers treat pending specially (show "one left, reserved for another shopper" rather than "sold out"). It is a business rule, visible in the language, not a database lock.
Second, compensation is not undo. Releasing stock after someone else bought the last one cannot give the first customer their pens. The warehouse lead already had the answer: offer an alternative or a refund. That is eventual consistency as the business practises it. Pat Helland's phrase is that the world runs on memories, guesses and apologies: a system remembers what it promised, guesses that it can keep the promise, and apologises when it cannot. Gregor Hohpe's version is that Starbucks does not use two-phase commit: it writes your name on a cup and, if the machine breaks, gives you your money back. The design question is not "how do we prevent the intermediate state" but "what does the business do when it happens", and the answer is a rule the domain expert already knows, expressed as a state and an event, not a lock. The post takes a day, and the business already works that way.
Remember it as: a saga is a travel agent: book one thing at a time, know how to cancel each, and when the last falls through, apologise well.
Where the rule genuinely cannot tolerate an intermediate state, and the business says so, the honest options are to move the two steps into one aggregate (chapter 10's boundary question again) or to accept the cost of a synchronous, protected call for that one step. Both are choices with a price tag. A distributed lock across services is neither; it is the price tag without the choice.
What Ana learned
- Replace the distributed transaction with a sequence of local ones, each with a compensation the business already recognises.
- Choreograph short processes, orchestrate long ones, and keep the orchestrator ignorant of every context's rules.
- Ask the business what it does in the intermediate state; the answer is a rule, and the rule is a state and an event.
…but
The order flow now completes or apologises on its own. Then Lena asks for a weekly report: revenue by product category, by promotion, by courier, with returns netted off. The data lives in five contexts, and the last person who tried to join them was the promotions job from chapter 1.
Chapter 17 · Act IVThe Report Nobody Could Run#
The problem
Lena's weekly report used to be one SQL query against one database, written by Ana in year one and run every Monday. It joined orders to products to promotions to couriers to returns, and it was the most important query in the company because it was how Lena decided what to buy next.
Now orders are in Ordering, prices in Billing, categories in Catalogue, couriers in Fulfilment, promotions in the new Promotions context and returns in Support, each in its own schema with its own credentials. The query cannot be written. Priya's first attempt calls five interfaces in a loop, one call per order per week, and is killed after forty minutes. Her second attempt asks each context for a read-only database user "just for reporting", which is chapter 8's shared database reintroduced by the finance department.
Meanwhile the mobile team has a related complaint. Their order-detail screen needs status, payment, parcel and points, exactly the account page's problem from chapter 12, and their answer was to make four calls from the phone. On a train it takes eleven seconds.
Two problems, one shape: somebody needs a picture that spans several countries, and there is no world government.
The idea
There are three honest ways to assemble a view across contexts, and the trick is to match each to how fresh the data must be, how much of it there is, and who is asking.
API composition is a component that calls several contexts' interfaces and stitches the result, the concierge who phones three offices so the guest makes one call. It gives fresh data with no copies, and it inherits every cost of a synchronous call chain: latency adds up, availability multiplies down, and joining more than a page of results across services is hopeless. It is right for a single screen showing a single order, with timeouts and fallbacks per call. It is wrong for a report.
A backend for frontend (BFF) is API composition owned by a particular client. The mobile team gets a small service that makes the four calls server-side, near the contexts, and returns one payload shaped for the phone. The web team may have its own. Each is owned by the team that owns the screen, so it can change with the screen and nobody else's. Where many clients need flexible composition, GraphQL federation offers a schema stitched from each context's subgraph, at the price of an extra layer that can hide the same call chain behind a query language.
A CDC-fed read store is the pattern of chapter 12 applied across borders: each context publishes its facts (as integration events or through change data capture from its outbox), and a consumer builds a store shaped for the questions being asked. For Lena's report that consumer is an analytical warehouse, fed continuously, holding a copy of the facts from all six contexts in a schema designed for joining and aggregating. In the atlas, it is the national statistics office: it receives copies of every country's ledgers and produces the figures; it does not reach into their records offices.
| API composition | Backend for frontend | CDC-fed read store / warehouse | |
|---|---|---|---|
| Freshness | Live | Live | Seconds to minutes behind |
| Suits | One screen, one entity, few calls | One client's screens, owned by that client's team | Reports, lists, search, anything that joins many rows |
| Cost | Latency and availability of every call | The same, plus one more service to own | Copies, a pipeline, and honesty about lag |
| Breaks when | A context is slow or down | Same, but only that client suffers | The pipeline stalls and nobody notices |
Who owns the warehouse's data is its own political question. The traditional answer is a central data team that ingests everything and owns every pipeline, and it fails the way every central team fails: it becomes the bottleneck, it does not understand any one context's language, and when Billing changes a field the pipeline breaks a week later in a dashboard nobody attributes to Billing. Data mesh is the 2019 proposal that turns this around: each context owns and publishes its analytical data as a product, in a documented schema, with quality guarantees, and a shared platform makes publishing and joining cheap. It is the bounded-context idea applied to analytics, and it comes with the same warning: without a platform team making it easy, "you own your data product" is a chore nobody does.
One more diagnosis, because the finance department's request for read-only database users was not a stupid idea, just an old one. Two patterns from the 2000s keep returning wearing new clothes. The enterprise service bus put logic in the pipes, transforming and routing messages centrally, until the bus knew every context's business and every change went through the bus team; smart endpoints and dumb pipes was the 2014 reaction. The canonical data model defined one company-wide schema for "order" so everyone could integrate; it is exactly what a warehouse must not impose upstream, or every context becomes a conformist to the reporting schema. The warehouse may have its own canonical shape. It builds it from the published facts; it does not push it back.
Remember it as: phone three offices for one guest's answer; for the census, let every country send copies to the statistics office and never open their filing cabinets.
Lena's report runs on Monday from the warehouse, four minutes behind reality, and the mobile screen loads in under a second through its BFF.
What Ana learned
- Compose live calls for one screen's one entity; build a read store or warehouse for anything that lists, sorts or aggregates.
- Give each client team its own backend for frontend, owned with the screen.
- Feed the warehouse from published facts, never from other contexts' tables, and never push its schema upstream.
…but
Six contexts, one warehouse, and an org chart that still has three teams named "frontend", "backend" and "platform". The promotions engine, Bramble's core, is owned by nobody in particular and changed by everybody, and it is about to show.
Chapter 18 · Act VYou Ship Your Org Chart#
The problem
Bramble now has eleven developers, six bounded contexts, a warehouse, a BFF and three teams: Frontend, Backend and Platform. It made sense when the system was one monolith with three layers. It makes no sense now, and the promotions incident proves it.
A bundle promotion goes wrong: buy two notebooks, get a pen free, except the free pen is being applied to every order containing a pen. The Frontend team changed how bundles are displayed. The Backend team changed the rule engine last sprint. Platform changed the event schema between them. Each change was correct alone. Nobody owns Promotions, so nobody reviewed the three together, and the person who understands the bundle rules best is Lena, who is not on any team.
Meanwhile the new marketplace, a Seller context that lets third parties list stationery on Bramble, was built by a squad of two who now also own the warehouse pipeline, the BFF and, since last month, the payment provider integration, because "they know Kafka". They have nine services between them and cannot describe what any of them did last week.
Ana draws the context map from chapter 7 and, beside it, the org chart. They have nothing in common.
The idea
Melvin Conway's 1968 observation, which chapter 7 introduced, is that a system's structure copies the communication structure of the organisation that builds it. Three teams named after layers will produce a system whose real seams run between layers, whatever the context map says, because the layers are where people talk. Conway's law cannot be repealed. It can be used: if you want the system shaped like the context map, shape the teams like it. This is the inverse Conway manoeuvre, and it is the most powerful architectural tool most architects never touch, because it lives in a spreadsheet owned by someone else.
Team Topologies gives the manoeuvre a vocabulary. A stream-aligned team owns a flow of business value end to end, which at Bramble means a bounded context or a small group of them: Ordering, or Shopping and Promotions together. In the atlas, it is a country's own government. A platform team builds the roads and the post: the deployment pipeline, the event infrastructure, the observability, offered as a product so that stream-aligned teams do not each build their own. An enabling team is a small group of travelling advisers who help a stream-aligned team learn something (Event Storming, contract testing) and then leave. A complicated-subsystem team owns a piece that needs rare expertise, the specialist bureau; Bramble does not have one and should not invent one.
The sizing rule is cognitive load. A team can own what it can hold in its collective head: the rules, the code, the operations, the incidents. A stream-aligned team of five that owns two contexts is healthy. A squad of two that owns nine services and the warehouse is not a team; it is an outage schedule. When a context is more than a team can hold, the answer is not more services but a smaller boundary or a second team. When nine services are less than a team's worth of complexity, the answer is fewer services.
Redrawing the map with teams around the contexts shows the mismatch instantly. Fulfilment has no owner. The Orders team owns Ordering and Billing, which is fine because the two share a partnership-grade relationship and one team can hold both. Promotions, the core, sits with Shopping under one team whose product owner is, at last, Lena's deputy. The Marketplace squad owns one context, not nine services, and the payment integration goes back to whoever owns Billing.
Ownership is the operational form of all this: every context has exactly one team that reviews its changes, carries its pager and decides its roadmap. A service with no owner is not free; it is owned by whoever was last paged for it, at the worst possible time.
There is a trap on the other side of this chapter that deserves naming because it is the one most often walked into in good faith. Teams sometimes reach for microservices to solve a people problem: two groups that do not get on, a manager who wants a fiefdom, a dependency on a slow team that a new service would route around. The service boundary then encodes the argument, and the argument outlives the people. Conway's law works in that direction too: build a service to avoid talking to someone and you have built a treaty that says you never will.
Remember it as: the map ends up looking like the org chart, so draw the org chart like the map, and size each country to one team.
Bramble reorganises over a quarter, not a weekend. The three-layer teams become four, three of them named after what they own. Fulfilment gets a team on the day the warehouse lead joins it as product owner.
What Ana learned
- Align teams to bounded contexts, not layers, and give every context exactly one owning team.
- Size a team's ownership by what it can hold in its head, and reduce services before adding people.
- Fund a platform team so stream-aligned teams do not each build the roads.
…but
The Shopfront team inherits the promotions engine and discovers it is still the ball of mud from chapter 1, walled off behind an anti-corruption layer in chapter 7 and never actually replaced. Lena wants it rebuilt. Everyone remembers the last time somebody said "rewrite".
Chapter 19 · Act VStrangling the Monolith#
The problem
The promotions engine is eight thousand lines of the original monolith: percentage discounts, bundles, loyalty multipliers, free shipping thresholds, a "VIP" flag nobody can explain, and a table of exceptions added one Black Friday at a time. It is Bramble's core subdomain. It is also the code nobody wants to touch, and it still reads the old customer table directly, which is the last reason that table exists.
The Shopfront team's first plan is the natural one: build Promotions v2 properly, with aggregates and events and a clean model, run it in parallel for a while, then switch. Six months. Kofi asks what happens in month four when Lena needs a new bundle type for the autumn range. "We'd add it to both." "And in month five, when v2 slips?" Nobody answers, because everyone has lived through a rewrite that slipped, and the room's memory is that the old system kept getting features until the new one was quietly abandoned.
"Don't replace it," Kofi says. "Surround it."
The idea
The strangler fig is a plant that grows around a host tree, sends roots down, and eventually stands on its own when the host has rotted away. Applied to software: put a facade in front of the old system, route one capability at a time through new code, and let the old system shrink until nothing routes to it. The old system is never switched off in one night; it is starved.
Three techniques do the actual work, and each has a place.
Strangler fig proper works at the edge: callers hit a facade, and the facade decides, per request or per capability, whether the old or the new code answers. It needs a seam, a place where the old system can be intercepted without modifying it, which is why the anti-corruption layer from chapter 7 was worth building early: it is the seam.
Branch by abstraction works inside a codebase that cannot be fronted so cleanly. Introduce an interface over the old capability, make all callers use the interface, write the new implementation behind it, switch the binding, delete the old. The new bridge is built beside the old before the signpost moves. It is how you strangle a module inside a monolith, where there is no network edge to put a facade on.
Parallel run is the safety net: for a period, send every request to both old and new, use the old answer, and log every disagreement. Two clerks do the same sum for a month until they agree. For a promotions engine, where a wrong answer costs money, the parallel run is what makes the switch a non-event.
| Strangler fig | Branch by abstraction | Parallel run | |
|---|---|---|---|
| Works at | The edge: a facade in front of the system | Inside the code: an interface over the capability | Alongside either: both implementations live |
| Needs | A seam callers can be routed through | All callers to go via the abstraction first | A way to compare answers and tolerate double cost |
| Gives you | Incremental replacement, capability by capability | Replacement inside a monolith with no network edge | Confidence, and a record of what the old one did |
The order in which capabilities move matters more than the technique. Move first what hurts most and changes most and is least entangled: percentage discounts change every week and touch nothing else, so they go first and deliver value in a fortnight. Move last what is most coupled, the VIP flag that reads the old customer table and three other things. Extracting the most coupled piece first is the most common way a strangling stalls: it takes months, delivers nothing visible, and the sponsor's patience is spent before anything easy has shipped.
Then the database. Code is strangled in weeks; data takes longer, and it is the part that decides whether the old system can actually die. Each capability that moves takes its data with it into the new schema, with a period of dual-reading or a one-way sync from old to new until the last reader of the old table is gone. Database decomposition is the hardest and last step, and "we'll leave the customer table shared for now" is how chapter 2's temporary arrangement becomes a decade. The strangling is finished when the old table has no readers, not when the old code has no callers.
And the plan the team started with, the big-bang rewrite, deserves its reputation. It freezes features or forks them, it delivers nothing until it delivers everything, and it is estimated by people who have not yet found the VIP flag. The fig is slower per capability and infinitely faster to first value.
Remember it as: don't replace the tree, surround it: route one branch at a time, compare the sap for a month, cut the last root at the database.
Percentage discounts move in two weeks. Bundles take two months and a parallel run that finds three long-standing bugs in the old engine. The VIP flag turns out to mean "Lena knows them", and becomes a loyalty tier in the new model. The old customer table's last reader is deleted eleven months after the first facade.
What Ana learned
- Surround the old system with a facade and move one capability at a time, easiest and most valuable first.
- Run old and new in parallel and read the differences; that log is the test.
- The migration ends when the old tables have no readers, not when the old code has no callers.
…but
Promotions v2 ships a change to its published event on a Thursday, and nothing breaks until Monday, when the warehouse shows every promotion at zero and nobody can trace which of six services stopped understanding which message.
Chapter 20 · Act VCan You See It?#
The problem
Promotions v2 renamed a field in PromotionAppliedV1 from discount to discountMinorUnits on Thursday, because the old name had been ambiguous about units. Its own tests passed. The warehouse consumer read a missing field as zero, the Billing consumer ignored it and kept charging full price, and the account-page projection crashed and stopped. Nothing paged anyone, because nothing was wrong with any single service. On Monday Lena's report showed a week of promotions at zero and the support queue filled with people charged full price.
Finding out which consumer broke took the morning, because a request's journey through six services exists only as six unrelated log files. Finding out why took the afternoon, because the end-to-end test suite that was supposed to catch this takes four hours, needs all six services deployed together in a staging environment, and had been red for so long that its failures were routinely ignored.
Ana realises that Bramble has built a system it cannot see, and tests it in the one way that reintroduces the lockstep it spent a year escaping.
The idea
Independent services need three things the monolith gave away for free: a way to know a change will not break a neighbour before shipping it, a way to follow one request across every border, and a way to ship each service on its own without a shared rehearsal. Each is a discipline, not a tool.
Contract tests replace the all-system suite for the question "will this break a consumer". A consumer writes down, as an executable expectation, exactly what it needs from a producer's event or interface: these fields, these types, this meaning. The producer's build runs every consumer's expectations before deploying. Thursday's rename would have failed Promotions' pipeline at ten in the morning, naming the warehouse as the consumer that depends on discount. The contract is the customs form both sides signed, and it is checked at the producer's border, where the change originates.
// In the warehouse consumer's repository: what it relies on. Published to the producer's build.
final class PromotionAppliedV1Contract {
@Override public String toString() { return "warehouse -> promotions.promotion-applied.v1"; }
void expectations(Contract c) {
c.field("orderId").isString().required();
c.field("discount").isInteger().required() // renaming this breaks THIS consumer
.meaning("minor units, positive means money off");
c.field("promotionCode").isString().required();
c.toleratesUnknownFields(); // tolerant reader, stated as a rule
}
}
// Promotions' pipeline: for each registered consumer contract, verify against the real event.
Tests as a whole take the shape of a honeycomb rather than the classic pyramid: many integration tests at each service's own boundary (does this service, alone, honour its contracts and its own rules against a real database), a solid layer of fast unit tests inside the domain, and very few, or no, tests that require the whole system to be deployed together. The all-system end-to-end suite is the trap: slow, flaky, always red, and, worst, it requires every service to be present in one environment at one version, which is the lockstep release in a test harness.
Correlation ids make a request visible across borders. Every inbound request or message either carries an id or is given one at the edge, and every log line, every outbound call and every published event carries it forward. It is the tracking number that follows one parcel through every office. Distributed tracing builds on it: each service records a span (started, ended, called what) tagged with the id, and a tracing backend draws the request as a tree across services. Thursday's incident becomes a query: show me a request that touched Promotions and the warehouse consumer, and here is the span where the field went missing.
Independent pipelines are the point of the whole exercise: each service has its own build, its own contract verification, its own deploy, and ships when it is ready. A shared release train, a fortnightly "integration window" or a staging environment where everything must be deployed together are all the same thing: the distributed monolith's release process, kept after its cause was removed.
Versioning is what makes independence survive change. The rule from chapter 15 applies to interfaces as well as events: change additively, tolerate unknown fields, and when a break is unavoidable, expand (publish v2 alongside v1), migrate consumers at their pace, then contract (retire v1 when its last consumer is gone, which the contracts tell you). A version number in a URL or a topic name is a fine label for this; it is not a substitute for it. Bumping to v2 and deleting v1 in one deploy is a breaking change with a number on it.
One paragraph on security, because it is the border control nobody budgets for. Services calling services must authenticate each other: mutual TLS or short-lived signed tokens carrying the caller's identity, issued by the platform, so that a compromised warehouse consumer cannot call Billing's refund endpoint. The platform team should make this the default so that stream-aligned teams never implement it themselves, and a request's identity should travel with its correlation id.
Remember it as: sign the customs form at the producer's border, put a tracking number on every parcel, and let every country ship on its own timetable.
Promotions restores discount, adds discountMinorUnits beside it, and the warehouse consumer's contract now runs in Promotions' build. The four-hour suite is deleted. Nobody misses it.
What Ana learned
- Verify consumers' contracts in the producer's build, and delete the all-system suite.
- Carry a correlation id on every request, message and log line, and trace across borders.
- Ship each service the moment its own pipeline is green; a shared release day is a distributed monolith.
…but
Bramble can now see itself, ship independently and change safely. Which is when Lena announces the subscription box, a new line of business, and Marcus returns with a proposal for it: twelve services, from day one.
Chapter 21 · Act VThe Consultant Who Said Microservices#
The problem
The subscription box is simple to describe: a curated bundle of stationery every month, at a fixed price, with a swap-in option. Lena wants it live by spring. Marcus's proposal is thorough: twelve services, from subscription and plan through curation, billing cycle, dunning, swap, pause and cancellation, each independently deployable with its own database, on the platform Bramble already has.
He is more careful than two years ago, and his argument is Tilkov's: Bramble now knows its domain, has a platform team, has practised the outbox and the saga, and a team that starts monolithic will reach across module boundaries by habit. Better to enforce the borders with a network from day one.
Ana asks for a week. She has two things to report first. The Seller and Catalogue services, split in year one, were merged back into one deployable this quarter: every change touched both, one team owned both, and their "independence" was two pipelines for one release. Nobody noticed except the on-call rota, which got shorter. And she has a spreadsheet of what the six current contexts cost to run as separate services, in pipelines, pagers, dashboards, contract suites and the platform team's time. It is not small.
The idea
Microservices have a price. Martin Fowler called it the microservices premium in 2015: the operational and cognitive cost of a distributed system, paid on every service whether or not it needed to be distributed. Pipelines, contracts, outboxes, inboxes, sagas, tracing, per-service on-call, and the platform to make it bearable. In the atlas, every new country needs an embassy, a customs house and translators. The premium buys independent deployability and team autonomy, and it is worth paying only when the system's complexity or the organisation's size makes that independence worth more than it costs.
The same year, two people who respect each other disagreed in public. Fowler's MonolithFirst: almost every successful microservice system he had seen started as a monolith that was later split, and almost every system built as microservices from scratch had ended in trouble, because you do not know where the boundaries are until you have lived in the domain. Stefan Tilkov's reply, "Don't start with a monolith": if you do know the domain, a monolith teaches your team habits that make splitting harder, and the boundaries you know should be enforced from the start. Both are right, under their premises, and the premise that matters is whether the boundaries are known.
The merged Seller and Catalogue services are the point in miniature. Their boundaries were guessed in year one, and the guess put a border through one language. Merging them was not a failure but the correction a design process makes when evidence arrives; a team that cannot merge services back has stopped designing.
Now the honest table:
| Modular monolith | Distributed monolith | Microservices | |
|---|---|---|---|
| Deploy unit | One, with enforced module borders | Several, released together | Several, released alone |
| Database ownership | One instance, one schema per module, one writer per table | Shared schema across services | One per service |
| Failure blast radius | The process; a bad module can take it down | Every service, via the chain | One service, if the rest degrade |
| Team autonomy | High within modules; one release cadence | Low: everyone waits for everyone | High: each team ships alone |
| Operational cost | One pipeline, one pager, one set of dashboards | Many of each, for no benefit | Many of each, and a platform to bear it |
| Migration path | Modules extract along their borders when a reason appears | Must be re-bounded before anything improves | Services merge when a border proves wrong |
The middle column is where Bramble spent chapter 3, and it is where a twelve-service subscription box run by a team of four would end up within a year, however good the platform: four people cannot own twelve pagers, and the borders between pause, swap and cancellation are not yet a language anyone at Bramble speaks.
The design review is on Friday. Ana brings the questions the book has been collecting, and asks Marcus each in turn.
From chapter 5: "Name the subdomain that customers choose you for, in one sentence, and say who on the team is working on it this quarter." Curation, he says; and it is one of twelve services, sharing a team of four with eleven others.
From chapter 6: "Inside this proposed boundary, name one word; does it mean exactly one thing everywhere inside, and something different just outside?" For plan, subscription and box, nobody in the room can yet say.
From chapter 6 again: "If this context were a module inside one deployable rather than its own service, what specifically would be lost?" For ten of the twelve, the answer is nothing anyone can name.
From chapter 2: "Show me one change you could deploy to this service today without deploying, migrating or notifying any other service." For pause and cancellation, which share the subscription's state, there is none.
From chapter 3: "For this request, list every service that must be up and fast right now for it to succeed, and give the product of their availabilities." The monthly billing run touches seven.
From chapter 18: "List what this team owns; can every member describe each item's purpose and last incident?" Four people, twelve services.
Marcus takes it well, because the questions are fair. He concedes the four-person point at once and the language point after some thought. He holds his ground on one thing, rightly: if the box succeeds, curation and the billing cycle will need to move independently, and a monolith that does not enforce its borders will make that a rewrite.
So the decision is neither his proposal nor its opposite. The subscription box is built as a modular monolith: one deployable, with bounded-context modules whose borders are enforced by the build (chapter 9), one schema per module with one writer per table (chapter 8), integration events between modules through an in-process bus with an outbox from day one (chapter 15), and a context map drawn now and revisited quarterly (chapter 7). The billing cycle is a module that talks to the existing Billing service through its published language, because Billing is already a separate country with a separate team. Everything else is one country of federated provinces sharing one parliament building.
The triggers for extraction go in the design record, so the decision is revisited on evidence rather than fashion: a module whose release is blocked by another's more than twice a quarter; a measured need for a different scaling profile; a second team hired to own it; a failure that has taken down the rest more than once. Any of those, and it moves out along the border that already exists.
Remember it as: one nation of provinces until a province needs its own parliament; pay the embassy bill only for the countries that earn it.
Marcus writes the design record with her. It is a better document than either of their first drafts.
What Ana learned
- Budget the premium before proposing a service, and count pagers per person.
- Default to a modular monolith with enforced borders and events from day one, and write down what would trigger extraction.
- Merge services back when a border proves wrong; a team that cannot merge has stopped designing.
…but
There is no next crack, only the next quarter. Ana keeps the questions on a card by her desk. The epilogue is that card, at length.
EpilogueThe Atlas#
This is the card by Ana's desk, at length. Nothing here is new; everything here is somewhere in the twenty-one chapters. It is arranged for the second reading, the design review, and the moment someone says a word in a meeting and you want to know, quickly, what it is for and how it goes wrong.
The metaphor table
Every concept in the book lives in one world. If you can place a word on this map, you can usually explain it.
| Concept | Metaphor |
|---|---|
| Domain | The whole territory the business operates in |
| Ubiquitous language | The local tongue, spoken by everyone in the country, including the code |
| Core / supporting / generic subdomain | The capital, the farmland that feeds it, the utilities bought in (water, power) |
| Bounded context | A country: one language, one law; the border is where a word changes meaning |
| Context map | The atlas page showing the treaties between countries |
| Upstream / downstream | The river: upstream decides what flows down |
| Partnership | Allied nations planning together |
| Shared kernel | A jointly governed border town; nobody may change a street alone |
| Customer–supplier | A trade agreement in which the downstream has a say |
| Conformist | Adopting the neighbour's language wholesale |
| Anti-corruption layer | The customs house with translators |
| Open host service / published language | The port open to all ships, and the trade lingua franca |
| Separate ways | No treaty; the countries simply do not talk |
| Big ball of mud | A region with no borders where every dialect is shouted at once |
| Entity / value object | A person with a passport number, versus an amount of money |
| Aggregate / aggregate root / invariant | A household; one head signs for everyone; a house rule never broken even for a second |
| Repository | The land registry: fetch a whole household by its address |
| Domain service / application service | A notary who works across households; the front desk that receives a request |
| Domain event / integration event | A headline in the local paper; the foreign-language edition of the same paper |
| Command | An order given to one named person |
| Layered / hexagonal / onion / clean | The walled city: the keep holds the domain, the gates are ports, the gatekeepers are adapters, every road leads inward |
| Vertical slice | One road from a gate straight to the keep, per use case |
| Modular monolith | One nation of federated provinces sharing one parliament building |
| Microservice | A country with its own parliament building (deployable) and its own records office (database) |
| Distributed monolith | A house whose rooms are in different cities but share one fuse box |
| Shared database | One records office for many countries; nobody can change a form |
| Synchronous call / asynchronous message | A phone call (both must be awake) / a letter (read when convenient) |
| Idempotency key | The reference number on the letter so it is acted on once |
| Circuit breaker / bulkhead | The fuse / the ship's watertight compartments |
| Transactional outbox | The letter tray by the door, written in the same sitting as the diary entry |
| Inbox / idempotent consumer | The mailroom log of letters already opened |
| Change data capture | Reading the diary to write the letters |
| Saga / compensation | The travel agent booking flight, hotel and car, with cancellation rules |
| Choreography / orchestration | Dancers watching each other / a conductor |
| CQRS | The restaurant: order slips go one way to the kitchen; the menu board is what you read |
| Event sourcing / projection | The accountant's ledger, never edited; the summary card written from it |
| Eventual consistency | The post takes a day, and the business already works that way |
| API composition / BFF | The concierge who phones three offices so the guest makes one call |
| Analytics / data mesh | The national statistics office receiving copies of every ledger |
| Conway's law | The map ends up looking like the org chart |
| Team Topologies | Stream-aligned = a country's own government; platform = roads and post; enabling = travelling advisers; complicated-subsystem = the specialist bureau |
| Strangler fig / branch by abstraction / parallel run | The fig; the new bridge built beside the old before the signpost moves; two clerks doing the same sum for a month |
| Contract test | The customs form both sides signed |
| Correlation id | The tracking number that follows one parcel through every office |
| Microservices premium | Embassies, customs houses and translators for every new country |
The restaurant and the ledger sit inside one country. Everything else is about the map.
The diagram legend
Every diagram in the book uses the same colours, so a purple box is always an aggregate and a red one is always a smell.
Context maps: the arrow points from upstream to downstream; the label reads U: <upstream pattern> → D: <downstream pattern> with OHS (open host service), PL (published language), ACL (anti-corruption layer), CF (conformist) and CS (customer–supplier); a two-headed arrow is a partnership; a small purple node joined to two contexts is a shared kernel; separate ways is said in the caption and never drawn. A subgraph means "deployed together"; a dashed arrow means "subscribes to or reads"; a solid arrow means "calls or commands".
The review checklist
Every trap in the book, by name, with its tell-tale sign. This list is generated from the chapters' The traps boxes by the build, so it is always the same list. Read it before a design review; read it again before signing off.
Strategic: acts I and II (chapters 1–8)
- Big Ball of Mud (ch. 1, One Database, One Dream): any change may break any other part, and estimates are measured in "what it might touch" rather than "what it does".
- Blaming the monolith for the mud (ch. 1, One Database, One Dream): a proposal to split the system that names no business boundary.
- Layers as the only structure (ch. 1, One Database, One Dream): folders called
controllers,servicesandrepositoriesand nothing named after the business. - One word, many meanings (ch. 1, One Database, One Dream): a table with more than about twenty columns, or a
statusfield whose values mean different things to different teams. - Premature decomposition (ch. 2, Let's Do Microservices): a service list produced in an afternoon by a team that could not name its business boundaries the week before.
- Entity services (ch. 2, Let's Do Microservices): a service named after a noun that offers little more than create, read, update and delete on one table.
- Decomposition by technical layer (ch. 2, Let's Do Microservices): services called "web", "logic" and "data".
- Shared database kept "for now" (ch. 2, Let's Do Microservices): the phrase itself.
- Netflix envy (ch. 2, Let's Do Microservices): a slide of a company a thousand times your size.
- Distributed monolith (ch. 3, The Night Everything Deployed Together): a release calendar that lists several services under one change, or a schema migration that requires more than one deployment.
- Synchronous call chains (ch. 3, The Night Everything Deployed Together): a request that cannot complete unless four other services are up right now.
- Retry storms without timeouts (ch. 3, The Night Everything Deployed Together): "retries: 3" in a config file and no timeout beside it.
- Ignoring the fallacies (ch. 3, The Night Everything Deployed Together): a remote call written exactly like the method call it replaced, with no timeout, no failure branch and no thought about partial success.
- Shared thread pools (ch. 3, The Night Everything Deployed Together): one pool serving every outbound dependency.
- Modelling from the database schema (ch. 4, The Workshop With the Orange Stickies): a "domain model" whose classes are the tables and whose relationships are the foreign keys.
- Developers only in the room (ch. 4, The Workshop With the Orange Stickies): a workshop with no one from the warehouse, finance or support.
- Tidying too early (ch. 4, The Workshop With the Orange Stickies): a facilitator who sorts and groups notes in the first ten minutes.
- Treating the wall as the design (ch. 4, The Workshop With the Orange Stickies): a service per subgraph the following week.
- Misidentified core domain (ch. 5, Problem Space, Solution Space): the "core" is whatever is most critical or most complex, rather than what makes customers choose you.
- Building the generic subdomain (ch. 5, Problem Space, Solution Space): a home-grown version of something with three good vendors, justified by "our needs are special".
- Everything is core (ch. 5, Problem Space, Solution Space): a classification with no generic or supporting entries.
- Classifying once (ch. 5, Problem Space, Solution Space): a subdomain chart from three years ago on the wiki.
- One canonical model for the whole company (ch. 6, Drawing the Borders): an enterprise "Customer" schema that every team must conform to.
- Bounded context drawn around a team by accident (ch. 6, Drawing the Borders): a context whose border matches the current org chart rather than a change in language.
- Nanoservices (ch. 6, Drawing the Borders): a service so small that most changes need three of them.
- Rigid one-context-one-service rule (ch. 6, Drawing the Borders): a context split into a separate deployable on principle, when a module would have served.
- Shared kernel abuse (ch. 7, The Map of Treaties): a "common" or "shared" module that grows a new class every sprint.
- Accidental conformist (ch. 7, The Map of Treaties): a downstream that imports an upstream's classes because they were there.
- Skipping the anti-corruption layer against a legacy or third-party model (ch. 7, The Map of Treaties): a payment provider's field names, or the old monolith's
status2, appearing in your domain model. - Leaking another context's model into your own (ch. 7, The Map of Treaties): Billing's
Invoicetype in Ordering's code. - Drawing the map once (ch. 7, The Map of Treaties): a context map on a wiki page dated two years ago.
- Cross-service joins in application code (ch. 8, Nobody Owns the Product Table): a loop that calls one service per row of another service's result, or a SQL join across two schemas.
- One context owning another's data (ch. 8, Nobody Owns the Product Table): a "core data" or "master data" context that stores everyone's fields so that nobody duplicates them.
- Duplication panic (ch. 8, Nobody Owns the Product Table): a design review that rejects a second copy of a product name on principle.
- Shared schema, separate services (ch. 8, Nobody Owns the Product Table): three services with three connection strings to the same tables.
- Copy without a source of truth (ch. 8, Nobody Owns the Product Table): three contexts each editing the product name.
Tactical and architecture: act III (chapters 9–13)
- Seven layers for a CRUD context (ch. 9, Which Way Do the Arrows Point?): a request that passes through a controller, a service, a use case, a port, an adapter, a repository interface and an ORM to update one column.
- Repository interface that is a data-access object in disguise (ch. 9, Which Way Do the Arrows Point?): a "repository" with
findByStatusAndCreatedAtBetweenand a method per query the UI needs. - Domain package that imports the framework (ch. 9, Which Way Do the Arrows Point?):
@Entity,@JsonPropertyor@Transactionalon a domain class. - Interface for everything (ch. 9, Which Way Do the Arrows Point?): an interface with exactly one implementation, in the same package, that will never have another.
- Vertical slices with no shared domain discipline (ch. 9, Which Way Do the Arrows Point?): the same order-total rule implemented three slightly different ways in three feature packages.
- Anaemic domain model (ch. 10, The Order That Ate the Database): a class of getters and setters and a
Servicefor every verb. - Giant aggregate (ch. 10, The Order That Ate the Database): a root that owns everything about the thing rather than everything that must be true with it.
- Aggregates referencing other aggregates by object (ch. 10, The Order That Ate the Database):
order.getBuyer().getAddress(). - Modifying several aggregates in one transaction (ch. 10, The Order That Ate the Database): a handler that loads two roots and saves both.
- ORM concerns shaping the model (ch. 10, The Order That Ate the Database): a bidirectional relationship, a public no-arg constructor or a field added "because the mapper needs it".
- Rich model in a generic subdomain (ch. 10, The Order That Ate the Database): an aggregate with a factory and a domain service for a settings table.
- Leaking internal domain events as the public contract (ch. 11, Something Happened): another context deserialising your aggregate's own event class.
- Fat events that carry a whole aggregate to everyone (ch. 11, Something Happened): an event with forty fields, most unused by any consumer.
- Thin events that force every consumer to call back (ch. 11, Something Happened): an event containing only an identifier, and consumers that immediately call the producer's API.
- Commands disguised as events (ch. 11, Something Happened): an event named
CheckCreditorSendInvoice. - Publishing without a version (ch. 11, Something Happened): an event type with no version in its name or envelope.
- CQRS everywhere (ch. 12, Two Models Are Cheaper Than One): a projection and a read table for a context whose screen is the aggregate.
- A read model treated as the source of truth (ch. 12, Two Models Are Cheaper Than One): a command handler that checks a read table before deciding.
- Projections that call back (ch. 12, Two Models Are Cheaper Than One): a projection that, on receiving an event, queries three services to enrich the row.
- No rebuild path (ch. 12, Two Models Are Cheaper Than One): a read model that cannot be dropped and regenerated.
- Hiding the lag (ch. 12, Two Models Are Cheaper Than One): a UI that reads the model immediately after a command and shows stale data with no indication.
- Event sourcing as the integration mechanism (ch. 13, The Ledger Never Lies): another context subscribing directly to your event store.
- Ignoring event versioning (ch. 13, The Ledger Never Lies): event classes with no version and no upcasters.
- Personal data in immutable events (ch. 13, The Ledger Never Lies): a name or address inside an event payload.
- Premature snapshots (ch. 13, The Ledger Never Lies): snapshotting every aggregate from day one.
- Event sourcing a supporting subdomain (ch. 13, The Ledger Never Lies): an event-sourced settings module or returns desk.
- Rules running during replay (ch. 13, The Ledger Never Lies): validation or side effects inside
apply.
Integration: act IV (chapters 14–17)
- No timeouts (ch. 14, The Payment Taken Twice): an HTTP client built with defaults, or a timeout of "infinite" in a config file.
- Retries without idempotency (ch. 14, The Payment Taken Twice): a retry policy on a call that creates, charges or sends.
- Request/response over a message broker (ch. 14, The Payment Taken Twice): a "reply queue" and a caller that blocks waiting for it.
- Chatty APIs (ch. 14, The Payment Taken Twice): a screen or a handler that makes ten small calls to the same neighbour.
- One pool for everything (ch. 14, The Payment Taken Twice): a single connection or thread pool shared by all outbound dependencies.
- Retrying on the wrong errors (ch. 14, The Payment Taken Twice): a retry on a 400 or a validation failure.
- Dual write (ch. 15, The Message That Never Left): a save to the database followed by a publish to a broker (or the reverse) in the same handler with no outbox between them.
- Publishing from memory after commit (ch. 15, The Message That Never Left): "publish after the transaction succeeds" in a comment.
- Non-idempotent consumers (ch. 15, The Message That Never Left): a consumer with no inbox and a side effect (charge, email, increment).
- Assuming global ordering (ch. 15, The Message That Never Left): a consumer that relies on events for different aggregates arriving in send order.
- Breaking schema changes (ch. 15, The Message That Never Left): a renamed or removed field in a published event and a single deploy.
- No consumer contracts (ch. 15, The Message That Never Left): a producer whose only knowledge of its consumers is a wiki page.
- Two-phase commit across services (ch. 16, Who Says the Order Is Done?): a distributed transaction coordinator between databases owned by different contexts.
- Saga without compensations (ch. 16, Who Says the Order Is Done?): a happy path with events flowing forward and no handler for a failure event after step one.
- Orchestrator that becomes a god service (ch. 16, Who Says the Order Is Done?): a process manager that validates stock levels, computes tax and checks credit itself.
- Choreography spaghetti (ch. 16, Who Says the Order Is Done?): nobody can draw the saga, and finding out what state an order is in means reading four contexts' inboxes.
- Technical lock where a business rule was needed (ch. 16, Who Says the Order Is Done?): a row lock or a distributed lock introduced to hide an intermediate state the business already has a word for.
- Non-idempotent steps (ch. 16, Who Says the Order Is Done?): a saga step without an idempotency key.
- Reporting by querying every service's database (ch. 17, The Report Nobody Could Run): a "read-only reporting user" on every schema.
- The enterprise service bus and smart pipes (ch. 17, The Report Nobody Could Run): transformation, routing rules or business logic configured in the message broker or gateway.
- The shared canonical data model, revisited (ch. 17, The Report Nobody Could Run): a reporting schema that contexts are asked to emit directly, "to make ingestion easier".
- Composition for lists (ch. 17, The Report Nobody Could Run): an API composer paginating across three services.
- A pipeline nobody watches (ch. 17, The Report Nobody Could Run): a warehouse with no freshness metric.
Organisation and migration: chapters 18, 19 and 21
- Three developers, forty services (ch. 18, You Ship Your Org Chart): a team that cannot list what it owns without opening a dashboard.
- Microservices to fix a people problem (ch. 18, You Ship Your Org Chart): a new service whose justification is a team you would rather not talk to.
- No platform team (ch. 18, You Ship Your Org Chart): every stream-aligned team running its own pipeline, broker and dashboards.
- No service ownership (ch. 18, You Ship Your Org Chart): a context whose last three changes came from three teams.
- Teams named after layers (ch. 18, You Ship Your Org Chart): Frontend, Backend and Data.
- Reorganising by decree (ch. 18, You Ship Your Org Chart): a new org chart announced on Monday with the old codebase on Tuesday.
- Big-bang rewrite (ch. 19, Strangling the Monolith): a plan whose first user-visible delivery is the last milestone.
- Extracting the most coupled component first (ch. 19, Strangling the Monolith): a first migration that touches four other systems and takes a quarter.
- Leaving the database shared indefinitely (ch. 19, Strangling the Monolith): new code with a clean model reading the old tables through an adapter, "until we get to the data".
- Strangling without a facade (ch. 19, Strangling the Monolith): callers pointed directly at the new service, one by one, with no single routing point.
- No comparison in the parallel run (ch. 19, Strangling the Monolith): both systems running and nobody reading the diff log.
- Declaring victory at the code (ch. 19, Strangling the Monolith): a retirement party for the legacy service while its database is still queried nightly.
- Résumé-driven architecture (ch. 21, The Consultant Who Said Microservices): a proposal whose service count exceeds its team size and whose justification cites companies a thousand times larger.
- Ignoring the microservices premium (ch. 21, The Consultant Who Said Microservices): a plan that budgets for building services and not for running them: no line for pipelines, pagers, contracts, tracing or the platform team's time.
- Never merging services back (ch. 21, The Consultant Who Said Microservices): two services that always change together, owned by one team, kept separate "because splitting was hard".
- Treating the monolith as a stage (ch. 21, The Consultant Who Said Microservices): "we'll start with a monolith and split later" with no module borders enforced.
Operations: chapter 20
- The all-system end-to-end suite (ch. 20, Can You See It?): a test stage that needs every service deployed together and takes hours.
- No contract tests (ch. 20, Can You See It?): a producer whose only knowledge of its consumers is a wiki page and hope.
- No correlation ids (ch. 20, Can You See It?): an incident investigated by matching timestamps across six log files.
- Lockstep deploys (ch. 20, Can You See It?): a release train, an integration window, or a change that lists four services.
- Version in the URL without expand/contract (ch. 20, Can You See It?): a
/v2/endpoint that appeared the day/v1/disappeared. - Testing through the UI (ch. 20, Can You See It?): a browser-driven suite as the main safety net for service boundaries.
The timeline
Three eras. In the first, the ideas were written down and mostly ignored. In the second, the microservices movement made them urgent. In the third, the evidence came back and the field settled on judgement over fashion.
Before 2004: the ideas exist. Conway, Parnas and the saga paper are older than most practitioners. Evans's 2003 book collected the strategic and tactical patterns, and the messaging patterns arrived the same year. Almost nobody built services yet; the ideas waited.
2004 to 2014: the architecture styles and the name. Hexagonal, onion and clean each reacted to domain logic welded to infrastructure. Helland and Hohpe argued that big systems cannot use distributed transactions. CQRS and event sourcing were named, Event Storming was invented, and in 2014 the word "microservices" got its definition.
2015 onward: the evidence and the judgement. Newman's book made independent deployability the test; Fowler and Tilkov disagreed in public about where to start. Then the outbox, contracts and data-intensive theory arrived, Team Topologies named the organisation, and Segment, Shopify and Prime Video reported what happened when the premium was and was not worth paying.
The reading list, by year
Where a source is freely available online it is named by its title; the rest are books.
- 1968 Melvin Conway, "How Do Committees Invent?", Datamation.
- 1972 David Parnas, "On the Criteria To Be Used in Decomposing Systems into Modules", CACM.
- 1987 Hector Garcia-Molina and Kenneth Salem, "Sagas", SIGMOD.
- 1994 Peter Deutsch, the fallacies of distributed computing (eighth by James Gosling, 1997).
- 1996 Buschmann, Meunier, Rohnert, Sommerlad and Stal, Pattern-Oriented Software Architecture, Volume 1.
- 1999 Brian Foote and Joseph Yoder, "Big Ball of Mud" (PLoP 1997).
- 2002 Martin Fowler, Patterns of Enterprise Application Architecture.
- 2003 Eric Evans, Domain-Driven Design: Tackling Complexity in the Heart of Software. Read the strategic chapters first.
- 2003 Martin Fowler, "AnemicDomainModel" (bliki). Gregor Hohpe and Bobby Woolf, Enterprise Integration Patterns.
- 2004 Martin Fowler, "StranglerFigApplication" (bliki). Gregor Hohpe, "Starbucks Does Not Use Two-Phase Commit".
- 2005 Alistair Cockburn, "Hexagonal architecture". Martin Fowler, "Event Sourcing".
- 2007 Pat Helland, "Life beyond Distributed Transactions" (CIDR) and "Memories, Guesses and Apologies".
- 2008 Jeffrey Palermo, "The Onion Architecture".
- 2009 Alberto Brandolini, "Strategic Domain Driven Design with Context Mapping" (InfoQ).
- 2010 Greg Young, "CQRS Documents".
- 2011 Steve Yegge's account of the Amazon API mandate; Netflix's early microservices talks.
- 2012 Robert C. Martin, "The Clean Architecture" (book 2017).
- 2013 Alberto Brandolini, Introducing EventStorming. Vaughn Vernon, Implementing Domain-Driven Design.
- 2014 James Lewis and Martin Fowler, "Microservices". Chris Richardson, microservices.io.
- 2015 Sam Newman, Building Microservices. Martin Fowler, "MicroservicePremium" and "MonolithFirst". Stefan Tilkov, "Don't start with a monolith". Scott Millett and Nick Tune, Patterns, Principles, and Practices of Domain-Driven Design.
- 2016 Debezium 0.1. David Heinemeier Hansson, "The Majestic Monolith".
- 2017 Martin Fowler, "What do you mean by 'Event-Driven'?". Martin Kleppmann, Designing Data-Intensive Applications.
- 2018 Jimmy Bogard, "Vertical Slice Architecture". Alexandra Noonan, "Goodbye Microservices" (Segment). Chris Richardson, Microservices Patterns. Scott Wlaschin, Domain Modeling Made Functional.
- 2019 Matthew Skelton and Manuel Pais, Team Topologies. Sam Newman, Monolith to Microservices. Zhamak Dehghani, "How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh". Kirsten Westeinde, "Deconstructing the Monolith" (Shopify).
- 2020 Adam Gluck, "Introducing Domain-Oriented Microservice Architecture" (Uber).
- 2021 Vlad Khononov, Learning Domain-Driven Design. Sam Newman, Building Microservices, second edition. Vaughn Vernon and Tomasz Jaskuła, Strategic Monoliths and Microservices.
- 2022 Spring Modulith (1.0 in 2023).
- 2023 Marcin Kolny, "Scaling up the Prime Video audio/video monitoring service and reducing costs by 90%".
- 2024 Vlad Khononov, Balancing Coupling in Software Design.
Two package layouts for one context
Both trees hold the same use case, placing an order, in Bramble's Ordering context. The hexagonal layout groups by role and suits a context with rich rules and several adapters; the vertical-slice layout groups by feature and suits a context where most use cases are simple and should stay small. In both, domain depends on nothing.
ordering/ hexagonal: ports and adapters
├── domain/ the keep: depends on nothing
│ ├── Order.java aggregate root; place(), addLine(), cancel()
│ ├── OrderId.java OrderLine.java Money.java
│ ├── OrderPlaced.java OrderCancelled.java domain events (internal shape)
│ ├── OrderRepository.java port (outbound): load/save whole aggregates
│ └── PaymentGateway.java port (outbound): take(orderId, amount)
├── application/ use cases: orchestrate, no rules
│ ├── PlaceOrder.java PlaceOrderHandler.java
│ └── CancelOrder.java CancelOrderHandler.java
├── adapters/
│ ├── in/
│ │ ├── web/OrderController.java driving: HTTP -> PlaceOrder command
│ │ └── messaging/PaymentEventsConsumer.java driving: PaymentTakenV1 -> command
│ └── out/
│ ├── persistence/JpaOrderRepository.java implements OrderRepository
│ ├── persistence/OutboxAppender.java same transaction as the aggregate
│ ├── payments/StripePaymentGateway.java implements PaymentGateway
│ └── events/OrderingPublisher.java domain event -> OrderPlacedV1
└── published/
└── OrderPlacedV1.java the published language: stable, versioned
ordering/ vertical slices
├── domain/ still the keep; shared by every slice
│ ├── Order.java OrderId.java OrderLine.java Money.java
│ ├── OrderPlaced.java OrderCancelled.java
│ ├── OrderRepository.java PaymentGateway.java ports, still owned here
├── features/
│ ├── placeorder/
│ │ ├── PlaceOrder.java request
│ │ ├── PlaceOrderHandler.java the whole use case, top to bottom
│ │ └── PlaceOrderEndpoint.java its own HTTP mapping
│ ├── cancelorder/
│ │ ├── CancelOrder.java CancelOrderHandler.java CancelOrderEndpoint.java
│ └── accountpage/
│ ├── AccountPageProjection.java read side: events -> flat rows
│ └── AccountPageQuery.java one indexed SELECT
├── infrastructure/ adapters shared by slices that need them
│ ├── JpaOrderRepository.java OutboxAppender.java StripePaymentGateway.java
│ └── OrderingPublisher.java
└── published/
└── OrderPlacedV1.java
Prefer the first when the context is core and its adapters are many; prefer the second when most features are a handler and a query. Either way, a build rule (Spring Modulith or its equivalent) should fail if domain imports anything from adapters, infrastructure or another context.
The card
If you keep only one page, keep this one.
- Which business boundary, in business words, does the proposed split follow? (ch. 1)
- Show me one change you could deploy to this service today without deploying, migrating or notifying any other service. (ch. 2)
- For this request, list every service that must be up and fast right now for it to succeed, and give the product of their availabilities. (ch. 3)
- Where on the timeline did a word change meaning, and what are the two meanings? (ch. 4)
- Name the subdomain that customers choose you for, in one sentence, and say who on the team is working on it this quarter. (ch. 5)
- If this context were a module inside one deployable rather than its own service, what specifically would be lost? (ch. 6)
- For this border, which side is upstream, and what does the downstream do when the upstream changes its model? (ch. 7)
- For each table this context reads, which context writes it, and how would this context learn of a schema change? (ch. 8)
- Can the domain package compile with no framework on the classpath? (ch. 9)
- Which rule in this aggregate must be true this instant, and which could be true within a minute? (ch. 10)
- Which class does the consumer deserialise, and does it live in the producer's domain package? (ch. 11)
- Does any command handler read from a read model before deciding? (ch. 12)
- Who outside this context reads the event store, and why is that not a published event? (ch. 13)
- For each outbound call, what is the timeout, and what does the user see when it fires? (ch. 14)
- What happens when this consumer receives the same event twice, and where is that recorded? (ch. 15)
- For each step in this process, what is the compensating action, and which context performs it? (ch. 16)
- How fresh must this view be, and how many rows does it join? (ch. 17)
- For each context, which single team owns its code, its pager and its roadmap? (ch. 18)
- What is the first capability to move, and what user-visible value does it deliver in the first month? (ch. 19)
- When this service changes its event or interface, whose build fails first? (ch. 20)
- What evidence, written down now, would trigger extracting a module, and who reviews it? (ch. 21)