AWS for Azure Architects and Developers#

Design enterprise AWS systems end to end, whatever you know of Azure

Compiled bykodebot

September 2026 edition Facts checked 12 September 2026 about 2 hours 55 minutes of reading

Read this first#

Your next systems run on AWS. This book teaches AWS from first principles, with Azure as an optional bridge, so that after about 2 hours 55 minutes of reading you can design an enterprise-grade AWS system end to end.

Who it is for#

Architects and developers. You need no Azure knowledge: where Azure helps, it appears beside AWS in a figure pair, a table column, a profile's "On Azure" line or a false-friend box, and every AWS idea stands without it. You need no AWS account to read it either; the one hands-on step, a first sign-in in chapter 1, uses an account your organisation gives you.

How to read it#

  • One story. MegaCorp's payments team moves to AWS, chapter by chapter. Each tool arrives when the team needs it, and its profile tells what it does, how it works, when to use it and when not to, its limits, cost, security and gotchas.
  • Figures carry facts. Architecture figures use the official AWS and Azure icons, and numbered steps match the legend under each figure. Every figure has an "As text" version. chapter 1 shows how to read them.
  • Nothing is folded away. What you see is the whole book. Each section ends with links to the official documentation, where the fine detail lives.
  • Code only where it teaches. Java shows the AWS SDK and AWS Lambda. Terraform appears at most once a chapter, Azure beside AWS, where one tool shows how the clouds differ.
  • Search with ⌘K or Ctrl+K. Type an Azure name, such as Event Grid, to reach its AWS counterpart, or paste an error message.

What it covers#

Part 0 gets you started. Part I orients you: AWS on one page, its history, the six shifts in thinking and the false friends. Part II lays the foundations a company decides once: accounts, identity, networks and the landing zone. Part III covers the building blocks each workload chooses from. Part IV asks what every design must answer: security, failure, cost, delivery and operations. Part V puts it together in three worked designs and a set of review drills.

The teaching parts take about 2 hours 55 minutes to read, as the build measures them. The appendices, from the service map to the cheat card, are for lookup and sit outside that time.

Facts were checked on 12 September 2026, and a statement about what is current carries its date, because AWS changes every week. The official icons are used under each vendor's terms, credited in chapter 46.

Before you start#

Part 0 · Before you start · Chapter 01·5 min read

Three things make the rest of the book easier: the shape of AWS on one page, how to read its figures, and a first sign-in that tells you exactly who and where you are.

The cloud on one page#

On its first day, MegaCorp's payments team asks what every newcomer asks: where does everything actually live? Almost always, the answer has two parts: an account and a Region.

An AWS account holds resources and is their security boundary. Nothing in one account can reach another unless someone explicitly allows it, and costs are counted per account. A Region is a geographic area, such as London, eu-west-2, made of several Availability Zones. What you create in one Region does not exist in any other unless you copy or replicate it.

From zero: Availability Zone

An Availability Zone is one or more data centres with their own power, networking and connectivity. The zones in a Region sit far enough apart, up to about 100 km, not to fail together, and close enough for synchronous replication within a few milliseconds. Every Region has three or more, as of September 2026.

A few services are global rather than Regional. IAM, which decides who may do what, and Route 53, which answers DNS queries, serve every Region, although each is managed from one Region, us-east-1.

Where things live: resources belong to an account; most sit in a Region, some in one of its Availability Zones, and a few services, such as IAM and Route 53, are globalAWS CloudAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPrivate subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24AWS IAMAmazon Route 53Amazon S3(statements)Amazon EC2(app)Amazon EC2(app)
Figure Where things live: resources belong to an account; most sit in a Region, some in one of its Availability Zones, and a few services, such as IAM and Route 53, are global#
As text
  • AWS Cloud
    • AWS Account payments-prod
      • AWS IAM
      • Amazon Route 53
      • Region eu-west-2
        • Amazon S3 (statements)
        • VPC 10.20.0.0/16
          • Availability Zone a
            • Private subnet 10.20.10.0/24
              • Amazon EC2 (app)
          • Availability Zone b
            • Private subnet 10.20.11.0/24
              • Amazon EC2 (app)

Read more What is an AWS account? · AWS Fault Isolation Boundaries: Regions · AWS Fault Isolation Boundaries: Availability Zones · AWS Fault Isolation Boundaries: global services

How to read the figures#

Most figures in this book are architecture diagrams drawn the way AWS draws them, with the official icons. Boxes are groups: whatever sits inside a box belongs to it. The colour and line of a box say what kind of group it is.

BoxLineWhat it shows
AWS Clouddark greythe whole of AWS
AWS accountpinka boundary for security, cost and quotas
Regionteal, dottedone geographic area
VPCpurplea private network in one Region
Availability Zoneteal, dashedone or more data centres within a Region
Public subnetgreena subnet with a route to the internet
Private subnetteala subnet without one
Security groupreda firewall around resources
Auto Scaling grouporange, dashedinstances that grow and shrink in number

Numbered black circles mark the steps of a flow, and the legend under the figure explains them in order. In the next figure, a customer's request arrives at a load balancer (step 1), which passes it to an instance in a private subnet (step 2).

Reading a figure: group boxes nest, official icons show services, and numbered steps follow the legendRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24CustomersApplicationLoad BalancerAmazon EC2(app)12
1A customer's request reaches the load balancer in a public subnet
2The load balancer passes it to an instance in a private subnet
Figure Reading a figure: group boxes nest, official icons show services, and numbered steps follow the legend#
As text
  • Customers
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • Application Load Balancer
        • Private subnet 10.20.10.0/24
          • Amazon EC2 (app)
  1. A customer's request reaches the load balancer in a public subnet
  2. The load balancer passes it to an instance in a private subnet

Where Azure helps, a figure pair shows Azure on the left and AWS on the right. Every figure also has an "As text" version beneath it, which a screen reader can read and the search can find.

Read more AWS Architecture Icons · Azure architecture icons

A first sign-in#

Alex joins the payments team as a developer. MegaCorp signs its people in to AWS through IAM Identity Center, so Alex never holds a long-term password or access key for AWS. Instead, the AWS CLI asks Identity Center for short-term credentials for a role: an identity, with its own permissions, that Alex can assume for a limited time.

aws configure sso
aws sso login --profile payments-dev
aws sts get-caller-identity --profile payments-dev

aws configure sso runs once. It asks for MegaCorp's Identity Center start URL and Region, opens a browser to sign in, lists the accounts and roles Alex may use, and saves the choice as a profile. Later, aws sso login signs in again when the session has expired.

A first sign-in: the AWS CLI gets short-term credentials for a role through IAM Identity Center, then asks AWS STS who is callingAlexAWS CLIIAM Identity CenterAWS STSaws sso login1sign in through thebrowser2a session token, cached ondisk3credentials for theprofile's role4short-term credentials5aws stsget-caller-identity6GetCallerIdentity7the account and the role's ARN8
Figure A first sign-in: the AWS CLI gets short-term credentials for a role through IAM Identity Center, then asks AWS STS who is calling#
As text
  1. Alex to AWS CLI: aws sso login
  2. AWS CLI to IAM Identity Center: sign in through the browser
  3. IAM Identity Center to AWS CLI: a session token, cached on disk
  4. AWS CLI to IAM Identity Center: credentials for the profile's role
  5. IAM Identity Center to AWS CLI: short-term credentials
  6. Alex to AWS CLI: aws sts get-caller-identity
  7. AWS CLI to AWS STS: GetCallerIdentity
  8. AWS STS to AWS CLI: the account and the role's ARN
{
    "UserId": "AIDASAMPLEUSERID",
    "Account": "123456789012",
    "Arn": "arn:aws:iam::123456789012:user/DevAdmin"
}

The documentation's example comes from an IAM user. Signed in through Identity Center, Alex's Arn names an assumed role instead, one that Identity Center created: arn:aws:sts::123456789012:assumed-role/AWSReservedSSO_PowerUserAccess_…/alex. Account says which account the command reached; check it before changing anything. The profile also stores a default Region. Commands go to that Region, and resources in any other Region do not appear in their results.

Gotcha

get-caller-identity always succeeds. It needs no permissions, and works even when a policy denies it, so it proves who you are, never what you may do.

Read more Installing or updating the AWS CLI · Configuring IAM Identity Center authentication with the AWS CLI · AWS CLI: aws sts get-caller-identity

Before you start
  1. Alex launches an instance in eu-west-2, then lists instances with a profile whose default Region is eu-west-1. The new instance is missing. Why?
    Answer
    The command went to eu-west-1, and resources in one Region do not exist in another. Run it with --region eu-west-2, or use a profile set to that Region.
  2. In a figure, a green box sits inside a teal dashed box, which sits inside a purple box. What are the three boxes?
    Answer
    A public subnet, inside an Availability Zone, inside a VPC.
  3. Why would the payments team run a copy of each tier in a second Availability Zone?
    Answer
    A zone is one or more data centres that can fail on its own, and the zones in a Region are built not to fail together. A copy in a second zone keeps running when the first fails.
  4. aws sts get-caller-identity succeeds for Alex. Can Alex now list the team's instances?
    Answer
    Not necessarily. The call needs no permissions. It shows who Alex is and which account the command reached, not what Alex may do.

AWS on one page#

Part I · Orientation · Chapter 02·3 min read

AWS is a long list of separate services, but they fall into a handful of categories, and each category has its own icon colour. Learn them, and a figure tells you what each box does before you read the label.

A handful of categories#

When MegaCorp's payments team first opens the AWS console, the list of services is long enough to discourage anyone. It becomes manageable once the team sees that every service belongs to one category, and that a design needs only a few services from each.

Every AWS service icon sits on a square tile in its category's colour. The first figure shows where workloads run and where data lives; the second shows how the parts connect, and how they are kept safe and visible.

Where workloads run and data lives: compute and containers are orange, storage green, databases magentaComputeContainersStorageDatabasesAWS LambdaAWS FargateAmazon S3Amazon DynamoDB
Figure Where workloads run and data lives: compute and containers are orange, storage green, databases magenta#
As text
  • Compute
    • AWS Lambda
  • Containers
    • AWS Fargate
  • Storage
    • Amazon S3
  • Databases
    • Amazon DynamoDB
How the parts connect and stay safe: networking is purple, application integration and management pink, security redNetworkingIntegrationSecurityManagementAmazon VPCAmazon SQSAWS KMSAmazonCloudWatch
Figure How the parts connect and stay safe: networking is purple, application integration and management pink, security red#
As text
  • Networking
    • Amazon VPC
  • Integration
    • Amazon SQS
  • Security
    • AWS KMS
  • Management
    • Amazon CloudWatch

The table lists every category this book teaches, with the services it covers and where. Each service is its own product, with its own API, console pages, quotas and pricing; chapter 4 explains why that matters.

CategoryIcon colourServices this book teachesChapters
ComputeorangeAmazon EC2, Amazon EC2 Auto Scaling, AWS Lambdachapter 17, chapter 19
ContainersorangeAmazon ECR, Amazon ECS, AWS Fargate, Amazon EKSchapter 18
StoragegreenAmazon S3, Amazon EBS, Amazon EFS, Amazon FSx, AWS Backupchapter 20, chapter 30
DatabasesmagentaAmazon RDS, Amazon Aurora, Amazon DynamoDB, Amazon ElastiCachechapter 21, chapter 22
Networking and content deliverypurpleAmazon VPC, Amazon Route 53, Amazon CloudFront, Elastic Load Balancing, Amazon API Gatewaychapter 12, chapter 13, chapter 14, chapter 24
Application integrationpinkAmazon SQS, Amazon SNS, Amazon EventBridge, AWS Step Functionschapter 23, chapter 19
AnalyticspurpleAmazon Athena, AWS Glue, Amazon Redshift, Amazon Kinesis Data Streams, Amazon MSKchapter 25, chapter 23
Artificial intelligencetealAmazon Bedrockchapter 26
Security, identity and complianceredAWS IAM, AWS IAM Identity Center, AWS KMS, Amazon GuardDuty, AWS WAFchapter 8, chapter 27, chapter 28
Management and governancepinkAWS Organizations, AWS Control Tower, Amazon CloudWatch, AWS CloudTrail, AWS Systems Managerchapter 7, chapter 29, chapter 34
Developer toolsmagentaAWS CloudFormation, AWS CDK, AWS CodePipeline, AWS CodeBuildchapter 32, chapter 33
Cloud financial managementgreenAWS Cost Explorer, AWS Budgets, Savings Planschapter 31
Gotcha

A colour names a category, not a service, and two colours are shared. Pink marks both management and application integration, and purple both networking and analytics, so read the label before you assume.

Read more Overview of Amazon Web Services · AWS Architecture Icons

Colours in a design#

Once the colours are familiar, a design reads at a glance. A customer's request to MegaCorp enters through CloudFront (step 1), purple because it moves traffic, and passes a load balancer (step 2). It reaches the orange compute that runs the payments code (step 3) and ends at a magenta database (step 4).

A payments request crosses the categories: purple networking, orange compute, a magenta databaseRegion eu-west-2VPC 10.20.0.0/16CustomersAmazonCloudFrontApplicationLoad BalancerAWS Fargate(payments)Amazon Aurora1234
1A customer's request reaches CloudFront at the edge
2CloudFront forwards it to the load balancer
3The load balancer passes it to a payments task
4The task reads and writes the database
Figure A payments request crosses the categories: purple networking, orange compute, a magenta database#
As text
  • Customers
  • Amazon CloudFront
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Application Load Balancer
      • AWS Fargate (payments)
      • Amazon Aurora
  1. A customer's request reaches CloudFront at the edge
  2. CloudFront forwards it to the load balancer
  3. The load balancer passes it to a payments task
  4. The task reads and writes the database

Read more AWS Architecture Icons

How AWS got here#

Part I · Orientation · Chapter 03·4 min read

AWS began in 2006 as a handful of separate services and grew one service at a time; Azure followed from 2008. That history explains why AWS feels like many products, and why older accounts look different from new ones.

Two clouds, two decades#

AWS launched in the spring of 2006 with Amazon S3, a service for storing objects, and added Amazon EC2, for renting servers, a few months later. Each was its own product with its own API. That pattern held: AWS grew by adding separate services, each launched when it was ready.

Two clouds, two decades: AWS from 2006 and Azure from 2008, with the launches that shape today's designsAWSAzure2006Amazon S3,then AmazonEC22008Windows Azureannounced2009Amazon VPC2010Windows Azuregenerallyavailable2011AWSCloudFormation2014AWS LambdaRenamedMicrosoftAzure2016AzureFunctionsgenerallyavailable2017AWS Fargate2018Amazon EKSgenerallyavailableAzureKubernetesServicegenerallyavailable2023Amazon BedrockgenerallyavailableAzure OpenAIServicegenerallyavailableAzure ADrenamedMicrosoftEntra ID
Figure Two clouds, two decades: AWS from 2006 and Azure from 2008, with the launches that shape today's designs#
As text
  1. 2006. AWS: Amazon S3, then Amazon EC2.
  2. 2008. Azure: Windows Azure announced.
  3. 2009. AWS: Amazon VPC.
  4. 2010. Azure: Windows Azure generally available.
  5. 2011. AWS: AWS CloudFormation.
  6. 2014. AWS: AWS Lambda. Azure: Renamed Microsoft Azure.
  7. 2016. Azure: Azure Functions generally available.
  8. 2017. AWS: AWS Fargate.
  9. 2018. AWS: Amazon EKS generally available. Azure: Azure Kubernetes Service generally available.
  10. 2023. AWS: Amazon Bedrock generally available. Azure: Azure OpenAI Service generally available; Azure AD renamed Microsoft Entra ID.

The strip below follows one thread through that history: how much of the machine you run yourself. Each step hands more of it to AWS, and each is still on sale, so designs choose among them.

Running code on AWS, from servers you manage to none: Amazon EC2, AWS Lambda, AWS Fargate and Amazon EKSAmazon EC2: serversyou manage2006AWS Lambda: code, noservers2014AWS Fargate:containers, noservers2017Amazon EKS: managedKubernetes2018
Figure Running code on AWS, from servers you manage to none: Amazon EC2, AWS Lambda, AWS Fargate and Amazon EKS#
As text
  1. 2006: Amazon EC2: servers you manage
  2. 2014: AWS Lambda: code, no servers
  3. 2017: AWS Fargate: containers, no servers
  4. 2018: Amazon EKS: managed Kubernetes

Read more Overview of Amazon Web Services · Amazon VPC User Guide: document history

Eras you will meet#

History matters because accounts keep what they were given. The payments team inherits an AWS account from a company MegaCorp bought, and its age shows in three places.

  • No default VPC. Accounts created after 4 December 2013 got a default VPC in every Region. An older account may have none, and its first instances ran in EC2-Classic, a flat network AWS has since retired.
  • NAT instances. Before NAT gateways arrived in December 2015, private subnets reached the internet through NAT instances: ordinary instances the team had to run, patch and scale itself.
  • No servers. From 2014, AWS Lambda, and later AWS Fargate, ran code and containers with no servers to manage. Newer parts of an estate lean on them.
Dating an inherited AWS account from what it containsnoyesyesnoyesnoDoes every Region have a default VPC?Created before 4 December 2013, orsomeone deleted themDo private subnets reach the internetthrough NAT instances?Built before December 2015, ornever updatedDo any subnets use a regional NATgateway?Networking updated since November2025Date the rest with the appendix on theAWS you inherit
Figure Dating an inherited AWS account from what it contains#
As text
  1. Does every Region have a default VPC? No: Created before 4 December 2013, or someone deleted them. Yes: the next step.
  2. Do private subnets reach the internet through NAT instances? Yes: Built before December 2015, or never updated. No: the next step.
  3. Do any subnets use a regional NAT gateway? Yes: Networking updated since November 2025. No: the next step.
  4. Date the rest with the appendix on the AWS you inherit
Gotcha

An account's age sets its defaults. A script that assumes a default VPC fails in an account created before December 2013, or in one where the default VPC was deleted.

Read more Default VPCs · NAT instances

Why AWS feels like many products#

Because each service launched on its own, each still has its own API, console pages, quotas, pricing and documentation. Services overlap: there are several ways to run a container or send a message, each from a different year and for a different need. Old features retire slowly; EC2-Classic took from 2021 to 2023 to go. chapter 44 dates what an estate contains.

Read more EC2-Classic is retiring: here's how to prepare

Six shifts in thinking#

Part I · Orientation · Chapter 04·3 min read

Six ideas change how you design on AWS: accounts as boundaries, roles as identities, policies evaluated together, subnets tied to zones, services as separate products, and moving data as a cost. Each gets a chapter later; here is the map.

1. The account is the boundary#

MegaCorp gives payments three AWS accounts: development, test and production. An account holds resources, and it is also the boundary for security, cost and quotas. Nothing crosses it unless someone allows it, costs are counted per account, and quotas apply per account and Region. Many small accounts, grouped with AWS Organizations, take the place of a few large ones. chapter 7 develops this.

On Azure
Where the boundaries sit: on Azure, subscriptions inside management groups hold resource groups; on AWS, accounts inside organizational units are the boundary (on Azure)Microsoft Entra tenant MegaCorpManagement group WorkloadsSubscription payments-prodResource group rg-paymentsStorageaccounts(statements)
On AWS
Where the boundaries sit: on Azure, subscriptions inside management groups hold resource groups; on AWS, accounts inside organizational units are the boundary (on AWS)OrganizationOU WorkloadsAWS Account payments-prodRegion eu-west-2Amazon S3(statements)
Figure Where the boundaries sit: on Azure, subscriptions inside management groups hold resource groups; on AWS, accounts inside organizational units are the boundary#
As text

On Azure:

  • Microsoft Entra tenant MegaCorp
    • Management group Workloads
      • Subscription payments-prod
        • Resource group rg-payments
          • Storage accounts (statements)

On AWS:

  • Organization
    • OU Workloads
      • AWS Account payments-prod
        • Region eu-west-2
          • Amazon S3 (statements)

Read more What is an AWS account? · Terminology and concepts for AWS Organizations · Azure management groups

2. A role is an identity you assume#

Alex, from chapter 1, has no AWS password at all. Alex assumes a role, an identity with its own permissions, and receives credentials that last hours, not years. Code does the same: an instance or a Lambda function assumes a role, so no key sits in a configuration file. chapter 8 develops this.

Assuming a role: the role's trust policy says who may assume it, and AWS STS returns short-term credentialsSettlement jobAWS STSpayments-prodAssumeRole for the settlement role1checks the role's trust policy2short-term credentials3calls AWS with the role's permissions4
Figure Assuming a role: the role's trust policy says who may assume it, and AWS STS returns short-term credentials#
As text
  1. Settlement job to AWS STS: AssumeRole for the settlement role
  2. AWS STS: checks the role's trust policy
  3. AWS STS to Settlement job: short-term credentials
  4. Settlement job to payments-prod: calls AWS with the role's permissions

Read more IAM roles

3. Policies are evaluated together#

Every request starts denied. It succeeds only if some policy allows it and no policy denies it: an explicit deny anywhere, in the caller's policies, the resource's policy or an organization's guardrail, always wins. A request from one account to another must be allowed on both sides. chapter 9 develops this.

Read more Policy evaluation logic

4. A subnet lives in one zone#

A subnet sits in exactly one Availability Zone, and nothing on it says public or private: its route table decides. Resilience is therefore a matter of subnet layout, with one subnet per zone for each tier. chapter 12 develops this.

Read more Subnets for your VPC

5. Every service is its own product#

Each AWS service has its own API, console pages, quotas and pricing, and nothing makes two services name or tag things alike. Tags are optional and case sensitive, so Env and env are two different keys. A naming and tagging scheme is the designer's job, decided early. chapter 31 develops this.

Read more Tagging your AWS resources · Service Quotas documentation

6. Moving data costs money#

Data moving between zones is billed as it leaves and again as it arrives. Data through a NAT gateway is billed per gigabyte processed, and data leaving AWS for the internet is billed too, while data arriving from the internet is free. The figure marks the billed paths of one payments service. chapter 31 develops this.

Which paths are billed: traffic between zones in both directions, NAT processing, and data out to the internetRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24InternetgatewayNAT gatewayAmazon EC2(app)Amazon Aurora(primary)123
1Zone a to zone b: billed as it leaves and as it arrives
2Through the NAT gateway: billed per gigabyte processed
3Out to the internet: billed as data transfer out
Figure Which paths are billed: traffic between zones in both directions, NAT processing, and data out to the internet#
As text
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Internet gateway
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • NAT gateway
        • Private subnet 10.20.10.0/24
          • Amazon EC2 (app)
      • Availability Zone b
        • Private subnet 10.20.11.0/24
          • Amazon Aurora (primary)
  1. Zone a to zone b: billed as it leaves and as it arrives
  2. Through the NAT gateway: billed per gigabyte processed
  3. Out to the internet: billed as data transfer out
Gotcha

Resilience has a running cost. Spreading tiers across zones, as shift 4 asks, creates the zone-to-zone traffic that shift 6 bills. Design for both at once.

Read more Understanding data transfer charges · Amazon VPC pricing

False friends#

Part I · Orientation · Chapter 05·3 min read

Azure and AWS use some of the same words for different things. Six cause most of the confusion in design reviews: role, policy, security group, resource group, gateway and endpoint.

Identity words#

In the payments team's first design review, an architect who knows Azure asks who holds "the Contributor role" on the payments account. On AWS the question has no answer, because the word means something else.

False friend: role
On Azure

A set of permissions, a role definition, assigned to a user, group or managed identity at a scope. Assignments add up.

On AWS

An identity with its own permissions and no password or keys. People and services assume it for a limited time.

False friend: policy
On Azure

Azure Policy checks how resources are configured, such as allowed regions or required tags, whoever makes the change.

On AWS

A JSON document of permissions. IAM, resource-based and organization policies are evaluated together. The nearest match to Azure Policy is AWS Config rules.

Someone says 'policy': which AWS policy do they mean?yesnoyesnoyesnoDoes it say which API actions a user orrole may call?An IAM policyIs it attached to a resource, such as abucket or a queue?A resource-based policyDoes it cap what every identity in anaccount may do?A service control policy, set inAWS OrganizationsTo check how resources are configured,use AWS Config rules
Figure Someone says 'policy': which AWS policy do they mean?#
As text
  1. Does it say which API actions a user or role may call? Yes: An IAM policy. No: the next step.
  2. Is it attached to a resource, such as a bucket or a queue? Yes: A resource-based policy. No: the next step.
  3. Does it cap what every identity in an account may do? Yes: A service control policy, set in AWS Organizations. No: the next step.
  4. To check how resources are configured, use AWS Config rules

Read more IAM roles · Policy evaluation logic · Evaluating resources with AWS Config rules · What is Azure role-based access control? · Overview of Azure Policy

Network words#

Three network words trip people up next.

False friend: security group
On Azure

A network security group: allow and deny rules in priority order, attached to a subnet, a network interface, or both.

On AWS

Allow rules only, attached to a resource's network interface, never to a subnet. A rule can name another security group.

False friend: Application Gateway and API Gateway
On Azure

Application Gateway is a layer 7 load balancer for web traffic, routing on URL path and host, with an optional web application firewall.

On AWS

API Gateway is a front door for REST, HTTP and WebSocket APIs. For Application Gateway's job, AWS uses an Application Load Balancer with AWS WAF.

On Azure
Application Gateway's match on AWS is an Application Load Balancer with AWS WAF, not Amazon API Gateway (on Azure)Virtual network 10.20.0.0/16Subnet 10.20.1.0/24Subnet 10.20.10.0/24ApplicationgatewaysAzure VirtualMachines(app)
On AWS
Application Gateway's match on AWS is an Application Load Balancer with AWS WAF, not Amazon API Gateway (on AWS)Region eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24AWS WAFApplicationLoad BalancerAmazon EC2(app)
Figure Application Gateway's match on AWS is an Application Load Balancer with AWS WAF, not Amazon API Gateway#
As text

On Azure:

  • Virtual network 10.20.0.0/16
    • Subnet 10.20.1.0/24
      • Application gateways
    • Subnet 10.20.10.0/24
      • Azure Virtual Machines (app)

On AWS:

  • Region eu-west-2
    • AWS WAF
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • Application Load Balancer
        • Private subnet 10.20.10.0/24
          • Amazon EC2 (app)
False friend: service endpoint
On Azure

A subnet setting that sends traffic for services such as Azure Storage over the Azure backbone, and lets those services admit the subnet.

On AWS

The URL of a service's API, such as https://dynamodb.us-west-2.amazonaws.com. The match for the Azure feature is a VPC endpoint.

Read more Security groups · Azure network security groups · What is Azure Application Gateway? · AWS service endpoints

Organising words#

The last false friend matters most for estate design. On Azure the team would keep the payments resources in one resource group and delete them together. AWS has resource groups too, but they hold nothing.

False friend: resource group
On Azure

A container for resources with a shared lifecycle. Each resource belongs to exactly one, and deleting the group deletes everything in it.

On AWS

A saved query, over tags or a CloudFormation stack, listing resources in one Region. Resources do not live in it; the account is the container.

On Azure
On Azure a resource group holds its resources; on AWS it is a query over tagged resources that live in the account (on Azure)Subscription payments-prodResource group rg-paymentsAzure VirtualMachines(app)Storageaccounts(statements)
On AWS
On Azure a resource group holds its resources; on AWS it is a query over tagged resources that live in the account (on AWS)AWS Account payments-prodRegion eu-west-2Resource group: app=paymentsAWS Lambda(settle)Amazon S3(statements)
Figure On Azure a resource group holds its resources; on AWS it is a query over tagged resources that live in the account#
As text

On Azure:

  • Subscription payments-prod
    • Resource group rg-payments
      • Azure Virtual Machines (app)
      • Storage accounts (statements)

On AWS:

  • AWS Account payments-prod
    • Region eu-west-2
      • Resource group: app=payments
        • AWS Lambda (settle)
        • Amazon S3 (statements)

Read more What are resource groups? · What is Azure Resource Manager?

Your first week#

Part I · Orientation · Chapter 06·3 min read

Joining an AWS estate, you need four answers fast: who and where you are, what you may do, what is already there, and who changed what. Each takes one command, and AWS's refusals are more helpful than they look.

Who and where you are#

Alex's first real task is the account MegaCorp inherited when it bought a smaller payments company. Before touching anything, Alex confirms the account and Region the profile reaches, and whether the account belongs to MegaCorp's organization.

aws sts get-caller-identity --profile acquired
aws organizations describe-organization --profile acquired
aws ec2 describe-vpcs --profile acquired --query "Vpcs[].{id:VpcId,cidr:CidrBlock,default:IsDefault}"

The commands name the account and role, the organization and its management account, and the VPCs in the profile's Region, flagging any default VPC.

Read more AWS CLI: aws sts get-caller-identity · Terminology and concepts for AWS Organizations · Default VPCs

What you may do#

Nobody hands Alex a list of permissions. Alex tries the task and reads the refusal: most access denied messages name the missing action and the kind of policy that refused it.

User: arn:aws:iam::123456789012:role/HR is not authorized to perform: codecommit:ListRepositories
because no identity-based policy allows the codecommit:ListRepositories action
Reading an access denied message: an explicit deny, a missing allow in a guardrail, or a missing allow in your own roleyesnoyesnoyesnoDoes the message say "with an explicitdeny"?A Deny statement blocks you:adding an allow will not helpDoes it say "no service control policyallows"?The organization's guardrail stopsthe action: its owners decideDoes it say "no identity-based policyallows"?Your role lacks the permission:ask for itAnother policy lacks an allow, such asthe resource's own policy
Figure Reading an access denied message: an explicit deny, a missing allow in a guardrail, or a missing allow in your own role#
As text
  1. Does the message say "with an explicit deny"? Yes: A Deny statement blocks you: adding an allow will not help. No: the next step.
  2. Does it say "no service control policy allows"? Yes: The organization's guardrail stops the action: its owners decide. No: the next step.
  3. Does it say "no identity-based policy allows"? Yes: Your role lacks the permission: ask for it. No: the next step.
  4. Another policy lacks an allow, such as the resource's own policy

Read more Troubleshoot access denied error messages

What is already there#

For an inventory, AWS Resource Explorer, on by default since October 2025, searches the current Region by name, tag or ID, at no charge.

aws resource-explorer-2 search --profile acquired --query-string "tag:app=payments"

Read more What is AWS Resource Explorer?

Who changed what#

On Wednesday the settlement job stops reaching its database. Someone changed a security group on Tuesday, and CloudTrail knows who. Its event history keeps 90 days of management calls in each Region, free and with no trail needed.

aws cloudtrail lookup-events --profile acquired --region eu-west-2 \
    --lookup-attributes AttributeKey=ResourceName,AttributeValue=sg-0123456789abcdef0
Who changed the security group? CloudTrail event history answers from the Region where the change happenedAlexAWS CLIAWS CloudTraillookup-events for the security group1LookupEvents in eu-west-22AuthorizeSecurityGroupIngress, by a rolesession, on Tuesday3who made the call, and when4
Figure Who changed the security group? CloudTrail event history answers from the Region where the change happened#
As text
  1. Alex to AWS CLI: lookup-events for the security group
  2. AWS CLI to AWS CloudTrail: LookupEvents in eu-west-2
  3. AWS CloudTrail to AWS CLI: AuthorizeSecurityGroupIngress, by a role session, on Tuesday
  4. AWS CLI to Alex: who made the call, and when
CloudTrail event history is kept per Region: 90 days of management events in each, with no trail neededAWS Account acquired-paymentsRegion eu-west-2Region eu-west-1AWS CloudTrail(event history)AWS CloudTrail(event history)
Figure CloudTrail event history is kept per Region: 90 days of management events in each, with no trail needed#
As text
  • AWS Account acquired-payments
    • Region eu-west-2
      • AWS CloudTrail (event history)
    • Region eu-west-1
      • AWS CloudTrail (event history)
Gotcha

Event history searches one Region at a time. A change made in eu-west-1 never appears in a search of eu-west-2. To keep events longer than 90 days, or to search them together, the estate needs a trail or an event data store.

Read more Working with CloudTrail event history

The same questions on each cloud#

QuestionAWS CLIAzure CLI
Who am I, and where?aws sts get-caller-identityaz account show
Which organization?aws organizations describe-organizationaz account management-group list
Which networks?aws ec2 describe-vpcsaz network vnet list
What is here?aws resource-explorer-2 searchaz resource list
Who changed what?aws cloudtrail lookup-eventsaz monitor activity-log list

Read more AWS CLI Command Reference

Orientation
  1. An architect asks who holds the Contributor role on the payments account. What do you answer?
    Answer
    That role is an Azure idea. On AWS, ask which roles exist, who may assume each one under its trust policy, and what its policies allow.
  2. A request fails "with an explicit deny in a service control policy". Will adding a permission to your role fix it?
    Answer
    No. An explicit deny always wins, and this one is an organization guardrail, so only its owners can change it.
  3. The payments service in zone a calls its database in zone b all day. What does that cost?
    Answer
    Data between zones is billed as it leaves and again as it arrives, so every call is charged in both directions.

Accounts and Organizations#

Part II · Foundations · Chapter 07·5 min read

On AWS the account is the boundary for security, cost and quotas, so an enterprise runs many small accounts. AWS Organizations groups them, pays for them together and caps what anyone in them may do.

The account is the unit#

MegaCorp's platform team starts with a question that shapes everything else: how many AWS accounts? The answer is more than anyone expects. Each workload gets an account per environment, and shared jobs such as logging and security get accounts of their own.

AWS account
On Azuresubscriptions

What it does. An account holds MegaCorp's resources and is their boundary for security, cost and quotas. Payments gets three: development, test and production.

How it works. Every resource's ARN carries its account ID, and nothing crosses to another account unless policies on both sides allow it. Each account also has a root user with complete access, which no IAM policy in the account can restrict.

When to use it. For each workload and environment, and for each shared job such as logging, security tooling and networking, so that a mistake or a breach stays inside one account.

When not to. For two teams that must share data all day, since every crossing needs policies on both sides. Nothing should run in the management account, which service control policies cannot restrict.

Limits. An organization starts with a quota of 10 accounts as of September 2026, raised on request up to 50,000. Most other quotas, such as 5 VPCs per Region, apply per account, so splitting accounts splits the quotas too.

Cost. Resources are billed to the account that holds them, and the organization's management account pays the whole bill.

Security. Never use the root user for daily work, and protect it with MFA. With centralized root access, the management account can remove root credentials from member accounts altogether.

Gotchas. A closed account can be reopened for 90 days. After that it is gone for good: its ID is never reused, its email address can never register another account, and until then it still counts against the organization's quota.

Read more What is an AWS account? · AWS account root user · Close an AWS account

An organization of accounts#

Twenty accounts need somewhere to live. MegaCorp creates an organization, keeps its management account empty, and groups the others by how they are governed: security accounts in one organizational unit, workloads in another, experiments in a third.

MegaCorp's organization: accounts grouped in organizational units under one root, with the management account holding nothing but the organizationOrganization rootAWS Account managementSecurity OUAWS Account log-archiveAWS Account security-toolingWorkloads OUAWS Account payments-devAWS Account payments-prodSandbox OUAWS Account sandbox-alexAWSOrganizationsAmazon S3(logs)AmazonGuardDutyAWS Fargate(payments)AWS Fargate(payments)Amazon EC2(experiments)
Figure MegaCorp's organization: accounts grouped in organizational units under one root, with the management account holding nothing but the organization#
As text
  • Organization root
    • AWS Account management
      • AWS Organizations
    • Security OU
      • AWS Account log-archive
        • Amazon S3 (logs)
      • AWS Account security-tooling
        • Amazon GuardDuty
    • Workloads OU
      • AWS Account payments-dev
        • AWS Fargate (payments)
      • AWS Account payments-prod
        • AWS Fargate (payments)
    • Sandbox OU
      • AWS Account sandbox-alex
        • Amazon EC2 (experiments)
AWS Organizations
On AzureAzure management groups

What it does. AWS Organizations groups MegaCorp's accounts under one management account, pays their bills together, and applies guardrails to whole groups of accounts at once.

How it works. Accounts sit under a single root, in organizational units nested up to five levels deep. A service control policy attached to the root or an OU caps what every user and role below it may do. It never grants anything.

When to use it. From the second account onwards. Group accounts by how they are governed rather than by the org chart; AWS suggests starting with Security, Infrastructure, Workloads and Sandbox OUs.

When not to. As a substitute for IAM. Service control policies only cap permissions: people and code still need IAM policies before they can do anything at all.

Limits. Default quotas as of September 2026: one root, 2,000 OUs, 10 service control policies attached to each root, OU or account, and 10,240 characters in each policy.

Cost. No additional charge; each account's resources are billed as usual, on one bill.

Security. Service control policies do not apply to the management account, so it holds nothing but the organization. Each account it creates gets OrganizationAccountAccessRole, which gives the management account full administrative control.

Gotchas. Every OU and account must keep at least one service control policy. Removing the default FullAWSAccess without a replacement makes every action in the member accounts fail.

Read more AWS Organizations documentation · Terminology and concepts for AWS Organizations · Quotas and service limits for AWS Organizations · Recommended OUs and accounts

Guardrails with service control policies#

Guardrails are where Organizations earns its place. The platform team attaches a policy to the Workloads OU that denies every Region except eu-west-1 and eu-west-2. A developer with full administrator rights in payments-dev still cannot start an instance in us-east-1, because every policy from the root down must allow an action, and none may deny it.

Can a role in payments-dev do this? Every service control policy from the root down must allow the action, none may deny it, and IAM must still grant ityesnonoyesnoyesDoes any service control policy betweenthe root and payments-dev deny theaction?Denied, even for the account'sroot userDoes every service control policy onthat path allow it?Denied by a guardrail above theaccountDoes an IAM policy grant it to therole?Denied: service control policiesnever grantAllowed
Figure Can a role in payments-dev do this? Every service control policy from the root down must allow the action, none may deny it, and IAM must still grant it#
As text
  1. Does any service control policy between the root and payments-dev deny the action? Yes: Denied, even for the account's root user. No: the next step.
  2. Does every service control policy on that path allow it? No: Denied by a guardrail above the account. Yes: the next step.
  3. Does an IAM policy grant it to the role? No: Denied: service control policies never grant. Yes: the next step.
  4. Allowed
Gotcha

Test a guardrail before it reaches the root. A policy that denies too much at the root locks out every member account at once, the security team's included. Try it first on an OU holding a few accounts.

Read more Service control policies (SCPs) · Creating a member account in an organization

MegaCorp's accounts#

The payments-prod account, in the Workloads OU, becomes the outer boundary of MegaCorp's running design. Every later chapter adds something inside it.

MegaCorp's running design, showing the accounts layer this chapter addsAWS Account payments-prodRegion eu-west-2
Figure MegaCorp's running design, showing the accounts layer this chapter adds#
As text
  • AWS Account payments-prod
    • Region eu-west-2

Read more What is an AWS account?

IAM identities and roles#

Part II · Foundations · Chapter 08·5 min read

Every request to AWS is signed by an identity, and the identities that matter are roles: people and code assume them and receive credentials that expire. A long-term key becomes the exception you have to justify.

Who is calling#

MegaCorp's settlement job has to write statements to Amazon S3. On a first attempt, a developer pastes an access key into the job's configuration. The platform team rejects the change: on AWS, code gets its access from a role, and a key in a file can leak and never expires.

AWS IAM
On AzureAzure role-based access control, Microsoft Entra ID

What it does. AWS IAM decides who may do what in an account. The settlement job may write statements to one bucket, and do nothing else.

How it works. IAM holds identities, mainly roles, and the policies attached to them. Every request is signed with an identity's credentials, and AWS checks every applicable policy before the service acts.

When to use it. Always: every account uses it. Give roles to code and to people signing in through IAM Identity Center, and keep IAM users for the rare tool that can use neither.

When not to. Managing people account by account; IAM Identity Center does that centrally, as chapter 10 shows. The customers of an application belong in Amazon Cognito, not in IAM.

Limits. Defaults as of September 2026: 1,000 roles per account, raised up to 10,000, and 20 managed policies per role, up to 25. A role's inline policies total at most 10,240 characters.

Cost. No additional charge. IAM Access Analyzer's unused-access analysis and custom policy checks are billed.

Security. Grant least privilege, and let IAM Access Analyzer propose policies from what a role actually used. Any human IAM user left needs MFA, ideally a passkey or security key.

Gotchas. IAM is eventually consistent, so a role created a moment ago may not work yet. Never create IAM resources on a critical path, such as during a failover.

Read more AWS IAM documentation · Security best practices in IAM · IAM and AWS STS quotas

Which identity should this caller use? People federate, code on AWS gets a role from its service, and long-term keys come lastyesnoyesnoyesnoyesnoIs it a person in MegaCorp's workforce?Sign in through IAM IdentityCenter and assume a roleIs it code running on an AWS service,such as Amazon EC2 or AWS Lambda?A role the service assumes for itIs it code outside AWS with an identityprovider or certificates?Federation with OpenID Connect orSAML, or IAM Roles AnywhereIs it another company's account, suchas a monitoring vendor?A cross-account role with anexternal IDOnly then an IAM user with access keys,rotated and watched
Figure Which identity should this caller use? People federate, code on AWS gets a role from its service, and long-term keys come last#
As text
  1. Is it a person in MegaCorp's workforce? Yes: Sign in through IAM Identity Center and assume a role. No: the next step.
  2. Is it code running on an AWS service, such as Amazon EC2 or AWS Lambda? Yes: A role the service assumes for it. No: the next step.
  3. Is it code outside AWS with an identity provider or certificates? Yes: Federation with OpenID Connect or SAML, or IAM Roles Anywhere. No: the next step.
  4. Is it another company's account, such as a monitoring vendor? Yes: A cross-account role with an external ID. No: the next step.
  5. Only then an IAM user with access keys, rotated and watched

Roles for code#

The settlement job gets a role called settlement. Its trust policy lets the compute service assume it, and its permissions policy allows one action on one bucket. The code carries no key: the AWS SDK finds the role's credentials, and they renew before they expire.

Code on an EC2 instance never holds a key: the SDK takes short-term role credentials from instance metadata, and they renew before they expireSettlement jobAWS SDK for JavaInstance metadataAmazon S3putObject1credentials for theinstance's role2short-term credentials3PutObject, signed with them4renews them before they expire5
Figure Code on an EC2 instance never holds a key: the SDK takes short-term role credentials from instance metadata, and they renew before they expire#
As text
  1. Settlement job to AWS SDK for Java: putObject
  2. AWS SDK for Java to Instance metadata: credentials for the instance's role
  3. Instance metadata to AWS SDK for Java: short-term credentials
  4. AWS SDK for Java to Amazon S3: PutObject, signed with them
  5. Instance metadata: renews them before they expire
AWS Security Token Service
On AzureMicrosoft Entra ID

What it does. AWS STS issues the short-term credentials behind every role. When the settlement role is assumed, STS returns a key, a secret and a session token, all of which expire.

How it works. A caller asks STS to assume a role. STS checks the role's trust policy and returns credentials valid for one hour unless the caller asks for longer, up to the role's maximum of at most 12 hours. Compute services make that call for your code.

When to use it. Whenever code or people need access: assuming a role in the same account or another, or exchanging a token from an identity provider for AWS credentials.

When not to. Never trade it for long-term keys to save effort. A vendor that monitors your accounts gets a cross-account role, not an IAM user of its own.

Limits. as of September 2026, 600 requests a second per account and Region, shared by AssumeRole, GetCallerIdentity and others. Calls that AWS services make for you do not count.

Cost. No additional charge.

Security. For a third party, require an external ID in the trust policy. It stops another customer of the same vendor from borrowing your role: the confused deputy problem.

Gotchas. Role chaining, one role assuming another, caps the session at one hour. Use Regional endpoints; the AWS SDK for Java 2.x and AWS CLI v2 already do.

Read more AWS STS API reference · AWS STS Regional endpoints · The confused deputy problem

Gotcha

Launching code with a role needs iam:PassRole. Without that check, anyone allowed to start an instance could hand it a more powerful role and borrow its rights. Grant iam:PassRole only for the roles a team should use.

Read more Use an IAM role for applications on Amazon EC2 · IAM roles

Crossing accounts#

The settlement job must also copy each statement into the log-archive account, which it cannot reach by default. Log-archive creates a role that trusts the settlement role: the job assumes it (step 1) and writes with the new credentials (step 2).

Crossing accounts: the settlement role in payments-prod assumes a role in log-archive whose trust policy names it, then writes with that role's permissionsAWS Account payments-prodAWS Account log-archiveIAM role(settlement)IAM role(statement-writer)Amazon S3(statements)12
1The settlement role assumes statement-writer, whose trust policy names it
2It writes the statement with statement-writer's permissions
Figure Crossing accounts: the settlement role in payments-prod assumes a role in log-archive whose trust policy names it, then writes with that role's permissions#
As text
  • AWS Account payments-prod
    • IAM role (settlement)
  • AWS Account log-archive
    • IAM role (statement-writer)
    • Amazon S3 (statements)
  1. The settlement role assumes statement-writer, whose trust policy names it
  2. It writes the statement with statement-writer's permissions
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "AWS": "arn:aws:iam::111111111111:role/settlement" },
    "Action": "sts:AssumeRole"
  }]
}

That is statement-writer's trust policy. The settlement role also needs permission of its own to call sts:AssumeRole on statement-writer: one account trusts, and the other allows.

Read more Cross-account policy evaluation logic

IAM policies#

Part II · Foundations · Chapter 09·5 min read

Permissions on AWS come from policies: JSON documents attached to identities, resources and whole accounts. Several kinds are evaluated together on every request, and knowing their order is how you design access and fix a refusal.

Nine kinds of policy#

The settlement role from chapter 8 needs one permission: to write objects to the statements bucket. The platform team writes it as an identity-based policy, the commonest kind, attached to the role.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": "s3:PutObject",
    "Resource": "arn:aws:s3:::megacorp-statements/*"
  }]
}

That is one of nine kinds of policy AWS evaluates. Most are JSON with the same shape: an effect, actions, resources and optional conditions. The kinds an architect meets most often:

KindAttached toWhat it does
Identity-baseda user, group or rolegrants the identity actions on resources
Resource-baseda resource, such as a bucket, queue or rolegrants named principals, even in other accounts, actions on that resource
Service and resource control policiesan organization root, OU or accountcaps what identities and resources below may do; never grants
Permissions boundarya user or rolecaps what that identity's own policies can grant; never grants
Session policyone role sessionnarrows a session below the role's permissions
VPC endpoint policya VPC endpointlimits what can pass through that endpoint

Prefer customer managed policies, which many roles can share and the team controls. AWS managed policies are quick to start with, but AWS changes them as services grow, and they rarely match least privilege.

Read more Policies and permissions in IAM · Managed policies and inline policies

How AWS decides#

When the settlement job calls PutObject, AWS gathers every policy that applies and starts from a refusal. Within one account, the request succeeds only if nothing denies it, the guardrails allow it, and the role's policy or the bucket's policy allows it.

How AWS evaluates a request within one account: an explicit deny wins, guardrails and caps must allow, and some policy must grantyesnonoyesnoyesnoyesDoes any policy that applies say Deny?Denied: an explicit deny alwayswinsDo the organization's service andresource control policies allow it?Denied by a guardrailDoes the caller's identity-basedpolicy, or the resource's own policy,allow it?Denied: nothing granted itDo the caller's permissions boundaryand session policy, where set, allowit?Denied by a cap on the callerAllowed
Figure How AWS evaluates a request within one account: an explicit deny wins, guardrails and caps must allow, and some policy must grant#
As text
  1. Does any policy that applies say Deny? Yes: Denied: an explicit deny always wins. No: the next step.
  2. Do the organization's service and resource control policies allow it? No: Denied by a guardrail. Yes: the next step.
  3. Does the caller's identity-based policy, or the resource's own policy, allow it? No: Denied: nothing granted it. Yes: the next step.
  4. Do the caller's permissions boundary and session policy, where set, allow it? No: Denied by a cap on the caller. Yes: the next step.
  5. Allowed
Where the policies on one request sit: guardrails on the organizational unit, an identity-based policy on the role, and a bucket policy on the bucketWorkloads OU: service control policiesAWS Account payments-prodRegion eu-west-2IAM role(settlement)Amazon S3(megacorp-statements)1
1PutObject: the guardrails, the role's policy and the bucket's policy all agree
Figure Where the policies on one request sit: guardrails on the organizational unit, an identity-based policy on the role, and a bucket policy on the bucket#
As text
  • Workloads OU: service control policies
    • AWS Account payments-prod
      • Region eu-west-2
        • IAM role (settlement)
        • Amazon S3 (megacorp-statements)
  1. PutObject: the guardrails, the role's policy and the bucket's policy all agree

Across accounts, as when the job writes to log-archive, AWS evaluates the request in both accounts, and each must allow it on its own.

Read more Policy evaluation logic · Cross-account policy evaluation logic

Tags as permissions#

MegaCorp soon has queues for payments, ledger and fraud, each with a developer role. Writing a policy per team per queue does not scale. Instead the team tags every role and resource with team, and one policy allows an action whenever the resource's team tag matches the caller's. This is attribute-based access control.

Attribute-based access control: one policy lets a role reach resources whose team tag matches its ownAWS Account payments-devRegion eu-west-2IAM role(team=payments)Amazon SQS(team=payments)IAM role(team=ledger)Amazon SQS(team=ledger)12
1Tags match: the payments role reads the payments queue
2Tags match: the ledger role reads the ledger queue
Figure Attribute-based access control: one policy lets a role reach resources whose team tag matches its own#
As text
  • AWS Account payments-dev
    • Region eu-west-2
      • IAM role (team=payments)
      • Amazon SQS (team=payments)
      • IAM role (team=ledger)
      • Amazon SQS (team=ledger)
  1. Tags match: the payments role reads the payments queue
  2. Tags match: the ledger role reads the ledger queue

A new queue tagged team=payments is reachable at once, with no policy change, which is the point: permissions scale as resources are added. The risk moves to the tags, so who may set them must itself be controlled.

Read more Define permissions based on attributes with ABAC

Delegating safely#

Developers in payments-dev want to create roles for their own Lambda functions without waiting for the platform team. A permissions boundary makes that safe: developers may create roles only if they attach the developer-boundary policy, so no role they create can do more than that boundary allows, whatever its own policy says.

Gotcha

A boundary caps, and never grants. A role with a boundary and no permissions policy can do nothing at all. It needs both, and it gets only what both allow.

Read more Permissions boundaries for IAM entities

Checking policies#

Policies drift: a bucket gets shared for one migration and stays shared. MegaCorp turns on an analyzer for the whole organization before the first workload goes live.

IAM Access Analyzer
On Azureno direct equivalent

What it does. IAM Access Analyzer finds access MegaCorp did not mean to grant: a bucket shared outside the organization, a role nobody uses, a policy broader than its job.

How it works. An analyzer takes the organization or one account as its zone of trust, and reasons over resource-based policies to flag any that let a principal outside it in. Other analyzers find unused roles and permissions, and it can draft a policy from CloudTrail activity.

When to use it. From the first account: run an organization-wide external access analyzer in every Region in use, and validate each new policy before it ships.

When not to. As proof that a policy suits its job. It catches over-sharing and mistakes, not a permission the job is missing.

Limits. External access analysis is Regional, so each Region needs its own analyzer. A changed policy is analysed within about 30 minutes.

Cost. External access analysis is free. Unused access analysis is billed for each role and user analysed each month, and custom policy checks for each request.

Security. Each finding shows who outside the zone of trust can reach what. Archive it if the access is intended; otherwise, remove the access.

Gotchas. A policy generated from CloudTrail holds only the actions used in the chosen period, so a quarterly job that has not yet run will be missing its permissions.

Read more Using IAM Access Analyzer

Identity for people and code#

Part II · Foundations · Chapter 10·6 min read

Three kinds of caller sign in to MegaCorp's systems: its people, its pipelines and its customers. Each has its own door: IAM Identity Center, OpenID Connect federation and Amazon Cognito.

People: one sign-in for every account#

Alex's first day on the payments team starts with a login MegaCorp already has: a Microsoft Entra ID account, used for email. Nobody creates an IAM user for Alex in each account. The platform team connected Entra ID to IAM Identity Center once, and Alex's group membership does the rest.

In 2013 IAM learned to trust a corporate identity provider over SAML, but account by account. From 2017, AWS Single Sign-On did it once for the whole organization. Renamed IAM Identity Center in 2022, it is why the CLI still says aws sso login.

Signing in to AWS, from federation set up in each account to one sign-in for the whole organizationIAM: SAML federation,account by account2013Amazon Cognito:identities for appusers2014AWS Single Sign-On:one sign-in acrossaccounts2017IAM Identity Center:the same service,renamed2022
Figure Signing in to AWS, from federation set up in each account to one sign-in for the whole organization#
As text
  1. 2013: IAM: SAML federation, account by account
  2. 2014: Amazon Cognito: identities for app users
  3. 2017: AWS Single Sign-On: one sign-in across accounts
  4. 2022: IAM Identity Center: the same service, renamed
Workforce sign-in: Entra ID copies the group to IAM Identity Center and signs its members in, and each permission set becomes a role in the accounts it is assigned toMicrosoft Entra IDAWS Account managementWorkloads OUAWS Account payments-devAWS Account payments-prodpayments-developersAWS IAMIdentity CenterIAM role(PaymentsDeveloper)IAM role(PaymentsReadOnly)123
1SCIM copies the group and its members; SAML signs them in
2The PaymentsDeveloper permission set becomes a role in payments-dev
3A read-only permission set becomes a role in payments-prod
Figure Workforce sign-in: Entra ID copies the group to IAM Identity Center and signs its members in, and each permission set becomes a role in the accounts it is assigned to#
As text
  • Microsoft Entra ID
    • payments-developers
  • AWS Account management
    • AWS IAM Identity Center
  • Workloads OU
    • AWS Account payments-dev
      • IAM role (PaymentsDeveloper)
    • AWS Account payments-prod
      • IAM role (PaymentsReadOnly)
  1. SCIM copies the group and its members; SAML signs them in
  2. The PaymentsDeveloper permission set becomes a role in payments-dev
  3. A read-only permission set becomes a role in payments-prod
AWS IAM Identity Center
On AzureMicrosoft Entra ID

What it does. IAM Identity Center gives MegaCorp's people one sign-in for every account. Alex signs in with Entra ID and picks payments-dev or payments-prod in the AWS access portal.

How it works. It trusts an identity source, here Entra ID, which signs people in over SAML and copies users and groups over SCIM. A permission set is a template of policies. Assign it to a group in chosen accounts, and Identity Center creates and maintains a matching role in each.

When to use it. For everyone who works in AWS, from the first account. It needs an organization instance, which always lives in the management account; administer it from a delegated member account.

When not to. For code, which uses roles, or for the customers of an application, who belong in Amazon Cognito.

Limits. Defaults as of September 2026: one instance per account; 3,500 permission sets, 500 of them in any one account; one inline policy per permission set; 100 groups per permission set in each account.

Cost. No extra charge.

Security. Assign groups, not people, so someone removed in Entra ID can no longer sign in. In the management account, assign people directly, or whoever edits the group decides who reaches it. Sessions last an hour by default, at most 12.

Gotchas. Entra ID provisions only the direct members of an assigned group, not members of nested groups. Removing a user's attribute in Entra ID leaves it in Identity Center, which matters when attributes drive access.

Read more AWS IAM Identity Center documentation · Manage AWS accounts with permission sets · Configure SAML and SCIM with Microsoft Entra ID and IAM Identity Center · Quotas and limits in IAM Identity Center

Pipelines: nothing to steal#

MegaCorp deploys the payments service from GitHub Actions. The first pipeline kept an access key in a repository secret, which works until the key leaks. Now each run gets a signed OpenID Connect token from GitHub and trades it with AWS STS for credentials that expire.

A pipeline without keys: GitHub signs a token for each run, and AWS STS exchanges it for a role's credentials if the role's trust policy accepts the tokenDeploy jobGitHubAWS STSa token for this run1signed: repository, branch, audience2AssumeRoleWithWebIdentity with the token3checks the signature and the deploy role'strust policy4credentials for the deploy role5
Figure A pipeline without keys: GitHub signs a token for each run, and AWS STS exchanges it for a role's credentials if the role's trust policy accepts the token#
As text
  1. Deploy job to GitHub: a token for this run
  2. GitHub to Deploy job: signed: repository, branch, audience
  3. Deploy job to AWS STS: AssumeRoleWithWebIdentity with the token
  4. AWS STS: checks the signature and the deploy role's trust policy
  5. AWS STS to Deploy job: credentials for the deploy role

The account registers GitHub as an identity provider once. The deploy role's trust policy then accepts only runs on the payments repository's main branch.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Federated": "arn:aws:iam::111111111111:oidc-provider/token.actions.githubusercontent.com" },
    "Action": "sts:AssumeRoleWithWebIdentity",
    "Condition": {
      "StringEquals": {
        "token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
        "token.actions.githubusercontent.com:sub": "repo:megacorp/payments:ref:refs/heads/main"
      }
    }
  }]
}
Gotcha

The sub condition is the lock. Without it, workflows in repositories MegaCorp does not control could assume the role. IAM refuses a GitHub trust policy with no sub condition or a bare wildcard, but repo:megacorp/* still admits every repository in the organization.

Read more Create a role for OpenID Connect federation

Customers: a directory of their own#

Next comes a merchant portal where shops download their settlement statements. Merchants are not staff, so they belong in neither Entra ID nor IAM, but in a directory of their own.

Amazon Cognito
On AzureMicrosoft Entra External ID

What it does. Amazon Cognito signs in an application's own users. Merchants sign up for the portal and sign in with multi-factor authentication, and the portal receives standard OpenID Connect tokens.

How it works. A user pool is the directory: it stores users, signs them in and issues tokens. It can also federate, so a large merchant signs in through its own SAML or OIDC provider. An optional identity pool exchanges a token for short-term AWS credentials, for apps that call AWS directly.

When to use it. Sign-up and sign-in for the customers or partners of a web or mobile app, when you want a managed directory rather than running one.

When not to. For MegaCorp's staff, who use IAM Identity Center. For machine-to-machine calls at volume, check the bill first: each token is charged.

Limits. Defaults as of September 2026, per account and Region: 1,000 user pools, 1,000 app clients per pool and 40 million users per pool, all adjustable. API request rates have quotas by category.

Cost. User pools are billed per monthly active user at the pool's feature plan: Lite, Essentials or Plus. Lite and Essentials have a free tier, and federated users have their own rate. Machine-to-machine tokens are billed per token. Identity pools are free.

Security. Require multi-factor authentication, and consider the Plus plan's threat protection when accounts hold money. Give each app client access only to the attributes it needs.

Gotchas. Some choices are fixed when the pool is created: username or email sign-in, which attributes are required, and every custom attribute. A wrong guess means a new pool and a migration. Identify users by sub, never by email.

Read more Amazon Cognito documentation · What is Amazon Cognito? · Amazon Cognito pricing · Quotas in Amazon Cognito · Working with user attributes

MegaCorp's design gains its identity layer: merchants sign in through a user pool in eu-west-2, and the pipeline deploys through the deploy role.

MegaCorp's running design with the identity layer this chapter adds: the merchants' user pool and the pipeline's deploy roleAWS Account payments-prodRegion eu-west-2IAM, global to the accountCustomersAmazon Cognito(merchants)IAM role(deploy)1
1Merchants sign in through the user pool, which issues their tokens
Figure MegaCorp's running design with the identity layer this chapter adds: the merchants' user pool and the pipeline's deploy role#
As text
  • Customers
  • AWS Account payments-prod
    • Region eu-west-2
      • Amazon Cognito (merchants)
    • IAM, global to the account
      • IAM role (deploy)
  1. Merchants sign in through the user pool, which issues their tokens

Regions and zones#

Part II · Foundations · Chapter 11·3 min read

Every resource lives somewhere: in one zone, one Region, or everywhere. Which of the three decides what a failure takes down, where data may sit, and why a Region in Virginia matters to all.

Choosing a Region#

MegaCorp's merchant data must stay in the UK, so the payments team builds in eu-west-2, London, after checking the Regional Services List for every service the design needs. Resources and data stay in their Region unless you copy them, and the console shows one Region at a time.

Regions launched after 20 March 2019, such as Europe (Spain), are opt-in: nobody can use one until it is enabled, and enabling copies the account's IAM data there, which can take hours.

Gotcha

Disabling a Region deletes nothing. The resources in a disabled opt-in Region remain and keep billing; only access to them is lost. Remove them first.

Read more Enable or disable AWS Regions in your account · AWS Regional Services List

Zone names differ between accounts#

The settlement job in payments-prod reads the ledger database in another account. Both teams chose eu-west-2a to keep traffic in one zone, yet the bill shows traffic between zones: AWS maps zone names to physical zones at random for each account.

Zone names are per account: both teams chose eu-west-2a, but the zone IDs show two physical zones, so every read crosses zonesAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone a (euw2-az2)AWS Account ledgerRegion eu-west-2VPC 10.40.0.0/16Availability Zone a (euw2-az1)Amazon EC2(settlementjob)Amazon RDS(ledger)1
1Both sit in eu-west-2a, yet in different zones: each read is billed as it leaves one zone and enters the other
Figure Zone names are per account: both teams chose eu-west-2a, but the zone IDs show two physical zones, so every read crosses zones#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Availability Zone a (euw2-az2)
          • Amazon EC2 (settlement job)
  • AWS Account ledger
    • Region eu-west-2
      • VPC 10.40.0.0/16
        • Availability Zone a (euw2-az1)
          • Amazon RDS (ledger)
  1. Both sit in eu-west-2a, yet in different zones: each read is billed as it leaves one zone and enters the other

A zone ID, such as euw2-az1, names the same physical zone in every account. When accounts must agree on a zone, plan and automate with zone IDs.

False friend: Availability Zone
On Azure

Zone numbers also map differently in each subscription, but data moving between zones in a region is not charged.

On AWS

Data moving between zones in a Region is billed, as it leaves one zone and as it arrives in the other.

Read more Availability Zone IDs for your AWS resources · What are Azure availability zones?

Global services, and us-east-1#

IAM, AWS Organizations, Amazon Route 53 and Amazon CloudFront are global: they answer everywhere, but their control planes sit in us-east-1, in northern Virginia.

A global service seen from London: creating a role goes through us-east-1, while requests signed with the role are checked in eu-west-2AlexIAM, us-east-1Settlement jobAmazon S3, eu-west-2CreateRole1copies the role to every Region, a momentlater2PutObject, signed with therole3checked in the Region4
Figure A global service seen from London: creating a role goes through us-east-1, while requests signed with the role are checked in eu-west-2#
As text
  1. Alex to IAM, us-east-1: CreateRole
  2. IAM, us-east-1: copies the role to every Region, a moment later
  3. Settlement job to Amazon S3, eu-west-2: PutObject, signed with the role
  4. Amazon S3, eu-west-2 to Settlement job: checked in the Region

So keep changes to global services out of recovery plans, as chapter 8 said of IAM, and manage CloudFront certificates in us-east-1.

Three scopes: a NAT gateway lives in one zone, a bucket in one Region, and IAM everywhereRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24AWS IAM(global)Amazon S3(regional)NAT gateway(zonal)
Figure Three scopes: a NAT gateway lives in one zone, a bucket in one Region, and IAM everywhere#
As text
  • AWS IAM (global)
  • Region eu-west-2
    • Amazon S3 (regional)
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • NAT gateway (zonal)
ScopeExamplesWhat it means for a design
Zonala subnet, an EC2 instance, a NAT gatewayfails with its zone, so run one in each of at least two zones
Regionala VPC, an S3 bucket, a Cognito user poolspans the Region's zones; another Region needs its own copy
GlobalIAM, Organizations, Route 53, CloudFrontanswers everywhere; changes go through one Region

Read more AWS Fault Isolation Boundaries: global services

Closer than a Region#

LocationWhere it isUse it for
Local Zonean extension of a Region, near a citylow latency, or data that must stay local; enable it, then add a subnet
Wavelength Zoneinside a carrier's 5G networkvery low latency to mobile devices
AWS OutpostsAWS racks or servers in your own data centreAWS services on premises, managed as part of a Region

Read more Regions and Zones · How AWS Local Zones work

VPC fundamentals#

Part II · Foundations · Chapter 12·9 min read

A VPC is your private network in one AWS Region. Three decisions shape every VPC: how to divide its addresses, which zone each subnet lives in, and how traffic leaves for the internet and for AWS services.

A network in one Region#

MegaCorp's payments team is moving its platform to AWS, into the London Region, eu-west-2. First it needs a network to put the platform in. MegaCorp's address plan, drawn up for the whole estate so that no two networks overlap, gives payments the range 10.20.0.0/16.

From zero: CIDR block

A CIDR block writes an address range as a base address and a prefix length. 10.20.0.0/16 fixes the first 16 of 32 bits and leaves 16 free: 65,536 addresses. Each extra bit of prefix halves the range, so a /24 holds 256. Ranges that share any address overlap.

Amazon VPC
On AzureAzure Virtual Network

What it does. An Amazon VPC (virtual private cloud) is the team's own network in one Region, isolated from every other network in AWS. Everything the platform runs takes its private address from it.

How it works. The team creates it with 10.20.0.0/16; any range from /16 to /28 is allowed. Subnets divide the range, one zone each, and route tables steer their traffic.

When to use it. Anything with a private address. Each environment gets its own VPC, so test can never touch production.

When not to. The default VPC that AWS created in the Region has only public subnets: fine for an experiment, not for payments.

Limits. By default as of September 2026: 5 VPCs per Region, and 5 IPv4 ranges and 200 subnets per VPC, all adjustable.

Cost. The VPC is free; traffic is not. Public IPv4 addresses cost by the hour, even when idle, and data between zones is billed as it leaves and again as it arrives.

Security. Without a route to an internet gateway, nothing on the internet can reach a subnet. VPC Flow Logs record who talked to whom; VPC Block Public Access can forbid internet access across the account.

Gotchas. A range cannot be resized, only added to, so the team took a /16 from the plan instead of guessing. It avoids 172.17.0.0/16, which some AWS services use.

Read more Amazon VPC documentation · Amazon VPC quotas · Amazon VPC pricing

A subnet lives in one zone#

Next come the subnets, and the first rule that shapes every AWS network: a subnet sits in exactly one Availability Zone and cannot span zones. To survive the loss of a zone, the team creates matching subnets in two zones and spreads each tier across them.

On Azure
One tier on each cloud: an Azure subnet spans the region's zones, while an AWS subnet sits in one zone, so subnets repeat (on Azure)UK SouthVirtual network 10.20.0.0/16Subnet 10.20.10.0/24Azure VirtualMachines(zone 1)Azure VirtualMachines(zone 2)
On AWS
One tier on each cloud: an Azure subnet spans the region's zones, while an AWS subnet sits in one zone, so subnets repeat (on AWS)Region eu-west-2VPC 10.20.0.0/16Availability Zone aPrivate subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24Amazon EC2(app)Amazon EC2(app)
Figure One tier on each cloud: an Azure subnet spans the region's zones, while an AWS subnet sits in one zone, so subnets repeat#
As text

On Azure:

  • UK South
    • Virtual network 10.20.0.0/16
      • Subnet 10.20.10.0/24
        • Azure Virtual Machines (zone 1)
        • Azure Virtual Machines (zone 2)

On AWS:

  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Private subnet 10.20.10.0/24
          • Amazon EC2 (app)
      • Availability Zone b
        • Private subnet 10.20.11.0/24
          • Amazon EC2 (app)
QuestionAWSAzure
Where a subnet livesone Availability Zonethe whole region, across its zones
IPv4 subnet sizes/16 to /28/29 to /2
Reserved addresses in each subnetthe first four and the lastthe first four and the last

Read more Subnets for your VPC · Azure Virtual Network FAQ

Routing makes a subnet public or private#

The team wants public subnets for the load balancer and the NAT gateways, and private ones for everything else. Nothing on a subnet says "public", though: a public subnet is one whose route table has a direct route to an internet gateway.

A subnet's type follows from its route tableyesnonoyesyesnoDoes its route table have a route to aninternet gateway?Public subnetDoes any route lead out of the VPC?Isolated subnet, reachable onlyinside the VPCDoes it route to a Site-to-Site VPNthrough a virtual private gateway?VPN-only subnetPrivate subnet: it reaches theinternet only through a NATdevice
Figure A subnet's type follows from its route table#
As text
  1. Does its route table have a route to an internet gateway? Yes: Public subnet. No: the next step.
  2. Does any route lead out of the VPC? No: Isolated subnet, reachable only inside the VPC. Yes: the next step.
  3. Does it route to a Site-to-Site VPN through a virtual private gateway? Yes: VPN-only subnet. No: the next step.
  4. Private subnet: it reaches the internet only through a NAT device

Read more Configure route tables · Internet gateways

Getting out: NAT gateways#

The settlement job is the first to need the internet. The figure shows the team's layout.

Each private subnet leaves through the NAT gateway in its own zone, and reaches Amazon S3 through a gateway endpointRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24Amazon S3InternetgatewayVPC endpoints(gateway)NAT gatewayAmazon EC2(app)NAT gatewayAmazon EC2(app)1234
1Internet traffic from zone a goes to the NAT gateway in zone a
2The NAT gateway sends it out through the internet gateway
3Zone b routes to its own NAT gateway, so a failure stays in one zone
4Traffic for Amazon S3 takes the gateway endpoint's route, not the NAT gateway
Figure Each private subnet leaves through the NAT gateway in its own zone, and reaches Amazon S3 through a gateway endpoint#
As text
  • Region eu-west-2
    • Amazon S3
    • VPC 10.20.0.0/16
      • Internet gateway
      • VPC endpoints (gateway)
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • NAT gateway
        • Private subnet 10.20.10.0/24
          • Amazon EC2 (app)
      • Availability Zone b
        • Public subnet 10.20.1.0/24
          • NAT gateway
        • Private subnet 10.20.11.0/24
          • Amazon EC2 (app)
  1. Internet traffic from zone a goes to the NAT gateway in zone a
  2. The NAT gateway sends it out through the internet gateway
  3. Zone b routes to its own NAT gateway, so a failure stays in one zone
  4. Traffic for Amazon S3 takes the gateway endpoint's route, not the NAT gateway
NAT gateway
On AzureAzure NAT Gateway

What it does. The settlement job runs in a private subnet but must call the card scheme's API. A NAT gateway lets it connect out, while nothing on the internet can connect in.

How it works. The job's internet traffic goes to the NAT gateway in its zone (step 1), which sends it out from its Elastic IP address (step 2). Regional NAT gateways, added in November 2025, span the zones themselves.

When to use it. Private workloads that call the internet: the card scheme, package repositories, partners.

When not to. Traffic for AWS services. Sending the nightly statements to Amazon S3 this way bills every gigabyte as processed; the next section avoids that.

Limits. It scales to 100 Gbps and ten million packets a second, then drops packets. One address holds 55,000 open connections to one destination, so the month-end batch gets a second address. Default quota as of September 2026: 5 per zone.

Cost. Billed per hour and per gigabyte processed, plus data transfer. A gateway in each of two zones doubles the hourly charge: the price of resilience.

Security. It accepts no connection from outside. It cannot have a security group; the resources behind it, and its subnet's network ACL, do the filtering.

Gotchas. A gateway shared by two zones fails with its own zone and takes the other's internet access with it, hence one per zone (step 3). It also ignores traffic arriving over VPC peering.

Read more NAT gateway basics · Regional NAT gateways

Reaching Amazon S3 without the internet#

The ledger service has a different need. It never talks to the internet, only to Amazon S3, where it writes the day's statements every night.

Gateway endpoint
On Azurevirtual network service endpoints

What it does. A gateway endpoint gives private subnets their own route to Amazon S3 or DynamoDB in the same Region, with no NAT gateway in the path.

How it works. It adds a route whose destination is S3's prefix list, the address ranges S3 uses in eu-west-2. Being more specific than 0.0.0.0/0, that route wins (step 4).

When to use it. Any VPC whose private subnets use S3 or DynamoDB. It is free, so the team adds one to every VPC.

When not to. Callers outside the VPC: the office network, peered VPCs, VPN and transit gateway traffic. They need an interface endpoint, which chapter 13 covers.

Limits. Its own Region only; a bucket elsewhere is still reached through the NAT gateway. Default quota as of September 2026: 20 per Region.

Cost. Nothing. Moving the nightly export off the NAT gateway removes the per-gigabyte charge on every statement.

Security. An endpoint policy says what may pass; the default allows everything, so the team limits it to MegaCorp's buckets. Bucket policies can insist on the endpoint with aws:sourceVpce.

Gotchas. Adding it drops open connections to S3, so the team adds it before go-live. Requests now come from private addresses, so aws:SourceIp conditions stop matching; they need aws:VpcSourceIp.

Read more Gateway endpoints · Gateway endpoints for Amazon S3

Security groups and network ACLs#

Two firewalls filter traffic inside a VPC. Security groups are the primary control; network ACLs add a coarse, stateless layer across a whole subnet.

CharacteristicSecurity groupNetwork ACL
Attached toa resource, such as an instancea subnet
Rulesallow rules onlyallow and deny rules
Evaluationevery rule, then one decisionin ascending order, until one matches
Return trafficallowed automatically: statefulallowed only by a rule: stateless
Rules can nameIP ranges, prefix lists, other security groupsIP ranges
Security group
On Azurenetwork security groups, application security groups

What it does. A firewall on each resource's network interface. The load balancer, the application instances and the database each get their own group.

How it works. Rules only allow. The application's group admits port 8080 from the load balancer's group rather than from addresses, so new instances are admitted with no rule change. Replies return automatically: the group is stateful.

When to use it. Always. It is the main control between tiers.

When not to. To deny something, such as a hostile address range. Security groups cannot deny; a network ACL on the subnet can, so the team uses one for that, sparingly.

Limits. By default as of September 2026: 60 inbound and 60 outbound rules per group, and 5 groups per network interface, adjustable to 16. Rules times groups may not exceed 1,000.

Cost. No charge.

Security. Only a few IAM principals may change groups, and no group opens port 22 or 3389 to the internet. Traffic to Amazon DNS and to instance metadata is never filtered.

Gotchas. A new group allows all outbound traffic until the team narrows it. A group belongs to one VPC, unless it is associated with other VPCs in the same Region.

Read more Security groups · Compare security groups and network ACLs · Azure network security groups

The same subnets in Terraform#

One Terraform comparison shows the zone rule in code. Every aws_subnet names its Availability Zone, so resilient AWS subnets come in sets; an azurerm_subnet has no zone to name.

Azure · hashicorp/azurerm 5.5.0azure/network/main.tfsubnets
resource "azurerm_subnet" "gateway" {
  name                 = "snet-gateway"
  resource_group_name  = azurerm_resource_group.payments.name
  virtual_network_name = azurerm_virtual_network.payments.name
  address_prefixes     = ["10.20.0.0/24"]
}

resource "azurerm_subnet" "app" {
  name                            = "snet-app"
  resource_group_name             = azurerm_resource_group.payments.name
  virtual_network_name            = azurerm_virtual_network.payments.name
  address_prefixes                = ["10.20.10.0/24"]
  default_outbound_access_enabled = false
}
AWS · hashicorp/aws 6.64.0aws/network/main.tfsubnets
data "aws_availability_zones" "available" {
  state = "available"
}

# 10.20.0.0/24 and 10.20.1.0/24: one public subnet in each of two zones
resource "aws_subnet" "public" {
  count             = 2
  vpc_id            = aws_vpc.payments.id
  cidr_block        = cidrsubnet(aws_vpc.payments.cidr_block, 8, count.index)
  availability_zone = data.aws_availability_zones.available.names[count.index]

  tags = { Name = "payments-public-${count.index}" }
}

# 10.20.10.0/24 and 10.20.11.0/24: one private subnet in each of the same zones
resource "aws_subnet" "private" {
  count             = 2
  vpc_id            = aws_vpc.payments.id
  cidr_block        = cidrsubnet(aws_vpc.payments.cidr_block, 8, count.index + 10)
  availability_zone = data.aws_availability_zones.available.names[count.index]

  tags = { Name = "payments-private-${count.index}" }
}
False friend: private subnet
On Azure

A subnet with default outbound access turned off. The azurerm provider still defaults default_outbound_access_enabled to true, so the specimen sets it.

On AWS

A subnet with no route to an internet gateway.

Read more Terraform: aws_subnet · Terraform: azurerm_subnet · Default outbound access in Azure

How the VPC got here#

Amazon VPC began in 2009 as a VPN bridge to company networks. The flat network that came before it, EC2-Classic, shared with other customers, is now retired.

How the VPC got here, 2009 to 2025Amazon VPC,reached over aVPN2009VPCs acrosszones, severalper account2011NAT gateways2015EC2-Classicretirementbegins2021RetiredRegional NATgateways2025
Figure How the VPC got here, 2009 to 2025#
As text
  1. 2009: Amazon VPC, reached over a VPN
  2. 2011: VPCs across zones, several per account
  3. 2015: NAT gateways
  4. 2021: EC2-Classic retirement begins (retired)
  5. 2025: Regional NAT gateways

Read more Amazon VPC User Guide: document history · EC2-Classic is retiring: here's how to prepare

MegaCorp's network#

The figure adds the payments network to MegaCorp's running design.

MegaCorp's running design: the accounts and network layers, with this chapter's network layer highlightedAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24InternetgatewayNAT gatewayNAT gateway1
1The NAT gateway sends traffic out through the internet gateway
Figure MegaCorp's running design: the accounts and network layers, with this chapter's network layer highlighted#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
  1. The NAT gateway sends traffic out through the internet gateway

Read more What is IPAM?

Connecting networks#

Part II · Foundations · Chapter 13·8 min read

MegaCorp's VPCs multiply with its accounts. Each must reach a few others, the data centre and shared services, but never everything.

NeedUseWhy
One service that many VPCs callAWS PrivateLinkconsumers reach the service, never each other's networks
Two VPCs that must talkVPC peeringone to one, with nothing in between
Teams in one trust boundary sharing a networka shared VPC, through AWS RAMone network, one owner of routes
Many VPCs, or the data centrea transit gatewayone attachment each, segmented by route tables

Two VPCs: peering#

The settlement job in payments-prod must reach the ledger database in another account. A peering connection joins the two VPCs once both owners agree and add routes. Ranges must not overlap, and peering is not transitive: ledger cannot reach a third VPC through it, nor use payments-prod's NAT gateway or VPN.

False friend: peering
On Azure

Peering can make a hub's VPN or ExpressRoute gateway a spoke's way to the data centre, a setting called gateway transit, and user-defined routes can send a spoke's traffic through an appliance in the hub.

On AWS

Peering joins exactly two VPCs. Neither can use the other's VPN or NAT gateway, or reach a third network through it; a hub on AWS is a transit gateway.

Read more How VPC peering connections work · Azure virtual network peering

Many VPCs: a transit gateway#

With fourteen VPCs, every pair would need 91 peering connections. The network team builds one transit gateway in the network account, shares it through AWS RAM, and each VPC attaches once.

One hub instead of a mesh: the network account shares a transit gateway through AWS RAM, each VPC attaches to it, a Site-to-Site VPN reaches the data centre, and route tables keep development away from productionWorkloads OUAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16AWS Account payments-devRegion eu-west-2VPC 10.30.0.0/16AWS Account networkRegion eu-west-2Customergateway(data centre)AWSSite-to-SiteVPNTransit gatewayattachment(attachment)Transit gatewayattachment(attachment)AWS TransitGateway(hub)AWS ResourceAccess Manager(share)12345
1The network account shares the hub with the Workloads OU through AWS RAM
2payments-prod sends traffic for other networks to the hub
3payments-dev attaches too, but its route table has no route to production
4Traffic for the data centre leaves the hub through the VPN
5Two IPsec tunnels cross the internet to the data centre's router
Figure One hub instead of a mesh: the network account shares a transit gateway through AWS RAM, each VPC attaches to it, a Site-to-Site VPN reaches the data centre, and route tables keep development away from production#
As text
  • Customer gateway (data centre)
  • AWS Site-to-Site VPN
  • Workloads OU
    • AWS Account payments-prod
      • Region eu-west-2
        • VPC 10.20.0.0/16
          • Transit gateway attachment (attachment)
    • AWS Account payments-dev
      • Region eu-west-2
        • VPC 10.30.0.0/16
          • Transit gateway attachment (attachment)
  • AWS Account network
    • Region eu-west-2
      • AWS Transit Gateway (hub)
      • AWS Resource Access Manager (share)
  1. The network account shares the hub with the Workloads OU through AWS RAM
  2. payments-prod sends traffic for other networks to the hub
  3. payments-dev attaches too, but its route table has no route to production
  4. Traffic for the data centre leaves the hub through the VPN
  5. Two IPsec tunnels cross the internet to the data centre's router
AWS Transit Gateway
On AzureAzure Virtual WAN, Azure virtual network peering

What it does. AWS Transit Gateway is a Regional router that VPCs, VPNs and Direct Connect attach to: one connection per network, not one per peer.

How it works. Each VPC attaches through one subnet per zone, and route tables decide who reaches whom.

When to use it. More than a handful of VPCs, or one way in for the data centre.

When not to. Two VPCs on their own: peering is simpler.

Limits. Defaults as of September 2026: 5 per account per Region and 5,000 attachments each, both adjustable; up to 100 Gbps per VPC attachment per zone.

Cost. Billed per attachment per hour, and per gigabyte sent into it, charged to the sender.

Security. Segment with route tables and blackhole routes, and share one hub through AWS RAM.

Gotchas. Only zones with an attachment subnet reach the hub. A stateful appliance behind it needs appliance mode, or replies are dropped.

Read more AWS Transit Gateway documentation · How AWS Transit Gateway works · AWS Transit Gateway quotas · AWS Transit Gateway pricing

One VPC, many accounts#

For a dozen small tools, the network team shares one VPC's subnets with their accounts through AWS RAM, and alone changes the routes.

AWS Resource Access Manager
On Azureno direct equivalent

What it does. AWS RAM shares a resource one account owns with other accounts or OUs: subnets, transit gateways, DNS rules and more.

How it works. The owner creates a resource share. Inside the organization no invitation is needed, and the resource appears in the other account as if it were its own.

When to use it. One network resource serving many accounts.

When not to. When the other account must fully control the resource: the owner keeps it.

Limits. Regional resources are shared within their Region, and global ones only from us-east-1.

Cost. No additional charge.

Security. The share's permission caps what other accounts can do, and their own policies and SCPs still apply.

Gotchas. Participants in a shared VPC cannot create NAT gateways, change routes or use the owner's default security group.

Read more AWS RAM documentation · Share your VPC subnets with other accounts · Responsibilities and permissions for owners and participants

Offering one service privately#

Twenty VPCs call the fraud team's scoring API. Instead of routing them all into its VPC, fraud publishes an endpoint service, and each consumer adds an interface endpoint.

PrivateLink: payments-prod reaches fraud's scoring API through an interface endpoint in its own VPC, and the networks never route to each otherAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16AWS Account fraudRegion eu-west-2VPC 10.50.0.0/16VPC endpoints(interfaceendpoint)AWS PrivateLink(scoring API)1
1Requests go one way, from the endpoint to the service, on the AWS network
Figure PrivateLink: payments-prod reaches fraud's scoring API through an interface endpoint in its own VPC, and the networks never route to each other#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • VPC endpoints (interface endpoint)
  • AWS Account fraud
    • Region eu-west-2
      • VPC 10.50.0.0/16
        • AWS PrivateLink (scoring API)
  1. Requests go one way, from the endpoint to the service, on the AWS network

Reaching the data centre#

Settlement files still go to a mainframe in MegaCorp's data centre. Site-to-Site VPN links the hub to it first; Direct Connect follows when volumes grow, with the VPN as backup.

AWS Site-to-Site VPN
On AzureAzure VPN Gateway

What it does. AWS Site-to-Site VPN joins a VPC or transit gateway to an on-premises network through IPsec tunnels over the internet.

How it works. Each connection has two tunnels, for high availability, and ends on a virtual private gateway or a transit gateway.

When to use it. A first link, small sites, and the backup for Direct Connect.

When not to. Large, steady volumes or strict latency: internet paths vary.

Limits. as of September 2026: 1.25 Gbps per tunnel, or up to 5 Gbps with large bandwidth tunnels on a transit gateway.

Cost. Billed per connection per hour, plus data transfer out.

Security. IPsec encrypts traffic between AWS and the customer's router.

Gotchas. Overlapping address ranges between the VPCs and the data centre cannot be routed apart, so keep them distinct.

Read more AWS Site-to-Site VPN documentation · What is AWS Site-to-Site VPN?

AWS Direct Connect
On AzureAzure ExpressRoute

What it does. AWS Direct Connect is a private fibre link from MegaCorp's network to AWS at a Direct Connect location.

How it works. A dedicated connection is a port for one customer; a hosted one comes through a partner. Virtual interfaces over it reach VPCs, transit gateways or public AWS services.

When to use it. Large, steady volumes, or latency that must be predictable.

When not to. As the only link for anything critical.

Limits. The 99.99% model needs separate devices in more than one location.

Cost. Billed per port hour, plus data transfer out, charged to the account that sends it.

Security. Not encrypted by default: use MACsec where supported, or VPN over it.

Gotchas. A private virtual interface reaches one VPC; to reach many, use a Direct Connect gateway.

Read more AWS Direct Connect documentation · What is Direct Connect? · AWS Direct Connect Resiliency Toolkit · Encryption in AWS Direct Connect

Inspecting traffic#

Payment data brings auditors, who want traffic from the data centre inspected and outbound traffic limited to known domains. Security groups cannot filter by domain, so MegaCorp adds AWS Network Firewall to the hub.

AWS Network Firewall
On AzureAzure Firewall

What it does. AWS Network Firewall is a managed, stateful firewall with intrusion prevention and domain allow-lists for VPC traffic.

How it works. An endpoint in a dedicated subnet per zone inspects what route tables send it; or the firewall attaches straight to a transit gateway.

When to use it. Central inspection between networks, to the internet, or from on premises.

When not to. As a replacement for security groups: it adds to them.

Limits. An endpoint cannot filter its own subnet, so firewall subnets hold nothing else.

Cost. Billed per endpoint per zone per hour and per gigabyte; a NAT gateway on the same path has its charges waived.

Security. Manage firewalls across accounts with AWS Firewall Manager.

Gotchas. Behind a transit gateway, appliance mode keeps both directions of a flow on one endpoint.

Read more AWS Network Firewall documentation · What is AWS Network Firewall? · AWS Network Firewall pricing

MegaCorp's design gains its connectivity layer.

MegaCorp's design with its connectivity layer: the VPC's attachment to the shared transit gateway, and the route to the data centreAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24Customergateway(data centre)AWS TransitGateway(shared hub)InternetgatewayTransit gatewayattachment(attachment)NAT gatewayNAT gateway123
1The NAT gateway sends traffic out through the internet gateway
2Traffic for MegaCorp's other networks leaves through the attachment to the shared hub
3The hub, owned by the network account, reaches the data centre over Site-to-Site VPN
Figure MegaCorp's design with its connectivity layer: the VPC's attachment to the shared transit gateway, and the route to the data centre#
As text
  • Customer gateway (data centre)
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS Transit Gateway (shared hub)
      • VPC 10.20.0.0/16
        • Internet gateway
        • Transit gateway attachment (attachment)
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
  1. The NAT gateway sends traffic out through the internet gateway
  2. Traffic for MegaCorp's other networks leaves through the attachment to the shared hub
  3. The hub, owned by the network account, reaches the data centre over Site-to-Site VPN

DNS#

Part II · Foundations · Chapter 14·3 min read

Names on AWS come from Amazon Route 53: public names for the internet, private names inside VPCs, and a resolver in every VPC that forwards to and from the data centre.

Public and private names#

The merchant portal needs portal.megacorp.com on the internet; the settlement job needs ledger.internal.megacorp.com inside MegaCorp's VPCs only. The first lives in a public hosted zone, the second in a private hosted zone associated with those VPCs.

Amazon Route 53
On AzureAzure DNS, Azure Traffic Manager

What it does. Amazon Route 53 is AWS's DNS: public and private zones, routing policies, and health checks.

How it works. A hosted zone holds a domain's records; a private zone answers only in its associated VPCs. Alias records point a name, even the apex, at an AWS resource and follow its address changes.

When to use it. Every domain on AWS, and routing by weight, latency, location or health, such as failover between Regions.

When not to. Moving a corporate domain that another provider must keep: delegate a subdomain instead.

Limits. Defaults as of September 2026: 500 hosted zones per account and 10,000 records per zone, both adjustable.

Cost. Billed per hosted zone per month and per query; alias queries to AWS resources are free.

Security. VPC Resolver can validate DNSSEC, and DNS Firewall blocks listed domains.

Gotchas. A private zone needs DNS hostnames and DNS support turned on in each VPC.

Read more Amazon Route 53 documentation · Considerations when working with a private hosted zone · Choosing a routing policy · Route 53 quotas · Amazon Route 53 pricing

How VPC Resolver answers a query from inside a VPC: a forwarding rule first, then a matching private zone, then the internetyesnoyesnoDoes a forwarding rule match the name?Forward the query, such as to thedata centreDoes a private hosted zone associatedwith the VPC match?Answer from that zone, or "no suchdomain" if it has no recordResolve the name on the internet
Figure How VPC Resolver answers a query from inside a VPC: a forwarding rule first, then a matching private zone, then the internet#
As text
  1. Does a forwarding rule match the name? Yes: Forward the query, such as to the data centre. No: the next step.
  2. Does a private hosted zone associated with the VPC match? Yes: Answer from that zone, or "no such domain" if it has no record. No: the next step.
  3. Resolve the name on the internet
Gotcha

A private zone hides the public names it overlaps. Had the team called the private zone megacorp.com, VPC Resolver would answer every megacorp.com query inside the VPCs from it. A name only the public zone holds, such as portal.megacorp.com, would get "no such domain". A separate internal subdomain avoids the clash.

Names across the data centre#

The mainframe's name lives in the data centre's DNS, which must also resolve internal.megacorp.com. The network team adds an outbound endpoint with a forwarding rule for corp.megacorp.com, and an inbound endpoint the data centre forwards to, and shares the rule through AWS RAM.

A query for a data centre name: VPC Resolver matches the forwarding rule and sends it through the outbound endpoint to the data centre's DNSSettlement jobVPC ResolverOutbound endpointData centre DNSmainframe.corp.megacorp.com?1the rule for corp.megacorp.com matches2forward the query3over the VPN or DirectConnect4the mainframe's address5
Figure A query for a data centre name: VPC Resolver matches the forwarding rule and sends it through the outbound endpoint to the data centre's DNS#
As text
  1. Settlement job to VPC Resolver: mainframe.corp.megacorp.com?
  2. VPC Resolver: the rule for corp.megacorp.com matches
  3. VPC Resolver to Outbound endpoint: forward the query
  4. Outbound endpoint to Data centre DNS: over the VPN or Direct Connect
  5. Data centre DNS to Settlement job: the mainframe's address

Read more What is Route 53 VPC Resolver? · Choosing between alias and non-alias records

MegaCorp's design gains its DNS layer.

MegaCorp's running design with the DNS layer this chapter adds: VPC Resolver answers internal names from the private hosted zoneAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24Route 53 hostedzone(internal.megacorp.com)InternetgatewayRoute 53 VPCResolver(VPC Resolver)NAT gatewayNAT gateway12
1The NAT gateway sends traffic out through the internet gateway
2VPC Resolver answers internal.megacorp.com from the private hosted zone
Figure MegaCorp's running design with the DNS layer this chapter adds: VPC Resolver answers internal names from the private hosted zone#
As text
  • AWS Account payments-prod
    • Route 53 hosted zone (internal.megacorp.com)
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Route 53 VPC Resolver (VPC Resolver)
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
  1. The NAT gateway sends traffic out through the internet gateway
  2. VPC Resolver answers internal.megacorp.com from the private hosted zone

Landing zones#

Part II · Foundations · Chapter 15·4 min read

MegaCorp's accounts were set up one at a time. A landing zone gives every account the same logging, access and controls from birth, and keeps them there.

From accounts to a landing zone#

By its second year MegaCorp has thirty hand-built accounts, and nobody can promise the auditors that every one logs every API call where its own administrators cannot change it. The platform team adopts AWS Control Tower.

MegaCorp's landing zone: Control Tower, the log-archive and audit accounts in the Security OU, and controls on every OUSecurity OUAWS Account auditAWS Account log-archiveWorkloads OUAWS Account payments-prodAWS ControlTowerAmazon SNS(drift alerts)Amazon S3(everyaccount's logs)AWS CloudTrail(trail)123
1The Workloads OU's controls govern payments-prod from the day it is made
2payments-prod's API calls are delivered to the log-archive account
3Drift and compliance alerts gather in the audit account
Figure MegaCorp's landing zone: Control Tower, the log-archive and audit accounts in the Security OU, and controls on every OU#
As text
  • AWS Control Tower
  • Security OU
    • AWS Account audit
      • Amazon SNS (drift alerts)
    • AWS Account log-archive
      • Amazon S3 (every account's logs)
  • Workloads OU
    • AWS Account payments-prod
      • AWS CloudTrail (trail)
  1. The Workloads OU's controls govern payments-prod from the day it is made
  2. payments-prod's API calls are delivered to the log-archive account
  3. Drift and compliance alerts gather in the audit account
False friend: landing zone
On Azure

The Cloud Adoption Framework's Azure landing zone is a platform landing zone plus one application landing zone per workload: that workload's development, test and production subscriptions.

On AWS

The landing zone is the whole multi-account environment, which Control Tower sets up and governs. A workload gets accounts in it, not a landing zone of its own.

Read more What is AWS Control Tower? · What is an Azure landing zone?

AWS Control Tower
On AzureAzure landing zone

What it does. AWS Control Tower sets up and governs a landing zone: shared log-archive and audit accounts, controls on every OU, and a factory for new accounts.

How it works. It orchestrates Organizations, IAM Identity Center, Service Catalog, CloudTrail and Config. Controls are preventive (SCPs and RCPs), detective (Config rules) or proactive (CloudFormation hooks).

When to use it. Starting a multi-account estate, or bringing an existing organization under one baseline.

When not to. When another tool edits the same SCPs: Control Tower treats that as drift.

Limits. as of September 2026: up to 10,000 accounts, ten SCPs per OU, and a home Region that cannot be changed.

Cost. No additional charge; AWS Config, CloudTrail, Amazon S3 and the other services it uses are billed.

Security. Preventive controls do not bind the management account, so keep it nearly empty.

Gotchas. Moving an account between OUs or editing a managed SCP outside Control Tower is drift, and enrolling accounts stops until it is repaired.

Read more AWS Control Tower documentation · Control behavior and guidance · Detect and resolve drift in AWS Control Tower · How AWS Regions work with AWS Control Tower · Limitations and quotas in AWS Control Tower · AWS Control Tower pricing

Vending accounts#

New accounts come from Account Factory: a staging account requested in the Workloads OU arrives with the baseline and the OU's controls in force. Terraform teams use Account Factory for Terraform instead.

Vending an account: created in the chosen OU, governed at once, and given the baseline before the team gets itPayments teamAccount FactoryAWS Organizationspayments-staginga new account in theWorkloads OU1create the account in thatOU2the OU's controls apply at once3the landing zone's baseline4ready5
Figure Vending an account: created in the chosen OU, governed at once, and given the baseline before the team gets it#
As text
  1. Payments team to Account Factory: a new account in the Workloads OU
  2. Account Factory to AWS Organizations: create the account in that OU
  3. AWS Organizations: the OU's controls apply at once
  4. Account Factory to payments-staging: the landing zone's baseline
  5. payments-staging to Payments team: ready

Read more Provision and manage accounts with Account Factory · Overview of Account Factory for Terraform (AFT)

MegaCorp's design gains its governance layer.

MegaCorp's design with its governance layer: the organization trail delivers payments-prod's API calls to log-archiveAWS Account payments-prodRegion eu-west-2AWS Account log-archiveAWS CloudTrail(organizationtrail)Amazon S3(trail logs)1
1Every API call in payments-prod is delivered to the log-archive account
Figure MegaCorp's design with its governance layer: the organization trail delivers payments-prod's API calls to log-archive#
As text
  • AWS Account payments-prod
    • AWS CloudTrail (organization trail)
    • Region eu-west-2
  • AWS Account log-archive
    • Amazon S3 (trail logs)
  1. Every API call in payments-prod is delivered to the log-archive account
Foundations
  1. A developer's new role can do nothing, although its permissions boundary allows everything. Why?
    Answer
    A boundary only caps. The role also needs a permissions policy, and gets what both allow.
  2. Two accounts both put their services in eu-west-2a to keep traffic inside one zone. Can the bill still show traffic between zones?
    Answer
    Yes: zone names map differently in each account. Compare zone IDs, such as euw2-az1.
  3. Fourteen VPCs and the data centre must reach each other, but development must never reach production. What do you build?
    Answer
    A transit gateway shared through AWS RAM, with separate route tables for production and development.

Choosing compute#

Part III · Building blocks · Chapter 16·3 min read

AWS runs code on servers you manage, in containers it schedules, or as functions it runs per event. The less you manage, the less you control: choose the most managed option each workload can live with.

Three ways to run code#

MegaCorp's payments team has three workloads. The payments API is a Java service that runs all day, the nightly settlement job takes about 40 minutes, and a handler takes bursts of small webhook calls from the card schemes. Each could run in any of three models, which differ in who does the work of running it.

ModelYou manageAWS managesServices
Instancesthe operating system, patching, scaling and security of each serverthe hardwareAmazon EC2 with EC2 Auto Scaling
Containersthe image and the size of each taskthe servers, with AWS Fargate, and the schedulingAmazon ECS or Amazon EKS, on Fargate or EC2
Functionsthe code and its memoryservers, operating system, capacity, scaling and loggingAWS Lambda

Read more Choosing an AWS compute service

Deciding, workload by workload#

The team asks the same questions of each workload, in order, and stops at the first yes.

Choosing compute for one workload: the first question answered yes decides, and most services end at containers on FargateyesnoyesnoyesnoDoes it need a particular operatingsystem, GPU, licence or agent on thehost?Amazon EC2, in an AutoScaling groupIs it short work triggered by events,each piece done within 15 minutes andkeeping no state?AWS LambdaDoes the organization already runKubernetes well?Amazon EKSAmazon ECS on AWS Fargate
Figure Choosing compute for one workload: the first question answered yes decides, and most services end at containers on Fargate#
As text
  1. Does it need a particular operating system, GPU, licence or agent on the host? Yes: Amazon EC2, in an Auto Scaling group. No: the next step.
  2. Is it short work triggered by events, each piece done within 15 minutes and keeping no state? Yes: AWS Lambda. No: the next step.
  3. Does the organization already run Kubernetes well? Yes: Amazon EKS. No: the next step.
  4. Amazon ECS on AWS Fargate

The webhook handler is short, event-driven work, so it goes to Lambda. The API runs all day and the settlement job runs for 40 minutes, too long for one Lambda invocation as of September 2026, so both go to ECS on Fargate. Only the fraud team's scoring model, which needs GPUs, lands on EC2.

MegaCorp's compute choices: the API and the settlement job as containers on Fargate, the webhooks on Lambda, and the GPU model on EC2AWS Account payments-prodRegion eu-west-2AWS Account fraudRegion eu-west-2AWS Fargate(payments API)AWS Fargate(settlementjob)AWS Lambda(webhooks)Amazon EC2(GPU scoring)
Figure MegaCorp's compute choices: the API and the settlement job as containers on Fargate, the webhooks on Lambda, and the GPU model on EC2#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS Fargate (payments API)
      • AWS Fargate (settlement job)
      • AWS Lambda (webhooks)
  • AWS Account fraud
    • Region eu-west-2
      • Amazon EC2 (GPU scoring)
Gotcha

Kubernetes is a commitment, not a default. It releases three versions a year and retires old ones, so every cluster needs regular upgrades. Choose Amazon EKS when the skills and the need are already there.

Read more Choosing an AWS container service · Lambda quotas

How AWS got here#

Each step in AWS's compute history hid more of the server. Simpler platforms remain: AWS Elastic Beanstalk runs your code on EC2 that it manages, and Amazon Lightsail bundles servers, databases and networking at a predictable monthly price. AWS App Runner, the simplest path for web apps, is closed to new customers, and AWS points new applications to Amazon ECS Express Mode instead.

Running code on AWS, each step hiding more of the server: from EC2 to Elastic Beanstalk, Lambda, ECS and FargateAmazon EC2: aserver you run2006AWS ElasticBeanstalk: yourcode on managed2011AWS Lambda:code, no servers2014Amazon ECS:containers on amanaged cluster2015AWS Fargate:containers, noservers2017
Figure Running code on AWS, each step hiding more of the server: from EC2 to Elastic Beanstalk, Lambda, ECS and Fargate#
As text
  1. 2006: Amazon EC2: a server you run
  2. 2011: AWS Elastic Beanstalk: your code on managed EC2
  3. 2014: AWS Lambda: code, no servers
  4. 2015: Amazon ECS: containers on a managed cluster
  5. 2017: AWS Fargate: containers, no servers

Read more AWS App Runner availability change · What is Amazon Lightsail?

EC2 and Auto Scaling#

Part III · Building blocks · Chapter 17·5 min read

Amazon EC2 rents virtual servers by the second, and EC2 Auto Scaling keeps a healthy number of them running. Use them when a workload needs the server itself, and put even one inside a group.

One instance#

The fraud team's scoring model needs GPUs and a driver installed on the host, so it runs on EC2. An instance boots from an Amazon Machine Image, which holds its operating system and software, and its instance type fixes the CPU, memory and accelerators. In c7gn.xlarge, c means compute optimized, 7 the generation, g Graviton, n extra networking, and xlarge the size.

Amazon EC2
On AzureAzure Virtual Machines

What it does. Amazon EC2 rents virtual servers, called instances, by the second. The fraud team's model runs on GPU instances with the drivers it needs.

How it works. An instance boots from an AMI into a subnet, with an instance type that sets its hardware and an IAM role for its credentials. Everything above the hypervisor is yours to run.

When to use it. A particular operating system, GPU, licence or host agent, or software that expects a whole server.

When not to. Work that a container on Fargate or a Lambda function can do: every instance is yours to patch, scale and secure.

Limits. as of September 2026: On-Demand and Spot capacity is capped per Region in vCPUs, by instance family; request increases well before a launch.

Cost. On-Demand is billed per second. Savings Plans lower the rate for a one- or three-year hourly commitment, and Spot sells spare capacity at a discount. Volumes and data transfer are billed apart.

Security. Require IMDSv2, which is optional by default, and reach instances through Session Manager: no inbound ports, bastion hosts or SSH keys.

Gotchas. An AMI belongs to one Region, so copy it before launching in another. A Spot Instance gets two minutes' notice before it is taken back.

Read more Amazon EC2 documentation · Amazon EC2 instance type naming conventions · Amazon Machine Images in Amazon EC2 · Amazon EC2 service quotas

Gotcha

The metadata service hands out the role's credentials. Applications on an instance get its role's temporary credentials from instance metadata. IMDSv2 guards that door with session tokens, AWS's defence against server-side request forgery bugs, but by default an instance also accepts IMDSv1. Require IMDSv2 on every instance.

A fleet: Auto Scaling groups#

One instance is one failure away from an outage, and scoring load triples on sale days. The team runs the instances in an Auto Scaling group across two zones, behind a load balancer: at least two, at most eight, and replaced when a health check fails.

The scoring fleet: an Auto Scaling group keeps instances in two zones, and the load balancer sends requests only to healthy onesAWS Account fraudRegion eu-west-2VPC 10.60.0.0/16Availability Zone aPrivate subnet 10.60.10.0/24Availability Zone bPrivate subnet 10.60.11.0/24ApplicationLoad Balancer(scoring)Amazon EC2 AutoScaling(scoring group)Amazon EC2(GPU instance)Amazon EC2(GPU instance)12
1The load balancer sends requests to healthy instances in both zones
2The group replaces an instance that fails its health check
Figure The scoring fleet: an Auto Scaling group keeps instances in two zones, and the load balancer sends requests only to healthy ones#
As text
  • AWS Account fraud
    • Region eu-west-2
      • VPC 10.60.0.0/16
        • Application Load Balancer (scoring)
        • Amazon EC2 Auto Scaling (scoring group)
        • Availability Zone a
          • Private subnet 10.60.10.0/24
            • Amazon EC2 (GPU instance)
        • Availability Zone b
          • Private subnet 10.60.11.0/24
            • Amazon EC2 (GPU instance)
  1. The load balancer sends requests to healthy instances in both zones
  2. The group replaces an instance that fails its health check
Amazon EC2 Auto Scaling
On AzureAzure Virtual Machine Scale Sets

What it does. EC2 Auto Scaling keeps the right number of healthy instances running: the scoring group never drops below two, and grows to eight under load.

How it works. A group launches instances from a launch template across chosen zones, balancing them evenly, registers them with a load balancer, and replaces any that fail a health check. Policies and schedules change the desired count.

When to use it. Every EC2 workload, even a single instance: a group of one still replaces it when it fails.

When not to. Containers on Fargate or functions on Lambda, where AWS runs the capacity.

Limits. as of September 2026: 500 groups per Region, and 50 scaling policies and 125 scheduled actions per group.

Cost. No additional fees; you pay for the instances, volumes and alarms it uses.

Security. Put the IAM role and the IMDSv2 requirement in the launch template, so every new instance starts with them.

Gotchas. Instances come and go, so keep nothing on them that must survive: logs, sessions and files belong in services. A lifecycle hook gives a terminating instance time to finish.

Read more Amazon EC2 Auto Scaling documentation · What is Amazon EC2 Auto Scaling? · Quotas for Auto Scaling resources and groups

Paying for instances#

The scoring fleet runs all year, while the nightly retraining job can stop and restart. The team pays for each differently.

Choosing how to pay for instances: interruptible work on Spot, steady use under a Savings Plan, and guaranteed capacity by reservationyesnoyesnoyesnoCan the work stop at two minutes'notice and pick up again?Spot InstancesWill this much compute run steadily fora year or more?A Savings Plan for the steady partMust capacity be certain in one zone,such as for a failover?A Capacity ReservationOn-Demand, by the second
Figure Choosing how to pay for instances: interruptible work on Spot, steady use under a Savings Plan, and guaranteed capacity by reservation#
As text
  1. Can the work stop at two minutes' notice and pick up again? Yes: Spot Instances. No: the next step.
  2. Will this much compute run steadily for a year or more? Yes: A Savings Plan for the steady part. No: the next step.
  3. Must capacity be certain in one zone, such as for a failover? Yes: A Capacity Reservation. No: the next step.
  4. On-Demand, by the second

Retraining goes to Spot. A Savings Plan covers the fleet's two steady instances, and the sale-day bursts run On-Demand.

Read more Amazon EC2 billing and purchasing options · Spot Instance interruption notices · Use the Instance Metadata Service to access instance metadata · AWS Systems Manager Session Manager

How EC2 grew#

EC2 over the years: rented servers, then fleets that follow demand, spare capacity, AWS's own processors and commitments to spendAmazon EC2: aserver you rent2006Auto Scaling:fleets thatfollow demand2009Spot Instances:spare capacity,cheaper2009AWS Graviton:Arm-basedinstances2018Savings Plans:commit to spend,not to a type2019
Figure EC2 over the years: rented servers, then fleets that follow demand, spare capacity, AWS's own processors and commitments to spend#
As text
  1. 2006: Amazon EC2: a server you rent
  2. 2009: Auto Scaling: fleets that follow demand
  3. 2009: Spot Instances: spare capacity, cheaper
  4. 2018: AWS Graviton: Arm-based instances
  5. 2019: Savings Plans: commit to spend, not to a type

Read more Amazon EC2 documentation

Containers#

Part III · Building blocks · Chapter 18·6 min read

A container image is built once and runs anywhere. On AWS, Amazon ECR stores the images, Amazon ECS or Amazon EKS runs them, and AWS Fargate supplies the capacity.

Images, and where they live#

The pipeline builds the payments API into a container image on every merge and pushes it to a private Amazon ECR repository; the same image then runs in development and in production.

Amazon ECR
On AzureAzure Container Registry

What it does. Amazon ECR stores the payments team's images privately, beside the services that run them.

How it works. Repositories hold images and a resource policy, and pushes and pulls use IAM. ECR can scan on push, replicate across Regions and accounts, cache public images and expire old ones.

When to use it. Every image that ECS, EKS or Lambda runs.

When not to. Pulling public images from the internet in production: cache them in ECR.

Limits. Repositories are Regional, so replicate images to each Region that runs them.

Cost. Storage, data transfer for pushes and pulls, and opted-in actions such as replication.

Security. Scan on push, make tags immutable, and grant other accounts pull access by policy.

Gotchas. Private tasks pull images through the NAT gateway, and pay per gigabyte, unless the VPC has ECR endpoints.

Read more Amazon ECR documentation · What is Amazon Elastic Container Registry?

Built once, run everywhere: the pipeline pushes each image to ECR in the tooling account, and development and production pull the same imageAWS Account toolingRegion eu-west-2Workloads OUAWS Account payments-devRegion eu-west-2AWS Account payments-prodRegion eu-west-2AWSCodePipeline(payments)Amazon ECR(payments-api)AWS Fargate(API tasks)AWS Fargate(API tasks)123
1The pipeline pushes the image once, tagged with its commit
2Development pulls it, as the repository's policy allows
3Production pulls the same image once the tests pass
Figure Built once, run everywhere: the pipeline pushes each image to ECR in the tooling account, and development and production pull the same image#
As text
  • AWS Account tooling
    • Region eu-west-2
      • AWS CodePipeline (payments)
      • Amazon ECR (payments-api)
  • Workloads OU
    • AWS Account payments-dev
      • Region eu-west-2
        • AWS Fargate (API tasks)
    • AWS Account payments-prod
      • Region eu-west-2
        • AWS Fargate (API tasks)
  1. The pipeline pushes the image once, tagged with its commit
  2. Development pulls it, as the repository's policy allows
  3. Production pulls the same image once the tests pass

Running containers with ECS#

The API runs as an ECS service of at least two tasks on Fargate behind a load balancer; the settlement job is a task that runs nightly and stops.

Amazon ECS
On AzureAzure Container Apps

What it does. Amazon ECS runs and scales containers with no control plane to operate.

How it works. A task definition names the containers, their size and two roles: the task role the code calls AWS with, and the execution role ECS uses to pull images and write logs. A service keeps tasks running behind a load balancer.

When to use it. Most container workloads, when nobody needs Kubernetes itself.

When not to. An organization committed to Kubernetes tooling: that is EKS.

Limits. as of September 2026: 5,000 services per cluster and 5,000 tasks per service.

Cost. Only the capacity the tasks use, on Fargate or EC2.

Security. Give each service its own task role, with least privilege.

Gotchas. ECS Express Mode, which builds a web service from one call, puts an internet-facing load balancer in the default VPC unless you give it subnets.

Read more Amazon ECS documentation · What is Amazon Elastic Container Service? · Amazon ECS task IAM role · Create your first Express Mode service using the AWS CLI · Amazon ECS endpoints and quotas

Gotcha

A task has two roles, and they are easily swapped. The task role is what the payments code calls AWS with; the execution role is what ECS uses to pull the image and write logs. Grant S3 to the execution role and the code is still refused, and a failed image pull points to the execution role, not the task role.

AWS Fargate
On AzureAzure Container Apps, Azure Container Instances

What it does. AWS Fargate runs ECS tasks and EKS pods on capacity AWS manages: no instances to patch.

How it works. You set CPU and memory per task. Each task gets its own kernel and network interface, on the latest patched platform.

When to use it. The default capacity for containers, and wherever tasks must be isolated.

When not to. GPUs, privileged containers or tasks above 32 vCPU.

Limits. as of September 2026: 0.25 to 32 vCPU per task; new accounts start with low vCPU quotas.

Cost. Per second for the vCPU, memory and storage requested. Fargate Spot is cheaper but can be interrupted.

Security. No task can reach another task's credentials.

Gotchas. AWS retires tasks on a platform revision with a security issue, so a service must tolerate any task being replaced.

Read more AWS Fargate for Amazon ECS · Architect for AWS Fargate for Amazon ECS · Amazon ECS task definition differences for Fargate · AWS Fargate pricing

ECS or EKS#

The risk team already runs Kubernetes in the data centre, so its model platform moves to Amazon EKS; the payments team, new to containers, stays on ECS.

Choosing how to run containers: EKS where Kubernetes is the standard, otherwise ECSyesnoyesnoDoes the organization already runKubernetes, or need its tools andportability?Amazon EKSDoes it need GPUs or particularinstance types?ECS on ECS ManagedInstancesAmazon ECS on AWS Fargate
Figure Choosing how to run containers: EKS where Kubernetes is the standard, otherwise ECS#
As text
  1. Does the organization already run Kubernetes, or need its tools and portability? Yes: Amazon EKS. No: the next step.
  2. Does it need GPUs or particular instance types? Yes: ECS on ECS Managed Instances. No: the next step.
  3. Amazon ECS on AWS Fargate
Amazon EKS
On AzureAzure Kubernetes Service

What it does. Amazon EKS runs the Kubernetes control plane, so the risk team keeps its tooling without running masters.

How it works. Nodes come from EKS Auto Mode, which AWS also runs, from managed node groups on EC2, or from Fargate. Pods get AWS credentials through EKS Pod Identity.

When to use it. Kubernetes skills, tooling or portability needs already in place.

When not to. A team new to containers: ECS does the job with far less to learn and upgrade.

Limits. as of September 2026: each version has 14 months of standard support and 12 of extended support, then an automatic upgrade.

Cost. Per cluster per hour, higher in extended support, plus the nodes and any Auto Mode fee.

Security. Give each service account its own role through Pod Identity, and restrict instance metadata on nodes.

Gotchas. A control plane upgrade leaves managed and self-managed nodes behind: upgrade them too.

Read more Amazon EKS documentation · What is Amazon EKS? · Understand the Kubernetes version lifecycle on EKS · Amazon EKS pricing · Learn how EKS Pod Identity grants pods access to AWS services

Read more Choosing an AWS container service

How AWS got here#

Containers on AWS, from a scheduler for your own instances to services made from one callAmazon ECS:containers onyour EC2 cluster2015AWS Fargate: nocluster to run2017Amazon EKS:managedKubernetes2018EKS Auto Mode:AWS runs thenodes too2024ECS ExpressMode: a servicefrom one call2025
Figure Containers on AWS, from a scheduler for your own instances to services made from one call#
As text
  1. 2015: Amazon ECS: containers on your EC2 cluster
  2. 2017: AWS Fargate: no cluster to run
  3. 2018: Amazon EKS: managed Kubernetes
  4. 2024: EKS Auto Mode: AWS runs the nodes too
  5. 2025: ECS Express Mode: a service from one call

MegaCorp's design gains its compute layer.

MegaCorp's design with its compute layer: Fargate tasks in both zones' private subnetsAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24InternetgatewayNAT gatewayAWS Fargate(payments)NAT gatewayAWS Fargate(payments)12
1Tasks reach the internet through the NAT gateway in their own zone
2The NAT gateway sends traffic out through the internet gateway
Figure MegaCorp's design with its compute layer: Fargate tasks in both zones' private subnets#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
            • AWS Fargate (payments)
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
            • AWS Fargate (payments)
  1. Tasks reach the internet through the NAT gateway in their own zone
  2. The NAT gateway sends traffic out through the internet gateway

Read more Amazon ECS documentation

Serverless#

Part III · Building blocks · Chapter 19·5 min read

AWS Lambda runs a function for each event and bills by the millisecond; AWS Step Functions joins functions and services into workflows that retry, wait and remember. Both remove servers, not the limits you design around.

Functions that run per event#

Card-scheme webhooks land on an Amazon SQS queue, and a Java function settles each message. Lambda reads the queue, hands the function batches of messages, and runs as many copies as the queue needs, up to the account's concurrency.

java/src/main/java/com/example/payments/SettlementHandler.javahandler
public class SettlementHandler implements RequestHandler<SQSEvent, SQSBatchResponse> {

    private final Settlements settlements;

    public SettlementHandler() {
        this(new Settlements());
    }

    SettlementHandler(Settlements settlements) {
        this.settlements = settlements;
    }

    @Override
    public SQSBatchResponse handleRequest(SQSEvent event, Context context) {
        List<SQSBatchResponse.BatchItemFailure> failures = new ArrayList<>();
        for (SQSEvent.SQSMessage message : event.getRecords()) {
            try {
                settlements.settle(message.getMessageId(), message.getBody());
            } catch (IllegalArgumentException e) {
                failures.add(new SQSBatchResponse.BatchItemFailure(message.getMessageId()));
            }
        }
        return new SQSBatchResponse(failures);
    }
}

The handler reports only the messages that failed, so only those are retried. The team reserves concurrency for it, so no other function can starve it and it cannot flood the ledger database.

AWS Lambda
On AzureAzure Functions

What it does. AWS Lambda runs a function for each event, from a queue, a bucket, an API or a schedule, and bills only while it runs.

How it works. Each concurrent request gets its own execution environment, created on demand and then reused. The first request to a new one waits for initialization: the cold start. Functions call AWS with an execution role.

When to use it. Short, event-driven work: webhooks, queue consumers, file processing, glue between services.

When not to. Work over 15 minutes, or steady heavy load that suits containers or Lambda Managed Instances.

Limits. as of September 2026: 15 minutes per invocation, up to 10,240 MB of memory, 6 MB synchronous payloads, and 1,000 concurrent executions per Region, shared by every function.

Cost. Per request and per GB-second, with a monthly free tier. Provisioned concurrency is billed while configured; SnapStart for Java costs nothing extra.

Security. One execution role per function, with least privilege. A function in a VPC reaches the internet only through a NAT gateway, even in a public subnet.

Gotchas. Events can arrive more than once, and a failed asynchronous event is retried twice: make handlers idempotent, and catch what still fails in a dead-letter queue.

Read more AWS Lambda documentation · Lambda quotas · Understanding Lambda function scaling · AWS Lambda pricing · Giving Lambda functions access to resources in an Amazon VPC · Handling errors for an SQS event source in Lambda

When a statement file lands in Amazon S3, S3 invokes an indexing function asynchronously, and Lambda owns the retries.

An asynchronous invocation that fails: Lambda queues the event, tries twice more, then hands it to a dead-letter queueAmazon S3LambdaIndexing functionDead-letter queuea new file, queued1invoke2error3two more attempts, one minute then twominutes apart4the event that still failed5
Figure An asynchronous invocation that fails: Lambda queues the event, tries twice more, then hands it to a dead-letter queue#
As text
  1. Amazon S3 to Lambda: a new file, queued
  2. Lambda to Indexing function: invoke
  3. Indexing function to Lambda: error
  4. Lambda: two more attempts, one minute then two minutes apart
  5. Lambda to Dead-letter queue: the event that still failed

Read more How Lambda handles errors and retries with asynchronous invocation

Java and cold starts#

A Java function initializes slowly: the JVM starts and frameworks load before the first request. SnapStart snapshots the initialized function when a version is published and resumes new environments from it, at no extra cost for Java. Provisioned concurrency keeps environments warm, for a fee.

Gotcha

SnapStart copies one initialized state into many environments. Anything unique made during initialization, such as a random seed or an ID, is duplicated: create it after restore.

Read more Improving startup performance with Lambda SnapStart · Building Lambda functions with Java

Workflows#

Settling a merchant's day takes several steps: validate the file, post to the ledger, notify the merchant, and undo if a step fails. The team models it as a Step Functions state machine instead of functions that call each other.

The settlement workflow: a new file starts an execution, and the state machine runs each step in turn, with retries and a path for undoingRegion eu-west-2stepsAmazon S3(settlementfiles)AWS StepFunctions(settlement)AWS Lambda(validate)AWS Lambda(post toledger)123
1An EventBridge rule starts an execution for each new file
2The file is validated first, with retries
3Posting runs exactly once, in a Standard workflow
Figure The settlement workflow: a new file starts an execution, and the state machine runs each step in turn, with retries and a path for undoing#
As text
  • Region eu-west-2
    • Amazon S3 (settlement files)
    • AWS Step Functions (settlement)
    • steps
      • AWS Lambda (validate)
      • AWS Lambda (post to ledger)
  1. An EventBridge rule starts an execution for each new file
  2. The file is validated first, with retries
  3. Posting runs exactly once, in a Standard workflow
AWS Step Functions
On AzureAzure Logic Apps, Durable Functions

What it does. AWS Step Functions runs workflows as state machines: steps, choices, retries, waits and parallel branches, with each execution's progress recorded.

How it works. A workflow is written in Amazon States Language and calls Lambda or over 220 other services directly. Standard workflows run up to a year, exactly once; Express workflows run up to five minutes, at least once.

When to use it. Multi-step processes that retry, wait for people or roll back, and must be auditable afterwards.

When not to. A single step, or logic the team would rather keep in Java: Lambda durable functions do that.

Limits. as of September 2026: 25,000 history events per Standard execution and 256 KiB per input or output. The workflow type cannot be changed later.

Cost. Standard is billed per state transition, retries included; Express per execution, duration and memory.

Security. Give each state machine its own role, allowed to call only its steps, and log executions to CloudWatch Logs.

Gotchas. Express workflows may run a step twice, so keep them to idempotent work; payments belong in Standard.

Read more AWS Step Functions documentation · Choosing workflow type in Step Functions · Step Functions service quotas · AWS Step Functions pricing

Choosing where multi-step work runs: one function, durable functions in code, or a Step Functions workflow of either typeyesnoyesnoyesnoDoes the work finish in one functioncall within 15 minutes?One Lambda functionWould the team rather write theworkflow in Java, inside Lambda?Lambda durable functionsIs it high volume, under five minutes,and safe to repeat?A Step Functions ExpressworkflowA Step Functions Standardworkflow
Figure Choosing where multi-step work runs: one function, durable functions in code, or a Step Functions workflow of either type#
As text
  1. Does the work finish in one function call within 15 minutes? Yes: One Lambda function. No: the next step.
  2. Would the team rather write the workflow in Java, inside Lambda? Yes: Lambda durable functions. No: the next step.
  3. Is it high volume, under five minutes, and safe to repeat? Yes: A Step Functions Express workflow. No: the next step.
  4. A Step Functions Standard workflow

Read more Lambda durable functions · Lambda Managed Instances

Object and file storage#

Part III · Building blocks · Chapter 20·4 min read

Amazon S3 keeps objects, Amazon EBS gives one instance a disk, and Amazon EFS and Amazon FSx share files. Choose by how data is reached, not how much there is.

Choosing storage by how it is reachedyesnoyesnoyesnoIs it a disk for one instance, such asa boot volume or a database?Amazon EBSMust many Linux clients share filesover NFS?Amazon EFSMust Windows clients share files overSMB, or does it need Lustre, ONTAP orZFS?Amazon FSxAmazon S3, over its API
Figure Choosing storage by how it is reached#
As text
  1. Is it a disk for one instance, such as a boot volume or a database? Yes: Amazon EBS. No: the next step.
  2. Must many Linux clients share files over NFS? Yes: Amazon EFS. No: the next step.
  3. Must Windows clients share files over SMB, or does it need Lustre, ONTAP or ZFS? Yes: Amazon FSx. No: the next step.
  4. Amazon S3, over its API

Read more Choosing an AWS storage service

Objects: Amazon S3#

Every merchant statement is a PDF in an S3 bucket, kept for seven years because the regulator says so.

Amazon S3
On AzureAzure Blob Storage

What it does. Amazon S3 stores objects, written whole and read by key, with strong read-after-write consistency.

How it works. Each object has a storage class; lifecycle rules move it colder or delete it, and versioning keeps old copies.

When to use it. Documents, backups, logs, data lakes and static content.

When not to. Files edited in place: use a file system, or mount the bucket with S3 Files.

Limits. as of September 2026: 10,000 buckets per account by default; a bucket's name and Region never change.

Cost. Storage by class, requests, retrievals and data transfer out.

Security. New buckets are private, block public access and encrypt every object by default.

Gotchas. Colder classes bill a minimum of 30 to 180 days, even if you delete sooner.

Read more Amazon S3 documentation · Understanding and managing Amazon S3 storage classes · Locking objects with Object Lock

A statement's life: moved to colder S3 Glacier classes, then deleted after seven yearsS3 StandardInstant RetrievalDeep Archiveafter 90 days1after a year2deleted after seven years3
Figure A statement's life: moved to colder S3 Glacier classes, then deleted after seven years#
As text
  1. S3 Standard to Instant Retrieval: after 90 days
  2. Instant Retrieval to Deep Archive: after a year
  3. Deep Archive: deleted after seven years
Gotcha

Object Lock can lock you out too. It makes objects write-once, and in compliance mode no one, the root user included, can delete one before its date. Test in governance mode first.

Disks: Amazon EBS#

The fraud team's GPU instances boot from EBS volumes and keep model files on a second gp3 volume.

Amazon EBS
On AzureAzure managed disks

What it does. Amazon EBS gives EC2 instances durable network disks.

How it works. A volume lives in one zone, replicated within it. gp3 sets size, IOPS and throughput separately; snapshots copy data to other zones and Regions.

When to use it. An instance's own disk: boot, database or working data.

When not to. Data that several instances must share.

Limits. as of September 2026: gp3 up to 64 TiB and 80,000 IOPS; io2 Block Express up to 256,000 IOPS.

Cost. What you provision, plus snapshot storage.

Security. Encryption by default is set per Region and skips existing volumes: turn it on early.

Gotchas. A volume cannot move zone: restore a snapshot in the new one.

Read more Amazon EBS documentation · Amazon EBS volume types · Enable Amazon EBS encryption by default

Shared files: EFS and FSx#

Settlement tasks in both zones share a directory of bank files on EFS.

One Regional EFS file system, mounted from both zonesRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPrivate subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24Amazon EFS(bank files)AWS Fargate(settlement)AWS Fargate(settlement)12
1Zone a mounts the file system
2Zone b mounts the same one
Figure One Regional EFS file system, mounted from both zones#
As text
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Amazon EFS (bank files)
      • Availability Zone a
        • Private subnet 10.20.10.0/24
          • AWS Fargate (settlement)
      • Availability Zone b
        • Private subnet 10.20.11.0/24
          • AWS Fargate (settlement)
  1. Zone a mounts the file system
  2. Zone b mounts the same one
Amazon EFS
On AzureAzure Files

What it does. Amazon EFS is shared NFS storage that grows and shrinks with its files.

How it works. A Regional file system spans zones, so clients in any zone mount it; lifecycle policies move cooling files to Infrequent Access and Archive classes.

When to use it. Linux instances, containers and functions that share files.

When not to. Windows clients, which EFS does not support.

Limits. A One Zone file system can be lost with its zone.

Cost. Storage by class, plus reads, writes and tiering.

Security. Encrypt in transit when mounting; control access with IAM, security groups and POSIX permissions.

Gotchas. Encryption at rest can only be chosen when the file system is created.

Read more Amazon EFS documentation · Amazon EFS pricing

The risk team's Windows reports read an SMB share on Amazon FSx for Windows File Server. FSx also runs NetApp ONTAP, OpenZFS and Lustre file systems. You provision storage, IOPS and throughput; a Windows file system is Single-AZ or Multi-AZ, and joins Active Directory when you create it.

Read more What is FSx for Windows File Server?

Relational databases#

Part III · Building blocks · Chapter 21·4 min read

Amazon RDS runs six familiar database engines for you; Amazon Aurora rebuilds MySQL and PostgreSQL on storage shared across three zones. Either way, the schema, the queries and their tuning stay yours.

The ledger is the payments team's most precious data: every posting, in PostgreSQL. It moves to Aurora PostgreSQL. The risk team's reports run on SQL Server, which Aurora does not offer, so that database moves to RDS for SQL Server.

Choosing a relational service: RDS for other engines, Aurora for PostgreSQL and MySQLyesnoyesnoMust it run SQL Server, Oracle, Db2 orMariaDB, or community PostgreSQL orMySQL?Amazon RDSIs the load spiky, or idle for hours ata time?Aurora Serverless v2Amazon Aurora with provisionedinstances
Figure Choosing a relational service: RDS for other engines, Aurora for PostgreSQL and MySQL#
As text
  1. Must it run SQL Server, Oracle, Db2 or MariaDB, or community PostgreSQL or MySQL? Yes: Amazon RDS. No: the next step.
  2. Is the load spiky, or idle for hours at a time? Yes: Aurora Serverless v2. No: the next step.
  3. Amazon Aurora with provisioned instances

Read more Choosing an AWS database service

Amazon RDS#

Amazon RDS
On AzureAzure SQL Database, Azure Database for PostgreSQL, Azure Database for MySQL

What it does. Amazon RDS runs Db2, MariaDB, SQL Server, MySQL, Oracle or PostgreSQL, and handles backups, patching and failover.

How it works. Pick an engine, instance class and storage. A Multi-AZ standby in another zone takes over on failure but serves no reads; read replicas do, copied asynchronously.

When to use it. Commercial engines, and community PostgreSQL or MySQL.

When not to. Spiky or mostly idle load: an instance bills for every hour it runs.

Limits. as of September 2026: 40 instances per Region, shared with Aurora; 15 read replicas each; backups kept up to 35 days.

Cost. Instance hours, storage, IOPS, backups and data transfer; each replica bills as an instance.

Security. Encryption can only be chosen at creation. Let RDS keep the master password in Secrets Manager, rotated every seven days.

Gotchas. A version past its standard support moves to paid Extended Support: plan upgrades.

Read more Amazon RDS documentation · Configuring and managing a Multi-AZ deployment for Amazon RDS · Amazon RDS Extended Support with Amazon RDS

Amazon Aurora#

Amazon Aurora
On AzureAzure Database for PostgreSQL, Azure SQL Database Hyperscale

What it does. Amazon Aurora is a MySQL- and PostgreSQL-compatible engine whose instances share one storage volume.

How it works. A writer and up to 15 readers share a volume kept on six storage nodes across three zones. The cluster endpoint always names the writer.

When to use it. PostgreSQL or MySQL systems needing fast failover, read scaling or a copy in another Region.

When not to. Other engines, which need RDS; writes in several Regions at once, which suit Aurora DSQL.

Limits. as of September 2026: a 256 TiB volume; a global database copies to up to 10 more Regions, typically under a second behind.

Cost. Instance hours or Serverless v2 capacity, storage used, and I/O unless the cluster is I/O-Optimized.

Security. Use IAM database authentication: tokens that last 15 minutes instead of stored passwords.

Gotchas. With no reader, a failed writer is recreated, typically within 10 minutes: keep a reader in another zone.

Read more Amazon Aurora User Guide · High availability for Amazon Aurora · Using Aurora serverless

An Aurora failover: the reader in zone b becomes the writer, typically within 60 seconds, and the cluster endpoint follows itPayments APIWriter (zone a)Reader (zone b)writes, through the cluster endpoint1fails2reconnects to the same endpoint, now this instance3
Figure An Aurora failover: the reader in zone b becomes the writer, typically within 60 seconds, and the cluster endpoint follows it#
As text
  1. Payments API to Writer (zone a): writes, through the cluster endpoint
  2. Writer (zone a): fails
  3. Payments API to Reader (zone b): reconnects to the same endpoint, now this instance
Gotcha

Failover changes DNS, and clients cache DNS. A client that keeps the old address keeps calling the failed instance. RDS Proxy bypasses DNS caches, cutting Aurora failover time by up to 66%, and pools connections so a surge of clients cannot overwhelm the database.

MegaCorp's design gains its data layer: the ledger's writer and reader sit in the private subnets of both zones.

MegaCorp's design with its data layer: the ledger's Aurora writer and reader in two zones, beside the statements bucketAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24InternetgatewayNAT gatewayAWS Fargate(payments)Amazon Aurora(ledger writer)NAT gatewayAWS Fargate(payments)Amazon Aurora(ledger reader)Amazon S3(statements)123
1Tasks reach the internet through the NAT gateway in their own zone
2The NAT gateway sends traffic out through the internet gateway
3Payments tasks write the ledger through the cluster endpoint
Figure MegaCorp's design with its data layer: the ledger's Aurora writer and reader in two zones, beside the statements bucket#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
            • AWS Fargate (payments)
            • Amazon Aurora (ledger writer)
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
            • AWS Fargate (payments)
            • Amazon Aurora (ledger reader)
      • Amazon S3 (statements)
  1. Tasks reach the internet through the NAT gateway in their own zone
  2. The NAT gateway sends traffic out through the internet gateway
  3. Payments tasks write the ledger through the cluster endpoint

Read more Using Amazon Aurora Global Database · Amazon RDS Proxy

NoSQL and caching#

Part III · Building blocks · Chapter 22·4 min read

Amazon DynamoDB trades joins for predictable speed at any size, and Amazon ElastiCache keeps hot data in memory. Both reward knowing how data will be read before you store it.

In the 2004 holiday season, outages at Amazon were traced to commercial technology pushed past its limits. Amazon built Dynamo, described it in a 2007 paper, and found it still hard to run. DynamoDB, launched in 2012, put Dynamo's scaling behind a service AWS operates.

Where data lives: a cache, DynamoDB, or a relational databaseyesnoyesnoIs it a copy kept only to make readsfaster?Amazon ElastiCacheIs every way of reading it known, anddone by key?Amazon DynamoDBA relational database: Aurora orRDS
Figure Where data lives: a cache, DynamoDB, or a relational database#
As text
  1. Is it a copy kept only to make reads faster? Yes: Amazon ElastiCache. No: the next step.
  2. Is every way of reading it known, and done by key? Yes: Amazon DynamoDB. No: the next step.
  3. A relational database: Aurora or RDS

Read more Choosing an AWS database service

Amazon DynamoDB#

Card schemes sometimes send a webhook twice. The settlement function writes each event ID to a DynamoDB table only if it is absent, so a repeat fails the condition and is skipped. Time to Live removes old IDs.

Skipping duplicate webhooks with a conditional writeCard schemeSettlement functionIdempotency tablewebhook for event 811put 81, only if absent2the same webhook again3put 81: condition fails, skip4
Figure Skipping duplicate webhooks with a conditional write#
As text
  1. Card scheme to Settlement function: webhook for event 81
  2. Settlement function to Idempotency table: put 81, only if absent
  3. Card scheme to Settlement function: the same webhook again
  4. Settlement function to Idempotency table: put 81: condition fails, skip
Amazon DynamoDB
On AzureAzure Cosmos DB

What it does. Amazon DynamoDB stores items by key and answers in single-digit milliseconds at any scale, with no servers or versions to manage.

How it works. A partition key is hashed to choose a partition, and a sort key orders items within it. Indexes add other keys. Data is copied across three zones.

When to use it. Known access by key at high or unpredictable volume: sessions, idempotency keys, carts.

When not to. Ad hoc queries and joins, which DynamoDB does not do.

Limits. as of September 2026: 400 KB per item; each partition serves 3,000 reads and 1,000 writes per second.

Cost. On-demand per request, the default, or provisioned per hour; plus storage and backups.

Security. Every call is authorized by IAM, with no passwords; data is encrypted at rest by default.

Gotchas. Reads are eventually consistent unless you ask for strong ones, which global secondary indexes cannot give.

Read more Amazon DynamoDB documentation · Core components of Amazon DynamoDB · DynamoDB read consistency

Gotcha

On-demand still has a ceiling on day one. A new on-demand table sustains 4,000 writes and 12,000 reads per second, and absorbs double its previous peak; grow faster within 30 minutes and requests are throttled. Pre-warm before a big launch.

Read more DynamoDB on-demand capacity mode

Amazon ElastiCache#

The merchant portal reads each merchant's settings on every page. The team caches them in Valkey, the open-source fork of Redis made in 2024: read the cache first, fall back to Aurora on a miss, and store the answer for five minutes.

Amazon ElastiCache
On AzureAzure Managed Redis, Azure Cache for Redis

What it does. Amazon ElastiCache runs Valkey, Memcached or Redis OSS as a managed in-memory cache.

How it works. Serverless caches scale themselves; node-based clusters are sized by node type and count, with replicas in other zones that take over in seconds.

When to use it. Hot reads in front of a database, session stores and leaderboards.

When not to. The only copy of important data: replication is asynchronous, so failover can lose recent writes.

Limits. as of September 2026: 40 serverless caches and 300 nodes per Region; Memcached has no replication.

Cost. Serverless by data stored and processing units, nodes by the hour; Valkey costs less than Redis OSS.

Security. Keep caches in private subnets, encrypt in transit, and authenticate with role-based access control.

Gotchas. Cached data goes stale when the database changes: give every key a time to live.

Read more Amazon ElastiCache documentation · Caching strategies

MegaCorp's design gains a NoSQL layer.

MegaCorp's NoSQL layer: the idempotency table in DynamoDB and the settings cache in ElastiCacheAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24InternetgatewayAmazonElastiCache(settings)NAT gatewayAWS Fargate(payments)NAT gatewayAWS Fargate(payments)Amazon DynamoDB(idempotencykeys)123
1Tasks reach the internet through the NAT gateway in their own zone
2The NAT gateway sends traffic out through the internet gateway
3Tasks read merchant settings from the cache before Aurora
Figure MegaCorp's NoSQL layer: the idempotency table in DynamoDB and the settings cache in ElastiCache#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Amazon ElastiCache (settings)
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
            • AWS Fargate (payments)
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
            • AWS Fargate (payments)
      • Amazon DynamoDB (idempotency keys)
  1. Tasks reach the internet through the NAT gateway in their own zone
  2. The NAT gateway sends traffic out through the internet gateway
  3. Tasks read merchant settings from the cache before Aurora

Messaging and events#

Part III · Building blocks · Chapter 23·5 min read

Queues hold work, topics copy messages, event buses route events, and streams keep an ordered record to replay. Choose by who needs a message, and for how long.

When a payment settles, the merchant wants a notice, statements need a line, and fraud wants every card event. The payments service publishes what happened, and each consumer subscribes.

Choosing how a message travels: a stream, an event bus, a topic, or a queueyesnoyesnoyesnoMust readers replay an ordered record,at high volume?Amazon Kinesis Data StreamsDo many services react to events,chosen by content?Amazon EventBridgeMust every subscriber get a copy pushedto it, people included?Amazon SNSAmazon SQS, for work oneconsumer does at its own pace
Figure Choosing how a message travels: a stream, an event bus, a topic, or a queue#
As text
  1. Must readers replay an ordered record, at high volume? Yes: Amazon Kinesis Data Streams. No: the next step.
  2. Do many services react to events, chosen by content? Yes: Amazon EventBridge. No: the next step.
  3. Must every subscriber get a copy pushed to it, people included? Yes: Amazon SNS. No: the next step.
  4. Amazon SQS, for work one consumer does at its own pace

Read more Amazon SQS, Amazon SNS, or Amazon EventBridge?

Queues: Amazon SQS#

Amazon SQS
On AzureAzure Service Bus, Azure Queue Storage

What it does. Amazon SQS holds messages until a consumer takes and deletes them: the buffer between producer and worker.

How it works. A received message hides for the visibility timeout and returns if it is not deleted. After several failed receives it moves to a dead-letter queue.

When to use it. Work one consumer does at its own pace, absorbing spikes.

When not to. Many consumers that each need a copy: put SNS or EventBridge in front.

Limits. as of September 2026: 1 MiB messages, kept up to 14 days; FIFO queues keep order within a message group.

Cost. Per request, each 64 KB counting as one; the first million each month are free.

Security. IAM and a queue policy decide who may send and receive; encryption uses KMS.

Gotchas. Standard queues deliver at least once, sometimes out of order: make consumers idempotent.

Read more Amazon SQS documentation · Using dead-letter queues in Amazon SQS

Gotcha

Dead letters keep their birthday. A message moved to a dead-letter queue keeps its original enqueue time, so it expires early unless the dead-letter queue keeps messages longer than the source queue.

Topics and buses: SNS and EventBridge#

Amazon SNS
On AzureAzure Service Bus, Azure Event Grid

What it does. Amazon SNS pushes each message published to a topic to every subscriber: queues, functions, HTTPS endpoints, email or SMS.

How it works. Subscribers can filter on attributes or body; FIFO topics keep order into SQS FIFO queues.

When to use it. Fan-out to a few known consumers, and notifications to people.

When not to. Routing on many rules across many services: EventBridge filters more richly.

Limits. as of September 2026: 256 KiB messages; in eu-west-2, 300 publishes a second per account by default.

Cost. Per publish and per delivery; delivery to SQS and Lambda is not charged per message.

Security. IAM and topic policies decide who publishes and who subscribes.

Gotchas. Messages are not kept: subscribe a queue for anything that must not be lost.

Read more Amazon SNS documentation · Amazon SNS message filtering

Amazon EventBridge
On AzureAzure Event Grid

What it does. Amazon EventBridge routes events from AWS services, your applications and SaaS partners to targets, by rules that match their content.

How it works. Rules on an event bus match event patterns and send each match to up to five targets. Pipes join one source to one target; Scheduler runs timed tasks.

When to use it. Events that many teams react to, across services and accounts.

When not to. Processing that depends on order, which EventBridge does not guarantee.

Limits. as of September 2026, in eu-west-2: 1,200 PutEvents requests a second and 300 rules per bus.

Cost. Events from AWS services on the default bus are free; custom events are charged per million.

Security. A resource policy on a bus lets other accounts send to it.

Gotchas. Delivery is at least once, retried for up to 24 hours: consumers must be idempotent.

Read more Amazon EventBridge documentation · Amazon EventBridge quotas

Streams: Kinesis Data Streams#

Amazon Kinesis Data Streams
On AzureAzure Event Hubs

What it does. Amazon Kinesis Data Streams keeps an ordered, replayable stream of records that many applications read independently.

How it works. A partition key picks a shard, and order holds within it; each shard takes 1 MB or 1,000 records a second.

When to use it. Clickstreams, telemetry and card events read by several consumers in near real time.

When not to. Work items that each need one consumer: that is a queue.

Limits. as of September 2026: records kept 24 hours by default, up to 365 days.

Cost. On-demand per GB written and read, or provisioned per shard-hour; longer retention costs more.

Security. IAM for producers and consumers, and server-side encryption with KMS.

Gotchas. A hot partition key is capped at one shard's limit, even on demand: choose keys that spread.

Read more Amazon Kinesis Data Streams documentation · What is Amazon Kinesis Data Streams?

Teams already running Apache Kafka can use Amazon MSK instead: the same Kafka APIs and tools, with AWS running the brokers.

Messaging on AWS, from the first queue to event busesAmazon SQS:queues, after a2004 beta2006Amazon SNS:publish once,deliver to many2010Amazon Kinesis:streams of datain real time2013EventBridge,from CloudWatchEvents2019EventBridgeScheduler andPipes2022
Figure Messaging on AWS, from the first queue to event buses#
As text
  1. 2006: Amazon SQS: queues, after a 2004 beta
  2. 2010: Amazon SNS: publish once, deliver to many
  3. 2013: Amazon Kinesis: streams of data in real time
  4. 2019: EventBridge, from CloudWatch Events
  5. 2022: EventBridge Scheduler and Pipes

MegaCorp's design gains its messaging layer.

MegaCorp's messaging layer: a queue, an event bus, a topic and a streamAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24InternetgatewayNAT gatewayNAT gatewayAmazon SQS(webhooks)AmazonEventBridge(payments)Amazon SNS(merchantalerts)Amazon KinesisData Streams(card events)12
1The NAT gateway sends traffic out through the internet gateway
2A rule sends settled payments to the merchant alerts topic
Figure MegaCorp's messaging layer: a queue, an event bus, a topic and a stream#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
      • Amazon SQS (webhooks)
      • Amazon EventBridge (payments)
      • Amazon SNS (merchant alerts)
      • Amazon Kinesis Data Streams (card events)
  1. The NAT gateway sends traffic out through the internet gateway
  2. A rule sends settled payments to the merchant alerts topic

Read more Welcome to the Amazon MSK Developer Guide

APIs and the edge#

Part III · Building blocks · Chapter 24·4 min read

Every outside request meets a front door: a load balancer, an API gateway or a content delivery network, with a firewall in front. Choose by what comes through.

Choosing a front door by what comes through ityesnoyesnoIs it web content or an API used farfrom its Region?Amazon CloudFront, in frontof the restIs it an API that needs keys, quotas orvalidation?Amazon API GatewayA load balancer: ALB for HTTP;NLB for TCP, UDP or fixed IPs
Figure Choosing a front door by what comes through it#
As text
  1. Is it web content or an API used far from its Region? Yes: Amazon CloudFront, in front of the rest. No: the next step.
  2. Is it an API that needs keys, quotas or validation? Yes: Amazon API Gateway. No: the next step.
  3. A load balancer: ALB for HTTP; NLB for TCP, UDP or fixed IPs

Read more What is Elastic Load Balancing?

Load balancers#

Elastic Load Balancing
On AzureAzure Load Balancer, Azure Application Gateway

What it does. Elastic Load Balancing spreads traffic across healthy targets in several zones, and scales itself.

How it works. An ALB routes HTTP by host, path or header; an NLB passes TCP and UDP at millions of requests a second.

When to use it. Web services and containers (ALB); other protocols or fixed IPs (NLB).

When not to. Classic Load Balancers, the previous generation: migrate them.

Limits. An NLB keeps each zone's traffic in that zone unless cross-zone balancing is on.

Cost. Hours running, plus capacity units for connections, bytes and rule evaluations.

Security. TLS ends at the balancer with ACM certificates; an ALB can take AWS WAF.

Gotchas. An NLB's targets see the client's IP address only for some target types.

Read more Elastic Load Balancing documentation · What is an Application Load Balancer?

APIs: Amazon API Gateway#

Amazon API Gateway
On AzureAzure API Management

What it does. Amazon API Gateway fronts Lambda, AWS services or private load balancers, throttling and authorizing each call.

How it works. REST APIs add keys, usage plans, validation, caching and WAF; HTTP APIs drop them for a lower price.

When to use it. Partner APIs such as the card-scheme webhooks, which need keys and quotas, and Lambda backends.

When not to. Plain web traffic to containers: an ALB is enough.

Limits. as of September 2026: 10 MB payloads; 10,000 requests a second per account and Region, shared by every API.

Cost. Per million calls, HTTP APIs costing less; a cache bills by the hour.

Security. IAM, Cognito or Lambda authorizers; a REST API can be private to your VPCs.

Gotchas. One busy API can use up the shared throttle: give each client a usage plan.

Read more Amazon API Gateway documentation · Choose between REST APIs and HTTP APIs

Gotcha

An integration gets 29 seconds. After that, API Gateway gives up even if the backend finishes. Regional REST APIs can wait longer at a cost to the account throttle; long work belongs on a queue.

The edge: CloudFront and AWS WAF#

Amazon CloudFront
On AzureAzure Front Door

What it does. Amazon CloudFront serves content from edge locations near users, caching what it can.

How it works. A distribution maps paths to origins such as S3, load balancers or API Gateway; objects stay cached 24 hours by default.

When to use it. Websites, downloads and APIs used far from their Region.

When not to. Non-HTTP traffic or fixed IPs: that is AWS Global Accelerator.

Limits. Edge code: CloudFront Functions for sub-millisecond work, Lambda@Edge for longer work.

Cost. Data out and requests, or a flat-rate plan; transfer from AWS origins is free.

Security. Keep S3 origins private with origin access control, and attach AWS WAF.

Gotchas. An S3 website endpoint cannot use origin access control.

Read more Amazon CloudFront documentation · Restrict access to an Amazon S3 origin

AWS WAF
On AzureAzure Web Application Firewall

What it does. AWS WAF inspects HTTP requests to CloudFront, ALBs and API Gateway REST APIs, and allows, blocks or counts them.

How it works. A web ACL holds rules on IPs, countries, headers, SQL injection or request rates, plus managed rule groups.

When to use it. Every public web entry point.

When not to. Non-HTTP traffic: use Network Firewall and security groups.

Limits. HTTP APIs cannot take a web ACL: put CloudFront in front of them.

Cost. Per web ACL and rule each month, and per million requests; bot rules cost extra.

Security. AWS Firewall Manager applies the same web ACLs across accounts.

Gotchas. Managed rules can block real customers: run them in Count mode first.

Read more AWS WAF documentation · What is AWS WAF?

AWS Shield Standard guards every AWS customer against common DDoS attacks at no extra charge; Shield Advanced is paid.

Front doors on AWS since 2009Elastic LoadBalancing2009Amazon APIGateway2015AWS WAF, firstfor CloudFront2015Application LoadBalancer2016AWS GlobalAccelerator2018
Figure Front doors on AWS since 2009#
As text
  1. 2009: Elastic Load Balancing
  2. 2015: Amazon API Gateway
  3. 2015: AWS WAF, first for CloudFront
  4. 2016: Application Load Balancer
  5. 2018: AWS Global Accelerator

MegaCorp's design gains its front door.

MegaCorp's front door: the web ACL, CloudFront and the load balancerAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPublic subnet 10.20.1.0/24Private subnet 10.20.11.0/24CustomersAWS WAF(web ACL)AmazonCloudFront(portal)InternetgatewayApplicationLoad Balancer(payments)NAT gatewayNAT gateway1234
1The NAT gateway sends traffic out through the internet gateway
2Requests meet the web ACL first
3CloudFront serves the requests it allows
4Dynamic requests go to the load balancer
Figure MegaCorp's front door: the web ACL, CloudFront and the load balancer#
As text
  • Customers
  • AWS WAF (web ACL)
  • Amazon CloudFront (portal)
  • AWS Account payments-prod
    • Region eu-west-2
      • VPC 10.20.0.0/16
        • Internet gateway
        • Application Load Balancer (payments)
        • Availability Zone a
          • Public subnet 10.20.0.0/24
            • NAT gateway
          • Private subnet 10.20.10.0/24
        • Availability Zone b
          • Public subnet 10.20.1.0/24
            • NAT gateway
          • Private subnet 10.20.11.0/24
  1. The NAT gateway sends traffic out through the internet gateway
  2. Requests meet the web ACL first
  3. CloudFront serves the requests it allows
  4. Dynamic requests go to the load balancer

Data and analytics#

Part III · Building blocks · Chapter 25·3 min read

Analytics on AWS starts with files in S3: AWS Glue reshapes and catalogues them, AWS Lake Formation grants access, and Amazon Athena or Amazon Redshift answers the questions.

Choosing where analytics queries runyesnoyesnoDo dashboards run the same queries overcurated data all day?Amazon RedshiftIs it SQL over files in S3, run now andthen?Amazon AthenaReshape the data first with anAWS Glue job
Figure Choosing where analytics queries run#
As text
  1. Do dashboards run the same queries over curated data all day? Yes: Amazon Redshift. No: the next step.
  2. Is it SQL over files in S3, run now and then? Yes: Amazon Athena. No: the next step.
  3. Reshape the data first with an AWS Glue job

Read more Choosing an AWS analytics service

Finance's data lake: Glue curates the payments data, and Athena and Redshift query itRegion eu-west-2S3 data lakeQueriesAWS Glue(jobs andcatalog)AWS LakeFormation(grants)Amazon S3(raw files)Amazon S3(curatedtables)Amazon Athena(ad hoc)Amazon Redshift(dashboards)123
1A Glue job reads the raw files
2It writes partitioned Parquet tables to the catalog
3Athena and Redshift query them within Lake Formation's grants
Figure Finance's data lake: Glue curates the payments data, and Athena and Redshift query it#
As text
  • Region eu-west-2
    • AWS Glue (jobs and catalog)
    • AWS Lake Formation (grants)
    • S3 data lake
      • Amazon S3 (raw files)
      • Amazon S3 (curated tables)
    • Queries
      • Amazon Athena (ad hoc)
      • Amazon Redshift (dashboards)
  1. A Glue job reads the raw files
  2. It writes partitioned Parquet tables to the catalog
  3. Athena and Redshift query them within Lake Formation's grants

AWS Glue runs serverless Spark jobs and keeps the Data Catalog that Athena, Redshift and Amazon EMR read. AWS Lake Formation adds grants to that catalog, down to columns, rows and cells.

Read more What is AWS Glue? · What is AWS Lake Formation?

Athena and Redshift#

Amazon Athena
On AzureMicrosoft Fabric, Azure Synapse Analytics

What it does. Amazon Athena runs standard SQL, or Spark, on data in S3, with nothing to provision.

How it works. Tables live in the Glue Data Catalog; Athena scans the files in parallel and returns results in seconds.

When to use it. Ad hoc questions, log analysis and exploring the lake.

When not to. Dashboards repeating the same queries all day: that is a warehouse's job.

Limits. as of September 2026: concurrent queries are capped per Region; a scan reads at most 1 million partitions.

Cost. Per terabyte scanned, 10 MB minimum per query.

Security. IAM and Lake Formation decide what each analyst reads.

Gotchas. Results land in an S3 bucket: whoever reads it sees every answer.

Read more Amazon Athena documentation · Amazon Athena pricing

Amazon Redshift
On AzureMicrosoft Fabric, Azure Synapse Analytics

What it does. Amazon Redshift is a fully managed, petabyte-scale data warehouse that BI tools query with SQL.

How it works. Serverless scales in seconds and bills only while queries run; provisioned clusters run RA3 nodes. Both can query S3 too.

When to use it. Curated data that many dashboards and analysts query all day.

When not to. Occasional questions over raw files, where Athena needs no warehouse.

Limits. as of September 2026: Python user-defined functions lose support after June 30, 2026.

Cost. Serverless per RPU-hour, by the second with a 60-second minimum; or node hours; plus storage.

Security. Lake Formation can govern its shared data down to rows and columns.

Gotchas. Every Serverless query bills at least 60 seconds, so floods of tiny queries add up.

Read more Amazon Redshift documentation · What is Amazon Redshift Serverless?

Gotcha

SELECT * is a spending decision. Athena bills the bytes it scans: store tables as compressed Parquet, partition them by date, and name only the columns you need.

Analytics on AWS, from a warehouse to serverless queriesAmazon Redshift2012Amazon Athena: SQL onS32016AWS Glue2017Redshift Serverless2022
Figure Analytics on AWS, from a warehouse to serverless queries#
As text
  1. 2012: Amazon Redshift
  2. 2016: Amazon Athena: SQL on S3
  3. 2017: AWS Glue
  4. 2022: Redshift Serverless

Generative AI#

Part III · Building blocks · Chapter 26·3 min read

Amazon Bedrock serves foundation models from many providers behind one API, with retrieval, guardrails and agents around them. The design questions stay familiar: data, identity, cost and limits.

MegaCorp's support assistant answers merchants' payout questions from the merchant handbook, with no investment advice and no card numbers.

Choosing how far to go beyond a promptyesnoyesnoMust answers draw on your owndocuments?A Bedrock knowledge base,for retrieval augmentedgenerationMust the model behave in ways promptingcannot reach?Customize: Bedrockfine-tuning, or SageMakerAI for full controlA Bedrock model and a goodprompt
Figure Choosing how far to go beyond a prompt#
As text
  1. Must answers draw on your own documents? Yes: A Bedrock knowledge base, for retrieval augmented generation. No: the next step.
  2. Must the model behave in ways prompting cannot reach? Yes: Customize: Bedrock fine-tuning, or SageMaker AI for full control. No: the next step.
  3. A Bedrock model and a good prompt

Read more Amazon Bedrock or Amazon SageMaker AI?

Amazon Bedrock
On AzureMicrosoft Foundry

What it does. Amazon Bedrock, generally available since 2023, runs foundation models from Amazon, Anthropic, OpenAI and others behind one serverless API.

How it works. Applications call a model by ID through the Converse or OpenAI-compatible APIs; knowledge bases add retrieval, guardrails filter, and AgentCore runs agents.

When to use it. Chat, summaries, extraction and agents built on existing models.

When not to. Training your own models: that is SageMaker AI.

Limits. as of September 2026: each model has token quotas per Region, and some models use tokens faster.

Cost. Tokens in and out on demand; batch at half price for some models; provisioned throughput by the hour.

Security. Model providers never see your prompts or completions; reach Bedrock privately through PrivateLink.

Gotchas. A global inference profile may run a request in any commercial Region: use a geographic one where data residency matters.

Read more Amazon Bedrock documentation · Amazon Bedrock pricing · Data protection in Amazon Bedrock

Knowledge bases and guardrails#

Answering from the handbook, with a guardrail on the way outSupport appKnowledge baseModelthe merchant's question1the question and the passages that matchit2an answer with citations, checked by the guardrail3
Figure Answering from the handbook, with a guardrail on the way out#
As text
  1. Support app to Knowledge base: the merchant's question
  2. Knowledge base to Model: the question and the passages that match it
  3. Model to Support app: an answer with citations, checked by the guardrail

The team's guardrail denies investment advice as a topic, masks card numbers, and blocks answers not grounded in the handbook. It checks prompts and answers alike, and the ApplyGuardrail API applies it to models outside Bedrock.

Gotcha

Retrieval ignores who is asking, unless you tell it. Only managed knowledge bases filter documents by each user's permissions; a vector store you manage returns any matching passage to anyone.

Read more Amazon Bedrock Knowledge Bases · Amazon Bedrock Guardrails

The assistant reaches Bedrock through an interface endpointRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPrivate subnet 10.20.10.0/24Amazon Bedrock(models andknowledge base)Amazon S3(merchanthandbook)AWS Fargate(supportassistant)VPC endpoints(interfaceendpoint)123
1The knowledge base indexes the handbook
2The assistant calls Bedrock through the endpoint
3which carries the request on the AWS network
Figure The assistant reaches Bedrock through an interface endpoint#
As text
  • Region eu-west-2
    • Amazon Bedrock (models and knowledge base)
    • Amazon S3 (merchant handbook)
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Private subnet 10.20.10.0/24
          • AWS Fargate (support assistant)
          • VPC endpoints (interface endpoint)
  1. The knowledge base indexes the handbook
  2. The assistant calls Bedrock through the endpoint
  3. which carries the request on the AWS network
Building blocks
  1. The nightly settlement job runs for 40 minutes. Lambda, or something else?
    Answer
    Not Lambda, which stops at 15 minutes: run an ECS task on Fargate, or split the work into a Step Functions workflow.
  2. Card-scheme webhooks sometimes arrive twice. How does the settlement function avoid posting twice?
    Answer
    It writes each event ID to DynamoDB only if the ID is absent; a repeat fails the condition and is skipped.
  3. Finance's Athena bill doubles every month. What do you check first?
    Answer
    How much each query scans: store tables as compressed Parquet, partition them by date, and select only the columns needed.

Keys, secrets and certificates#

Part IV · Enterprise-grade · Chapter 27·5 min read

AWS splits what Azure keeps in Key Vault: AWS KMS holds keys, AWS Secrets Manager holds secrets, and AWS Certificate Manager issues and renews TLS certificates.

False friend: Key Vault
On Azure

One key vault stores keys, secrets and certificates, with access granted through Azure RBAC or vault access policies.

On AWS

KMS keeps only keys, which never leave it unencrypted; passwords go to Secrets Manager and certificates to ACM, each with its own permissions.

The ledger's password used to sit in a configuration file. Alex moves it to Secrets Manager, where RDS rotates it every seven days, and encrypts the settlement files under a customer managed key, so the audit trail shows every decrypt.

Where a sensitive value belongsyesnoyesnoyesnoIs it a key that encrypts or signsdata?AWS KMSIs it a TLS certificate for a loadbalancer, CloudFront or API Gateway?AWS Certificate ManagerIs it a password, API key or token thatshould rotate?AWS Secrets ManagerPlain configuration: ParameterStore
Figure Where a sensitive value belongs#
As text
  1. Is it a key that encrypts or signs data? Yes: AWS KMS. No: the next step.
  2. Is it a TLS certificate for a load balancer, CloudFront or API Gateway? Yes: AWS Certificate Manager. No: the next step.
  3. Is it a password, API key or token that should rotate? Yes: AWS Secrets Manager. No: the next step.
  4. Plain configuration: Parameter Store

Read more AWS Systems Manager Parameter Store

Keys: AWS KMS#

Envelope encryption: KMS hands out a data key, and only the encrypted copy is stored with the dataSettlement jobAWS KMSAmazon S3GenerateDataKey under the payments key1a data key, in plaintext and encrypted2the file encrypted with the data key, plus the encrypted key3forgets the plaintext key4
Figure Envelope encryption: KMS hands out a data key, and only the encrypted copy is stored with the data#
As text
  1. Settlement job to AWS KMS: GenerateDataKey under the payments key
  2. AWS KMS to Settlement job: a data key, in plaintext and encrypted
  3. Settlement job to Amazon S3: the file encrypted with the data key, plus the encrypted key
  4. Settlement job: forgets the plaintext key
AWS KMS
On AzureAzure Key Vault

What it does. AWS KMS creates and controls the keys that encrypt and sign your data, in hardware security modules they never leave unencrypted.

How it works. Services such as S3, EBS and RDS ask KMS for data keys and encrypt the data themselves; KMS only wraps and unwraps the data keys.

When to use it. Customer managed keys where you must control policy, rotation and audit; AWS owned keys where convenience matters most.

When not to. Storing passwords or certificates: those belong in Secrets Manager and ACM.

Limits. as of September 2026: in eu-west-2, 20,000 symmetric requests a second, shared across the account.

Cost. A monthly fee per customer managed key, plus requests beyond a free tier; AWS owned keys are free.

Security. A key policy decides who may use each key, and CloudTrail records each use of your keys.

Gotchas. Deleting a key makes everything encrypted under it unrecoverable: disable it first, and the deletion itself waits 7 to 30 days.

Read more AWS KMS documentation · AWS KMS keys · AWS KMS request quotas

Gotcha

Your buckets spend your KMS quota. Every upload or download of an S3 object encrypted with SSE-KMS is a KMS request made on your behalf, counted against the same account quota as your own calls. A busy bucket can throttle an unrelated service.

Secrets: AWS Secrets Manager#

AWS Secrets Manager
On AzureAzure Key Vault

What it does. AWS Secrets Manager stores database credentials, API keys and tokens, and gives them to code at run time.

How it works. Each secret is encrypted with KMS; rotation replaces it on a schedule, managed for RDS or through a Lambda function.

When to use it. Any credential that would otherwise sit in code or configuration.

When not to. Plain configuration values, which Parameter Store's standard tier holds at no extra charge.

Limits. A secret lives in one Region unless you replicate it.

Cost. Per secret per month, replicas included, and per 10,000 API calls; rotation functions bill as Lambda.

Security. Let each role read only its own secrets.

Gotchas. Rotation breaks code that caches a password forever: fetch the secret again when a login fails.

Read more AWS Secrets Manager documentation · What is AWS Secrets Manager?

Certificates: AWS Certificate Manager#

AWS Certificate Manager
On AzureAzure Key Vault

What it does. AWS Certificate Manager issues, stores and renews TLS certificates for load balancers, CloudFront and API Gateway.

How it works. Request a certificate for your domains, or import one; ACM renews the certificates it issues.

When to use it. HTTPS on AWS's own front doors, where public certificates cost nothing.

When not to. Certificates on your own servers: use ACME automation, or pay for exportable ones.

Limits. Certificates are Regional and cannot be copied; CloudFront uses only us-east-1.

Cost. Free for public certificates on integrated services; exportable ones cost per domain.

Security. A wildcard certificate covers every subdomain, so issue it sparingly.

Gotchas. Every Region that serves the domain needs its own certificate, validated there.

Read more AWS Certificate Manager documentation · What is AWS Certificate Manager?

Keys, certificates and secrets on AWSAWS KMS2014AWS CertificateManager, freecertificates2016AWS Secrets Manager2018
Figure Keys, certificates and secrets on AWS#
As text
  1. 2014: AWS KMS
  2. 2016: AWS Certificate Manager, free certificates
  3. 2018: AWS Secrets Manager

MegaCorp's design gains its keys: the payments key and the ledger's credentials in eu-west-2, and the portal's certificate in us-east-1.

MegaCorp's keys layer: a customer managed key and a secret in eu-west-2, and CloudFront's certificate in us-east-1AWS Account payments-prodRegion eu-west-2Region us-east-1AWS WAF(web ACL)AmazonCloudFront(portal)AWS KMS(payments key)AWS SecretsManager(ledger)AWS CertificateManager(portal)123
1CloudFront serves the requests it allows
2CloudFront uses a certificate from us-east-1
3The ledger's credentials are encrypted under the payments key
Figure MegaCorp's keys layer: a customer managed key and a secret in eu-west-2, and CloudFront's certificate in us-east-1#
As text
  • AWS WAF (web ACL)
  • Amazon CloudFront (portal)
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS KMS (payments key)
      • AWS Secrets Manager (ledger)
    • Region us-east-1
      • AWS Certificate Manager (portal)
  1. CloudFront serves the requests it allows
  2. CloudFront uses a certificate from us-east-1
  3. The ledger's credentials are encrypted under the payments key

Read more Delete an AWS KMS key · AWS Certificate Manager pricing

Detection and posture#

Part IV · Enterprise-grade · Chapter 28·4 min read

Four questions keep an estate safe: who did what, how is it configured, is anyone attacking, and where do we stand. CloudTrail, Config, GuardDuty and Security Hub answer them.

Which service answers the questionyesnoyesnoyesnoWho did what, and when?AWS CloudTrailHow is a resource configured, and hasit drifted?AWS ConfigIs someone attacking us now?Amazon GuardDutyWhere do we stand, acrossaccounts? AWS Security Hub
Figure Which service answers the question#
As text
  1. Who did what, and when? Yes: AWS CloudTrail. No: the next step.
  2. How is a resource configured, and has it drifted? Yes: AWS Config. No: the next step.
  3. Is someone attacking us now? Yes: Amazon GuardDuty. No: the next step.
  4. Where do we stand, across accounts? AWS Security Hub

Read more What is AWS CloudTrail?

Records: CloudTrail and Config#

AWS CloudTrail
On AzureAzure Monitor

What it does. AWS CloudTrail records API calls made from the console, CLI, SDKs and AWS services.

How it works. Event history keeps 90 days of management events per Region; a trail delivers events to S3, for one account or the whole organization, for as long as you keep them.

When to use it. Always, with one organization trail into a separate log-archive account.

When not to. Application logs: those belong in CloudWatch Logs.

Limits. as of September 2026: event history covers 90 days, one Region at a time.

Cost. The first copy of management events is free; data events and extra copies are charged.

Security. Deliver to a bucket in another account that the workload's admins cannot change.

Gotchas. Event history is per Region: activity in a Region you never use shows up only there.

Read more AWS CloudTrail documentation · AWS CloudTrail pricing

AWS Config
On AzureAzure Policy, Azure Resource Graph

What it does. AWS Config records how each resource is configured, and every change, over time.

How it works. Rules check resources as they change and flag noncompliant ones; conformance packs bundle rules, and an aggregator gathers every account and Region into one view.

When to use it. Proving how something was configured on a given day, and catching drift.

When not to. Recording every resource type everywhere without a plan: each change is billed.

Limits. It records only the resource types you choose.

Cost. Per configuration item recorded and per rule evaluation.

Security. Remediation can fix a noncompliant resource automatically.

Gotchas. Rules detect after the fact; to stop a change happening at all, use a service control policy.

Read more AWS Config documentation · AWS Config pricing

Threats and posture: GuardDuty and Security Hub#

Amazon GuardDuty
On AzureMicrosoft Defender for Cloud

What it does. Amazon GuardDuty detects threats such as stolen credentials, cryptomining and data exfiltration.

How it works. Once enabled, it analyses CloudTrail management events, VPC flow logs and DNS logs with no agents; protection plans add S3, EKS, RDS, Lambda and runtime monitoring.

When to use it. Every account and Region, managed from one delegated administrator.

When not to. Checking configuration against a standard: that is Security Hub.

Limits. It watches only the Regions where you enable it.

Cost. Per million events and per GB of logs analysed, after a 30-day free trial.

Security. Route high-severity findings through EventBridge to whoever is on call.

Gotchas. Protection plans launched after you enabled GuardDuty stay off until you turn them on.

Read more Amazon GuardDuty documentation · Amazon GuardDuty pricing

AWS Security Hub
On AzureMicrosoft Defender for Cloud

What it does. AWS Security Hub scores the estate against security standards and gathers findings from GuardDuty, Inspector, Macie and partners in one place.

How it works. Controls from the AWS Foundational Security Best Practices, CIS, PCI DSS and NIST standards run as checks; findings share one format, and automation rules and EventBridge act on them.

When to use it. One view of security across every account, for security teams and auditors.

When not to. Detecting threats itself: it relies on GuardDuty and others for that.

Limits. It sees only Regions where it is enabled, and only findings from after that.

Cost. Per security check and per finding ingested, after a 30-day free trial.

Security. Run it from a delegated administrator account for the whole organization.

Gotchas. Most controls need AWS Config recording, which is billed separately.

Read more AWS Security Hub documentation · Introduction to AWS Security Hub CSPM

Amazon Inspector continually scans EC2 instances, ECR images and Lambda functions for known vulnerabilities, and reports to Security Hub too.

Gotcha

Watch the Regions you don't use. AWS recommends enabling GuardDuty in every supported Region, even where you run nothing, so it can report unusual activity there.

From finding to people: a high-severity finding pages the on-call engineerAmazon GuardDutyAmazon EventBridgeOn-calla finding: keys used from an unfamiliarnetwork1a rule matches high severity andpublishes to SNS2contain first, revoke the keys, theninvestigate3
Figure From finding to people: a high-severity finding pages the on-call engineer#
As text
  1. Amazon GuardDuty to Amazon EventBridge: a finding: keys used from an unfamiliar network
  2. Amazon EventBridge to On-call: a rule matches high severity and publishes to SNS
  3. On-call: contain first, revoke the keys, then investigate

AWS Security Incident Response can triage findings for you, escalating fewer than 1% of them, and engages engineers within 15 minutes on cases you raise.

MegaCorp's design gains its detection layer, run from the audit account.

MegaCorp's detection layer: Config in payments-prod, and GuardDuty and Security Hub in the audit accountAWS Account payments-prodRegion eu-west-2AWS Account auditdelegated administratorAWS Config(recorder)AWS SecurityHub(all findings)AmazonGuardDuty(organization)12
1Config records each change, and Security Hub checks it against its standards
2GuardDuty findings from every account land in Security Hub
Figure MegaCorp's detection layer: Config in payments-prod, and GuardDuty and Security Hub in the audit account#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS Config (recorder)
  • AWS Account audit
    • delegated administrator
      • AWS Security Hub (all findings)
      • Amazon GuardDuty (organization)
  1. Config records each change, and Security Hub checks it against its standards
  2. GuardDuty findings from every account land in Security Hub

Read more What is Amazon Inspector? · What is AWS Security Incident Response?

Observability#

Part IV · Enterprise-grade · Chapter 29·3 min read

Metrics say something is wrong, logs say what happened, and traces say where. Amazon CloudWatch holds the metrics, logs and alarms, and AWS X-Ray follows a request from service to service.

At two in the morning, settlement slows. An alarm on the function's 99th-percentile latency pages Alex; the trace shows the ledger query taking eight seconds; a Logs Insights query finds the lock timeouts behind it.

Which signal answers the questionyesnoyesnoyesnoMust you know within minutes thatsomething is wrong?A CloudWatch metric, withan alarmDo you need what happened, line byline?CloudWatch Logs, queriedwith Logs InsightsDo you need where the time went, acrossservices?A trace in X-Ray orApplication SignalsA dashboard that puts them sideby side
Figure Which signal answers the question#
As text
  1. Must you know within minutes that something is wrong? Yes: A CloudWatch metric, with an alarm. No: the next step.
  2. Do you need what happened, line by line? Yes: CloudWatch Logs, queried with Logs Insights. No: the next step.
  3. Do you need where the time went, across services? Yes: A trace in X-Ray or Application Signals. No: the next step.
  4. A dashboard that puts them side by side

Read more Metrics concepts

Metrics, logs and alarms: Amazon CloudWatch#

Amazon CloudWatch
On AzureAzure Monitor

What it does. Amazon CloudWatch collects metrics and logs from AWS services and your code, alarms on them, and draws dashboards.

How it works. Services publish their metrics free; your code adds custom metrics through OpenTelemetry or PutMetricData, and writes logs to log groups that Logs Insights queries.

When to use it. Monitoring everything on AWS, from one team's dashboard to a cross-account monitoring account.

When not to. Keeping every debug line forever: set a retention, or route old logs to S3.

Limits. as of September 2026: metrics stay in their Region, with 1-minute detail for 15 days and hourly data for 15 months.

Cost. Custom metrics per month, log ingestion and storage per GB, queries per GB scanned, and alarms.

Security. Mask sensitive data, such as card numbers, with a data protection policy on the log group.

Gotchas. Each unique combination of dimensions is a separate custom metric, billed on its own.

Read more Amazon CloudWatch documentation · What is Amazon CloudWatch Logs? · Amazon CloudWatch pricing

An alarm fires on a sustained change, not a single spikeSettlement functionCloudWatch alarmOn-callp99 latency above two seconds, threeminutes running1the state changes to ALARM, and SNS pages2open the trace, then query the logs3
Figure An alarm fires on a sustained change, not a single spike#
As text
  1. Settlement function to CloudWatch alarm: p99 latency above two seconds, three minutes running
  2. CloudWatch alarm to On-call: the state changes to ALARM, and SNS pages
  3. On-call: open the trace, then query the logs
Gotcha

Logs never expire unless you say so. A new log group keeps its logs forever by default, and storage is billed every month. Set a retention on every log group when you create it.

Traces: AWS X-Ray#

AWS X-Ray
On AzureApplication Insights

What it does. AWS X-Ray follows each request through your services and the AWS resources, databases and APIs they call.

How it works. Instrumented code sends segments, and integrated services such as Lambda add their own; X-Ray joins them into traces and a map of services.

When to use it. Finding which hop in a chain of services is slow or failing.

When not to. Counting things or keeping audit records: those are metrics and logs.

Limits. A trace covers only instrumented code and integrated services.

Cost. Trace volume drives the bill, and sampling rules control the volume.

Security. Keep secrets and card numbers out of trace annotations.

Gotchas. By default only the first request each second and 5% of the rest are traced, so a rare failure may be missed.

Read more AWS X-Ray documentation · What is AWS X-Ray? · Configuring sampling rules

CloudWatch Application Signals instruments Java, Python, Node.js and .NET services on EKS, ECS, EC2 and Lambda through OpenTelemetry, and shows latency, errors and service level objectives without a dashboard to build.

Read more Application Signals

MegaCorp's design gains its observability layer: payments-prod shares its signals with a monitoring account.

MegaCorp's observability layer: CloudWatch and X-Ray in payments-prod, seen from the monitoring accountAWS Account payments-prodRegion eu-west-2AWS Account monitoringAmazonCloudWatch(payments)AWS X-Ray(traces)AmazonCloudWatch(monitoring)12
1payments-prod shares its metrics and logs with the monitoring account
2and its traces
Figure MegaCorp's observability layer: CloudWatch and X-Ray in payments-prod, seen from the monitoring account#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • Amazon CloudWatch (payments)
      • AWS X-Ray (traces)
  • AWS Account monitoring
    • Amazon CloudWatch (monitoring)
  1. payments-prod shares its metrics and logs with the monitoring account
  2. and its traces

Read more Using Amazon CloudWatch alarms

Resilience and disaster recovery#

Part IV · Enterprise-grade · Chapter 30·7 min read

A zone fails, a Region fails, or someone deletes the data. Zones answer the first, a disaster recovery strategy the second, and backups kept out of reach the third.

MegaCorp's payments board sets two numbers. Payments must flow again within an hour: the recovery time objective, or RTO. At most a minute of payments may be lost: the recovery point objective, or RPO. Lower numbers cost more.

Read more Recovery objectives (RTO and RPO)

Losing a zone#

Most disasters hit only one zone, so running in several already covers much of the risk. Make the design statically stable: the capacity to survive a lost zone runs before the failure, because launching it during one relies on control planes and on spare capacity. In two zones, each must carry the whole load; in three, each carries half.

Amazon Application Recovery Controller
On Azureno direct equivalent

What it does. Amazon Application Recovery Controller (ARC) moves traffic away from a failing zone or Region.

How it works. A zonal shift takes a load balancer's or Auto Scaling group's traffic out of one zone; with zonal autoshift, AWS starts the shift when a zone is impaired. Routing controls fail over between Regions.

When to use it. A bad deployment or an impaired zone; a rehearsed Region failover.

When not to. Automatic Region failover on one alarm: a false alarm costs availability and data.

Limits. as of September 2026: a zonal shift lasts up to three days, extendable, for ALBs, NLBs, Auto Scaling groups and EKS.

Cost. Zonal shift is free; routing control clusters are billed.

Security. Safety rules can keep only one Region switched on at a time.

Gotchas. A load balancer that is failing open ignores a zonal shift.

Read more Amazon Application Recovery Controller documentation · Zonal shift in ARC · ARC pricing

Gotcha

Shift only onto capacity that exists. ARC moves traffic, not servers: before a zonal shift, the other zones must already be able to carry the extra load.

AWS Fault Injection Service
On AzureAzure Chaos Studio

What it does. AWS Fault Injection Service (FIS) breaks things on purpose, so you learn how the system copes before a real failure teaches you.

How it works. An experiment template names actions, such as stopping tasks or failing over a database; targets, chosen by tag; and stop conditions, CloudWatch alarms that end it early. Ready-made scenarios include AZ Availability: Power Interruption.

When to use it. Proving that failover, alarms and runbooks work.

When not to. Production, before the experiment has passed in pre-production.

Limits. A single-account experiment reaches only its own account.

Cost. Per action-minute, more for each extra target account.

Security. It acts through an IAM role you give it; limit that role to the targets.

Gotchas. The actions are real: without a stop condition, a test can become an outage.

Read more AWS Fault Injection Service documentation · AWS FIS scenarios reference · AWS FIS pricing

Read more REL11-BP05 Use static stability

Losing a Region: four strategies#

Guarding against the loss of a Region costs more. AWS names four strategies, in rising order of cost and falling RTO and RPO.

Choosing a disaster recovery strategy from the RTO and RPOyesnoyesnoyesnoCan the business wait up to a day, andlose hours of data?Backup and restoreCan it wait tens of minutes, and loseminutes of data?Pilot lightCan it wait minutes, and lose secondsof data?Warm standbyMulti-site active/active, nearzero
Figure Choosing a disaster recovery strategy from the RTO and RPO#
As text
  1. Can the business wait up to a day, and lose hours of data? Yes: Backup and restore. No: the next step.
  2. Can it wait tens of minutes, and lose minutes of data? Yes: Pilot light. No: the next step.
  3. Can it wait minutes, and lose seconds of data? Yes: Warm standby. No: the next step.
  4. Multi-site active/active, near zero
What waits in the recovery Region: each strategy keeps more running, so recovery is faster and costs moreIn the recovery RegionNo compute runningBackup and restorePilot lightCompute runningWarm standbyMulti-site active/activeAWS Backup(copiedbackups)Amazon Aurora(live copy)Amazon Aurora(live copy)AWS Fargate(a few tasks)Amazon Aurora(live copy)AWS Fargate(full size,serving)
Figure What waits in the recovery Region: each strategy keeps more running, so recovery is faster and costs more#
As text
  • In the recovery Region
    • No compute running
      • Backup and restore
        • AWS Backup (copied backups)
      • Pilot light
        • Amazon Aurora (live copy)
    • Compute running
      • Warm standby
        • Amazon Aurora (live copy)
        • AWS Fargate (a few tasks)
      • Multi-site active/active
        • Amazon Aurora (live copy)
        • AWS Fargate (full size, serving)

Backup and restore rebuilds from backups and infrastructure as code; pilot light keeps data live but deploys compute only when needed; warm standby only has to scale up; multi-site active/active already serves from every Region.

MegaCorp picks pilot light, in eu-west-1 and a separate account, payments-dr. An Aurora global database keeps the ledger there, typically under a second behind.

MegaCorp's failover runbook: data first, then compute, then trafficOn-callAurora global databaseARC routing controlpromote eu-west-1 to take writes, inunder a minute1deploy the payments tasks in eu-west-1, andscale them out2switch traffic from eu-west-2 to eu-west-13
Figure MegaCorp's failover runbook: data first, then compute, then traffic#
As text
  1. On-call to Aurora global database: promote eu-west-1 to take writes, in under a minute
  2. On-call: deploy the payments tasks in eu-west-1, and scale them out
  3. On-call to ARC routing control: switch traffic from eu-west-2 to eu-west-1

A person starts the failover, since a false alarm costs availability and data. The steps are scripted, and ARC's data plane is designed for higher availability than control planes.

False friend: region pair
On Azure

Microsoft pairs many regions, and geo-redundant storage copies data to the pair on its own. Many newer regions have no pair.

On AWS

No Region has a partner. Nothing leaves a Region unless a replication or copy feature, or your code, sends it, and you choose the recovery Region.

For whole servers, AWS Elastic Disaster Recovery replicates machines from a data centre or another cloud into a low-cost staging area, and launches them within minutes.

Read more Disaster recovery options in the cloud · REL13-BP02 Use defined recovery strategies · What is Elastic Disaster Recovery? · Azure region pairs and nonpaired regions

Losing the data: backups#

Replication copies mistakes too: a deleted ledger row is gone from eu-west-1 within a second. Only a point-in-time backup goes back to before the mistake.

AWS Backup
On AzureAzure Backup

What it does. AWS Backup schedules, copies and keeps backups of EC2, EBS, S3, RDS, Aurora, DynamoDB, EFS, FSx and more, from one place.

How it works. A backup plan sets frequency and retention, and resources join it by tag. Backups land in vaults, and copies go to other Regions and accounts.

When to use it. Every data store, with backup policies applying plans across the organization.

When not to. Alone, for a low RTO or RPO: restores take hours.

Limits. It governs only its own backups, and some resource types copy in full every time.

Cost. Storage per GB-month, restores per GB, and copies between Regions; no service fee.

Security. Copy to a locked vault in another account: not even the root user can delete a backup early.

Gotchas. A compliance-mode lock is permanent once its grace time, at least three days, ends.

Read more AWS Backup documentation · AWS Backup Vault Lock · AWS Backup pricing

Quotas#

Every account starts with default quotas, many of them per Region. Lambda functions in a Region share 1,000 concurrent executions by default as of September 2026. MegaCorp raised that in eu-west-2; in eu-west-1 it is still the default.

Service Quotas
On Azureno direct equivalent

What it does. Service Quotas shows each quota, the maximum for a resource or action, and requests increases.

How it works. Services set defaults; adjustable quotas rise by request, which support may approve, deny or partly approve. Automatic Management warns before a quota runs out.

When to use it. Before launch, before big events, and for the recovery Region.

When not to. Fixed quotas: design around them.

Limits. as of September 2026: a request template raises up to 10 quotas in each new account of an organization.

Cost. Free; alarms on quota usage bill as CloudWatch alarms.

Security. Alarm on usage: a sudden climb can mean a runaway job or stolen credentials.

Gotchas. Increases take time to review, and global quotas are raised only from us-east-1.

Read more Service Quotas documentation · What is Service Quotas? · Quota request templates

Gotcha

Raise quotas in the recovery Region too. A pilot light scales up during failover, so the recovery Region's quotas must already allow production capacity.

MegaCorp's design gains its resilience layer: a quarterly game day, and the ledger's live copy and locked backups in payments-dr.

MegaCorp's resilience layer: a game day in eu-west-2, and a pilot light with locked backups in payments-dr, eu-west-1AWS Account payments-prodRegion eu-west-2AWS Account payments-drRegion eu-west-1Vault Lock, compliance modeAWS FaultInjectionServiceAmazon Aurora(globaldatabase)Amazon Aurora(ledgersecondary)AWS Backup(ledgerbackups)123
1Each quarter, FIS fails over the ledger's writer to prove the tasks reconnect
2The global database keeps a secondary ledger in eu-west-1, typically under a second behind
3AWS Backup copies each night's ledger backup to a locked vault in payments-dr
Figure MegaCorp's resilience layer: a game day in eu-west-2, and a pilot light with locked backups in payments-dr, eu-west-1#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS Fault Injection Service
      • Amazon Aurora (global database)
  • AWS Account payments-dr
    • Region eu-west-1
      • Amazon Aurora (ledger secondary)
      • Vault Lock, compliance mode
        • AWS Backup (ledger backups)
  1. Each quarter, FIS fails over the ledger's writer to prove the tasks reconnect
  2. The global database keeps a secondary ledger in eu-west-1, typically under a second behind
  3. AWS Backup copies each night's ledger backup to a locked vault in payments-dr

Read more Service Quotas and CloudWatch alarms

Cost#

Part IV · Enterprise-grade · Chapter 31·5 min read

AWS bills by use: compute by the second or hour, storage by the gigabyte-month, and data by the gigabyte as it moves.

MegaCorp's first month on AWS brings a surprise line: NAT gateway data processing. The statements job writes every file to S3 through the NAT gateway, paying per gigabyte. Alex adds a gateway endpoint, which has no hourly or processing charge, and that traffic now moves for nothing.

The paths data travels#

Where data moves, and what it costs: billed between zones and between Regions, but free to S3 through a gateway endpointRegion eu-west-2VPC 10.20.0.0/16Availability Zone aAvailability Zone bRegion eu-west-1Amazon S3(statements)VPC endpoints(S3 gateway)AWS Fargate(statementsjob)Amazon Aurora(writer)Amazon S3(replica)1234
1Zone to zone: billed as it leaves and again as it arrives
2To S3 through a gateway endpoint: no charge to move it
3The endpoint routes straight to the bucket
4Region to Region: billed as it leaves eu-west-2
Figure Where data moves, and what it costs: billed between zones and between Regions, but free to S3 through a gateway endpoint#
As text
  • Region eu-west-2
    • Amazon S3 (statements)
    • VPC 10.20.0.0/16
      • VPC endpoints (S3 gateway)
      • Availability Zone a
        • AWS Fargate (statements job)
      • Availability Zone b
        • Amazon Aurora (writer)
  • Region eu-west-1
    • Amazon S3 (replica)
  1. Zone to zone: billed as it leaves and again as it arrives
  2. To S3 through a gateway endpoint: no charge to move it
  3. The endpoint routes straight to the bucket
  4. Region to Region: billed as it leaves eu-west-2

Data in from the internet is free, and so is data that stays in one zone. Data that crosses a zone, a Region or the edge of AWS is usually billed.

Gotcha

A NAT gateway charges twice. Every gigabyte through it pays a processing charge on top of the usual data transfer. Send S3 and DynamoDB traffic through gateway endpoints instead.

Read more Overview of data transfer costs for common architectures · Understanding data transfer charges

Pricing models#

On-Demand pays by the second; Spot sells spare capacity cheaply but can be reclaimed at two minutes' notice; a commitment cuts the rate for steady use, as chapter 17 showed.

Savings Plans
On AzureSavings plan for compute, Azure Reservations

What it does. Savings Plans lower the price of steady use in exchange for committing to spend an amount per hour for one or three years.

How it works. Compute Savings Plans cover EC2 in any family, size or Region, plus Fargate and Lambda. EC2 Instance plans save more on one family in one Region; Database plans cover Aurora, RDS, DynamoDB and more.

When to use it. Steady use, sized from Cost Explorer's recommendation and bought centrally.

When not to. Use that may shrink: a plan can't be cancelled or changed.

Limits. No capacity reservation, and no discount on Spot.

Cost. All, part or none paid upfront, at a rate fixed for the term.

Security. Only the management account decides which accounts share the discount.

Gotchas. Turning sharing off for an account can raise the bill.

Read more Savings Plans documentation · Savings Plans types · Reserved Instances and Savings Plans discount sharing

Reserved Instances, the older commitment to one instance configuration, still fill many bills; Savings Plans give similar discounts without exchanges.

Seeing and controlling spend#

Which tool answers the cost questionyesnoyesnoyesnoMust someone hear when spend passes aline you set?AWS BudgetsMust someone hear about spend nobodyplanned?AWS Cost Anomaly DetectionDo you need every line item, joinedwith your own data?A CUR 2.0 export to S3Explore and forecast in AWS CostExplorer
Figure Which tool answers the cost question#
As text
  1. Must someone hear when spend passes a line you set? Yes: AWS Budgets. No: the next step.
  2. Must someone hear about spend nobody planned? Yes: AWS Cost Anomaly Detection. No: the next step.
  3. Do you need every line item, joined with your own data? Yes: A CUR 2.0 export to S3. No: the next step.
  4. Explore and forecast in AWS Cost Explorer
AWS Cost Explorer
On AzureMicrosoft Cost Management

What it does. AWS Cost Explorer charts cost and usage by service, account, Region and tag, and forecasts ahead.

How it works. It refreshes at least daily, and recommends Savings Plans and Reserved Instances from your usage.

When to use it. Finding what drives the bill, and sizing a commitment.

When not to. Line-by-line analysis with your own data: export CUR 2.0 to S3.

Limits. as of September 2026: 13 months of history, and forecasts 18 months ahead.

Cost. The console is free; API requests are billed.

Security. Tags appear in cost reports, so keep secrets out of them.

Gotchas. Once enabled, Cost Explorer can't be turned off.

Read more AWS Cost Management documentation · Analyzing your costs with AWS Cost Explorer

AWS Budgets
On AzureMicrosoft Cost Management

What it does. AWS Budgets alerts when cost, usage or commitment coverage crosses a line you set, and can act.

How it works. A budget watches actual or forecast spend; alerts go by email or SNS, and actions apply an IAM policy or SCP, or stop EC2 and RDS instances.

When to use it. Every account, with a forecast alert; sandboxes, with an action.

When not to. A hard cap: it sees charges hours late.

Limits. Data updates up to three times a day.

Cost. Alerts are free, as are two action-enabled budgets a month.

Security. Actions run through a role you grant AWS Budgets.

Gotchas. A budget is visible only in the account that created it.

Read more AWS Cost Management documentation · Configuring budget actions · AWS Budgets pricing

A budget that acts: the sandbox's forecast crosses the line, developers hear, and new launches stopAWS BudgetsDevelopersSandbox accountthe month's forecast passes 90% of thebudget1an action applies an SCP that denies new instances2running instances need an action in thesandbox's own budget to stop3
Figure A budget that acts: the sandbox's forecast crosses the line, developers hear, and new launches stop#
As text
  1. AWS Budgets to Developers: the month's forecast passes 90% of the budget
  2. AWS Budgets to Sandbox account: an action applies an SCP that denies new instances
  3. Sandbox account: running instances need an action in the sandbox's own budget to stop

Cost Anomaly Detection flags unplanned spend with its likely cause; Data Exports sends CUR 2.0, every line of the bill, to S3.

Read more Detecting unusual spend with AWS Cost Anomaly Detection · What is AWS Data Exports?

Tags: who pays for what#

MegaCorp tags every resource with CostCenter and Owner. A tag reaches the bill only after the management account activates it, and earlier costs stay untagged unless backfilled, up to twelve months.

Keys are case sensitive, so CostCenter and costcenter split the bill. A tag policy standardizes them, though it ignores untagged resources; account tags cover everything in an account, even charges that cannot be tagged.

False friend: tag inheritance
On Azure

Cost Management can copy subscription and resource group tags onto the usage records of every resource beneath them.

On AWS

Only account tags flow down, to all usage in the account. Otherwise, tag each resource, and activate each key for cost allocation.

Read more Organizing and tracking costs using cost allocation tags · Backfill cost allocation tags · Tag policies · Using account tags for cost allocation · Group and allocate costs using tag inheritance

Infrastructure as code#

Part IV · Enterprise-grade · Chapter 32·5 min read

Every MegaCorp resource starts as reviewed code. On AWS that code is Terraform, with its state in S3, or CloudFormation, written as YAML or generated from Java by the AWS CDK.

MegaCorp's platform team already runs Terraform on Azure, so the accounts and networks stay in Terraform. The payments developers write Java all day, so they define their own buckets and queues with the CDK.

Choosing a tool for infrastructure as codeyesnoyesnoDo your teams already run Terraform, ormanage more than AWS?Terraform, with state in S3Do developers want loops, types andtests in Java?The AWS CDK, which writesCloudFormationCloudFormation templates, inYAML
Figure Choosing a tool for infrastructure as code#
As text
  1. Do your teams already run Terraform, or manage more than AWS? Yes: Terraform, with state in S3. No: the next step.
  2. Do developers want loops, types and tests in Java? Yes: The AWS CDK, which writes CloudFormation. No: the next step.
  3. CloudFormation templates, in YAML

Read more What is CloudFormation?

Terraform on AWS#

Terraform works on AWS as it does on Azure; the backend is what changes. State moves from a blob container to an S3 bucket, and locking to a lock file that S3 writes only if none exists.

Azure · hashicorp/azurerm 5.5.0azure/network/versions.tfbackend
backend "azurerm" {
  use_oidc             = true
  use_azuread_auth     = true
  storage_account_name = "megacorptfstate"
  container_name       = "tfstate"
  key                  = "payments/network.tfstate"
}
AWS · hashicorp/aws 6.64.0aws/network/versions.tfbackend
backend "s3" {
  bucket       = "megacorp-terraform-state"
  key          = "payments/network.tfstate"
  region       = "eu-west-2"
  use_lockfile = true
  encrypt      = true
}

Turn on versioning for the state bucket, so a bad write can be undone, and guard the state like a secret: it can hold values such as initial database passwords. Older configurations lock with DynamoDB, which is deprecated. The provider's default_tags block puts CostCenter and Owner on every resource that takes tags, except Auto Scaling groups.

Gotcha

Terraform removes the rule AWS adds. AWS gives a new security group an allow-all outbound rule, and the AWS provider deletes it, so the group sends nothing until you write rules. The payments load balancer needs one to reach its tasks on the listener and health check port.

Read more Terraform S3 backend · Terraform azurerm backend · Sensitive data in Terraform state · AWS provider: default_tags · Terraform: aws_security_group

CloudFormation#

AWS CloudFormation
On AzureAzure Resource Manager templates, Bicep

What it does. AWS CloudFormation creates, updates and deletes a set of AWS resources, called a stack, from a template.

How it works. The stack records what CloudFormation created. A change set previews an update, including any replacement, and nothing changes until you execute it; a failed deployment rolls back.

When to use it. AWS-only estates, StackSets that reach every account and Region, and everything the CDK generates.

When not to. Resources outside AWS, or teams already fluent in Terraform.

Limits. as of September 2026: 500 resources per template, and a 1 MB template body from S3.

Cost. No charge for AWS resource types; third-party types and custom hooks are billed per operation.

Security. Review each change set before executing it: it shows every deletion and replacement.

Gotchas. Deleting a stack deletes its resources, and fails on a bucket that still holds objects.

Read more AWS CloudFormation documentation · Update CloudFormation stacks using change sets · CloudFormation quotas

Resources:
  Statements:
    Type: AWS::S3::Bucket
    Properties:
      BucketEncryption:
        ServerSideEncryptionConfiguration:
          - ServerSideEncryptionByDefault:
              SSEAlgorithm: aws:kms
      PublicAccessBlockConfiguration:
        BlockPublicAcls: true
        BlockPublicPolicy: true
        IgnorePublicAcls: true
        RestrictPublicBuckets: true

Changes made by hand show up as drift: CloudFormation reports each changed property but leaves the fix to you. StackSets deploy one template to many accounts and Regions in a single operation.

Read more CloudFormation drift detection · AWS CloudFormation pricing

The AWS CDK, in Java#

AWS CDK
On Azureno direct equivalent

What it does. The AWS CDK defines infrastructure in Java, TypeScript, Python, C#, Go or JavaScript, and deploys it through CloudFormation.

How it works. Constructs, from single resources to whole patterns with secure defaults, make up stacks. The CDK synthesizes a CloudFormation template, then deploys it as a change set.

When to use it. Developers who want loops, types, tests and reuse in the language they already write.

When not to. Reviewers who must read every line that deploys: a few lines of Java can become hundreds of lines of template.

Limits. It deploys only through CloudFormation, so CloudFormation's quotas apply.

Cost. Open source; you pay for what it deploys.

Security. By default, a deployment that broadens permissions or security group rules waits for approval.

Gotchas. Each account and Region must be bootstrapped once before stacks with assets can deploy.

Read more AWS CDK documentation · What is the AWS CDK?

java/src/main/java/com/example/payments/infra/PaymentsStack.javastack
public PaymentsStack(final Construct scope, final String id, final StackProps props) {
    super(scope, id, props);

    Vpc.Builder.create(this, "Payments")
            .ipAddresses(IpAddresses.cidr("10.20.0.0/16"))
            .maxAzs(2)
            .natGateways(2)
            .build();

    Bucket.Builder.create(this, "Statements")
            .encryption(BucketEncryption.KMS_MANAGED)
            .blockPublicAccess(BlockPublicAccess.BLOCK_ALL)
            .enforceSsl(true)
            .build();
}

These lines define the payments VPC, across two zones with a NAT gateway in each, and the statements bucket, encrypted and closed to the public. One construct can expand into dozens of resources: AWS's own example of a VPC, cluster and service becomes more than 50.

From Java to resources: the CDK writes a template, and CloudFormation deploys it as a change setDeveloperAWS CDKAWS CloudFormationcdk deploy1synthesizes the Java app into aCloudFormation template2creates a change set, then executes it3builds the stack, and rolls back if aresource fails4
Figure From Java to resources: the CDK writes a template, and CloudFormation deploys it as a change set#
As text
  1. Developer to AWS CDK: cdk deploy
  2. AWS CDK: synthesizes the Java app into a CloudFormation template
  3. AWS CDK to AWS CloudFormation: creates a change set, then executes it
  4. AWS CloudFormation: builds the stack, and rolls back if a resource fails

How infrastructure as code grew#

Infrastructure as code on AWS and Azure: templates first, then one open-source tool for every cloud, then friendlier languagesAWS CloudFormation2011Terraform 0.1: AWSand DigitalOcean2014AWS CDK, generallyavailable2019Bicep 0.3, supportedin production2021
Figure Infrastructure as code on AWS and Azure: templates first, then one open-source tool for every cloud, then friendlier languages#
As text
  1. 2011: AWS CloudFormation
  2. 2014: Terraform 0.1: AWS and DigitalOcean
  3. 2019: AWS CDK, generally available
  4. 2021: Bicep 0.3, supported in production

Terraform's author built it after admiring CloudFormation in 2011, wanting the same idea for every cloud. On Azure, ARM's JSON templates gave way to Bicep, a shorter language that compiles to that JSON, much as the CDK's Java compiles to CloudFormation. Azure keeps Bicep's state, so there is no state file to guard.

resource statements 'Microsoft.Storage/storageAccounts@2023-05-01' = {
  name: 'megacorpstatements'
  location: resourceGroup().location
  sku: {
    name: 'Standard_LRS'
  }
  kind: 'StorageV2'
}

Read more What is Bicep? · Bicep frequently asked questions

Delivery pipelines#

Part IV · Enterprise-grade · Chapter 33·5 min read

A change reaches production in three moves: a pipeline builds and tests it without long-lived keys, a deployment exposes it a little at a time, and feature flags decide who sees it.

The payments service's first pipeline ran in GitHub Actions, trading a signed token for short-lived credentials, as chapter 10 showed. As more teams join, MegaCorp's platform team moves releases into a tooling account running CodePipeline, with GitHub still the source.

Read more Create a role for OpenID Connect federation

Pipelines across accounts#

One pipeline, many accounts: the pipeline assumes a deploy role in each account it releases to, and production waits for an approvalAWS Account toolingRegion eu-west-2Workload accountsAWS Account payments-devAWS Account payments-prodAWS CodeBuild(build andtest)AWSCodePipeline(payments)IAM role(deploy)IAM role(deploy)123
1Each commit is built and tested in CodeBuild
2The pipeline assumes the deploy role in payments-dev
3After an approval, it assumes the deploy role in payments-prod
Figure One pipeline, many accounts: the pipeline assumes a deploy role in each account it releases to, and production waits for an approval#
As text
  • AWS Account tooling
    • Region eu-west-2
      • AWS CodeBuild (build and test)
      • AWS CodePipeline (payments)
  • Workload accounts
    • AWS Account payments-dev
      • IAM role (deploy)
    • AWS Account payments-prod
      • IAM role (deploy)
  1. Each commit is built and tested in CodeBuild
  2. The pipeline assumes the deploy role in payments-dev
  3. After an approval, it assumes the deploy role in payments-prod
AWS CodePipeline
On AzureAzure Pipelines

What it does. AWS CodePipeline runs a release as stages of actions, such as source, build, approve and deploy, on every change.

How it works. A V2 pipeline starts on chosen branches, tags or paths, passes variables between stages, and can roll a stage back. Actions in other accounts run through roles in those accounts.

When to use it. Releases to many accounts from one tooling account, with no long-lived keys.

When not to. Teams already well served by GitHub Actions assuming roles through OIDC.

Limits. An artifact moves between accounts only through the pipeline's own account.

Cost. V2 bills per action execution minute, approvals free; builds and artifacts bill separately.

Security. Let each deploy role trust only the pipeline's service role.

Gotchas. Cross-account pipelines need a customer managed KMS key for artifacts; the default key won't work.

Read more AWS CodePipeline documentation · Create a pipeline that uses resources from another account · AWS CodePipeline pricing

AWS CodeBuild
On AzureAzure Pipelines

What it does. AWS CodeBuild compiles code, runs tests and produces artifacts, with no build servers to run.

How it works. Each build runs your commands in a prepackaged environment, with tools such as Maven and Gradle, or in your own image, and CodeBuild scales with the builds that arrive.

When to use it. Build and test actions in CodePipeline, and builds nobody wants to host.

When not to. Work that needs a long-running server: builds are time-limited.

Limits. as of September 2026: a build runs for at most 36 hours, and concurrent builds are capped per compute type, often at one by default.

Cost. Per build minute.

Security. Give each project its own IAM role, limited to what its build touches.

Gotchas. Builds beyond the concurrency quota fail; raise it in Service Quotas before a busy release.

Read more AWS CodeBuild documentation · Quotas for AWS CodeBuild · AWS CodeBuild pricing

Deployment strategies#

Choosing how a change reaches usersyesnoyesnoyesnoIs it a setting or a feature switch,not new code?An AppConfig rollout, 20%at a timeMust the old version keep running untilthe new one proves itself?Blue/green, with a baketimeCan a few users try it before everyone?Canary: 10% first, the restminutes laterAll at once, rolled back if it fails
Figure Choosing how a change reaches users#
As text
  1. Is it a setting or a feature switch, not new code? Yes: An AppConfig rollout, 20% at a time. No: the next step.
  2. Must the old version keep running until the new one proves itself? Yes: Blue/green, with a bake time. No: the next step.
  3. Can a few users try it before everyone? Yes: Canary: 10% first, the rest minutes later. No: the next step.
  4. All at once, rolled back if it fails

Since July 2025, Amazon ECS runs blue/green deployments itself, keeping the old revision ready until the new one proves itself.

ECS blue/green: the new revision starts beside the old, passes its tests, takes the traffic, and the old waits out a bake timeAmazon ECSBlue revisionGreen revisionstart its tasks beside the blue ones1test it through the lifecycle hooks2shift the production traffic to it3kept through the bake time, ready to takethe traffic back4
Figure ECS blue/green: the new revision starts beside the old, passes its tests, takes the traffic, and the old waits out a bake time#
As text
  1. Amazon ECS to Green revision: start its tasks beside the blue ones
  2. Amazon ECS to Green revision: test it through the lifecycle hooks
  3. Amazon ECS to Green revision: shift the production traffic to it
  4. Blue revision: kept through the bake time, ready to take the traffic back

For Lambda functions, and ECS services that use it, CodeDeploy shifts traffic in canary or linear steps and rolls back when a deployment fails or an alarm threshold is met.

A canary with a guardrail: a tenth of the traffic first, and an alarm sends it all backAWS CodeDeploySettlement functionCloudWatch alarmsend 10% of invocations to the newversion1errors cross the threshold within five minutes2roll back: redeploy the previous version3
Figure A canary with a guardrail: a tenth of the traffic first, and an alarm sends it all back#
As text
  1. AWS CodeDeploy to Settlement function: send 10% of invocations to the new version
  2. CloudWatch alarm to AWS CodeDeploy: errors cross the threshold within five minutes
  3. AWS CodeDeploy to Settlement function: roll back: redeploy the previous version
Gotcha

Blue/green needs room for two. ECS runs both revisions until the bake time ends, which can double the service's tasks, so capacity and quotas must allow both.

Read more Amazon ECS blue/green deployments · Deployment configurations in CodeDeploy · Redeploy and roll back a deployment with CodeDeploy

Configuration and feature flags#

The new instant-refund feature ships hidden behind a flag. The team turns it on for a fifth of its targets every six minutes, AWS's recommended strategy, and an alarm on refund errors would roll it back.

AWS AppConfig
On AzureAzure App Configuration

What it does. AWS AppConfig changes how an application behaves, through feature flags and settings, without a redeployment.

How it works. Each change is validated, then rolled out by a deployment strategy; the AppConfig Agent caches it beside the application, and an alarm during the rollout or bake time rolls it back.

When to use it. Feature flags, kill switches, allow and block lists, and tuning values.

When not to. Values fixed at build time, which belong in the code or the environment.

Limits. as of September 2026: AWS's recommended production strategy takes 30 minutes, then watches alarms for 30 more.

Cost. Per retrieval of configuration; the agent's cache keeps retrievals down.

Security. IAM decides who may change or deploy a flag, and CloudTrail records each change.

Gotchas. Rollback on alarms works only once AppConfig has permission to watch them.

Read more AWS AppConfig documentation · Predefined AppConfig deployment strategies

False friend: App Configuration
On Azure

Azure App Configuration stores settings and feature flags in one place; applications read them through client libraries and pick up changes without a restart.

On AWS

AWS AppConfig treats each change as a release: validated, rolled out to a share of targets at a time, and rolled back on a CloudWatch alarm.

Read more What is AWS AppConfig? · What is Azure App Configuration?

Operations#

Part IV · Enterprise-grade · Chapter 34·4 min read

Machines still need patches and, now and then, a person. Systems Manager provides both without an open port.

At three in the morning a disk fills on a ledger batch host. Once, an engineer reached it through a bastion on port 22; now Alex opens a session from the console, and no inbound port is open.

Which Systems Manager tool fits the jobyesnoyesnoyesnoDo you need a shell on one machine,now?Session ManagerMust one command run on many machines?Run CommandIs it a multi-step fix, with approvals,across accounts?An Automation runbookPatches on a schedule: a PatchManager policy
Figure Which Systems Manager tool fits the job#
As text
  1. Do you need a shell on one machine, now? Yes: Session Manager. No: the next step.
  2. Must one command run on many machines? Yes: Run Command. No: the next step.
  3. Is it a multi-step fix, with approvals, across accounts? Yes: An Automation runbook. No: the next step.
  4. Patches on a schedule: a Patch Manager policy
AWS Systems Manager
On AzureAzure Automation, Azure Update Manager, Azure Arc

What it does. AWS Systems Manager operates fleets of machines, on EC2, on premises and in other clouds, without logging in to each one.

How it works. SSM Agent on each machine connects to the service; tools built on it open sessions, run commands, patch and automate fixes, across accounts and Regions.

When to use it. Any fleet of instances or servers: access, patching and routine fixes.

When not to. Major operating system upgrades, which Patch Manager does not perform.

Limits. as of September 2026: 100 automations run at once per account, with up to 5,000 queued.

Cost. Session Manager, Patch Manager and Run Command cost nothing extra on EC2; Automation bills per step.

Security. IAM decides who may open a session to which machines, and sessions can be logged to S3 or CloudWatch Logs.

Gotchas. A machine counts as managed only when its agent can reach the service, through a route out or VPC endpoints.

Read more AWS Systems Manager documentation · What is AWS Systems Manager? · AWS Systems Manager pricing

Access without open ports#

A session without SSH: IAM decides, the agent connects out, and the session is loggedAlexSession ManagerBatch hostS3 bucketstart a session to thebatch host1checks Alex's IAM permissions and thesession settings2SSM Agent opens thetwo-way channel3the session's log4
Figure A session without SSH: IAM decides, the agent connects out, and the session is logged#
As text
  1. Alex to Session Manager: start a session to the batch host
  2. Session Manager: checks Alex's IAM permissions and the session settings
  3. Session Manager to Batch host: SSM Agent opens the two-way channel
  4. Session Manager to S3 bucket: the session's log

Sessions that forward a port or carry SSH are not logged, because Session Manager only tunnels them. Hosts without public addresses reach the service through VPC endpoints.

Read more AWS Systems Manager Session Manager

Patching#

One patch policy covers every account and Region in MegaCorp's organization. It scans daily and installs weekly, against a patch baseline that approves security updates a few days after release.

A patch policy: scan often, install in a window, and report compliancePatch policyPayments hostsS3 bucketscan every day against the baseline1install approved patches in the weeklywindow2compliance report, as CSV3
Figure A patch policy: scan often, install in a window, and report compliance#
As text
  1. Patch policy to Payments hosts: scan every day against the baseline
  2. Patch policy to Payments hosts: install approved patches in the weekly window
  3. Patch policy to S3 bucket: compliance report, as CSV
Gotcha

Compliant is not the same as secure. Patch Manager measures each machine against your baseline, not against every known flaw. AWS does not test patches first, and the predefined baselines are examples, not recommendations.

Read more AWS Systems Manager Patch Manager

Runbooks#

When the batch disk fills again, the fix becomes an Automation runbook: steps that snapshot the volume, grow it, and extend the file system, started by an EventBridge rule on the disk alarm.

False friend: runbook
On Azure

An Azure Automation runbook is a PowerShell or Python script, run in an Azure sandbox or on a Hybrid Runbook Worker.

On AWS

A Systems Manager runbook is a YAML or JSON document of steps, which can run scripts, wait for approval, and run across accounts and Regions with rate controls.

Read more AWS Systems Manager Automation · Azure Automation runbook types

Enterprise-grade
  1. An auditor asks who decrypted last month's settlement files. Where do you look?
    Answer
    CloudTrail, which records every use of the customer managed payments key, with the caller.
  2. The board wants payments back within an hour, losing at most a minute. Which strategy, and what must you check in the recovery Region?
    Answer
    Pilot light: data live there, compute deployed on failover. Check that its quotas allow production capacity.
  3. Someone deletes ledger rows by mistake. Why doesn't the global database save you?
    Answer
    Replication copies the deletion within a second. A point-in-time backup, in a locked vault in another account, brings the rows back.
  4. An engineer needs a shell on a host in a private subnet. How, without opening port 22?
    Answer
    Session Manager: IAM grants the session, SSM Agent connects out through VPC endpoints, and the session is logged.

The design method#

Part V · Designing · Chapter 35·3 min read

Every design in this book answers the same thirteen questions. The Well-Architected pillars say why one answer beats another, and a decision record keeps the reasons after the people who chose them move on.

When MegaCorp's payments team started on AWS, each choice was argued in a meeting and forgotten by the next one. Now every design walks the same thirteen questions: six decided once for the company, four for each workload, and three asked of every design.

QuestionWhat it decidesCovered in
1 Accountshow workloads split into accounts, and how the accounts are organizedchapter 7, chapter 15
2 Sign-inhow people, pipelines and customers prove who they arechapter 8, chapter 9, chapter 10
3 NetworksRegions, address plans, VPCs and how they connectchapter 11, chapter 12, chapter 13, chapter 14
4 Guardrailswhat no account may do, and how posture is checkedchapter 15, chapter 28
5 Logs and alertswhat is recorded, and who is wokenchapter 28, chapter 29
6 Deliveryhow code and infrastructure reach productionchapter 32, chapter 33
7 Where it runsinstances, containers or functionschapter 16, chapter 17, chapter 18, chapter 19
8 Datawhere each kind of data liveschapter 20, chapter 21, chapter 22, chapter 25, chapter 26
9 How parts talkqueues, topics, events and streamschapter 23
10 Traffic inload balancers, APIs, the edge and certificateschapter 24, chapter 27
11 Failurezones, Regions, backups and quotaschapter 30
12 Costwhat is billed, and who payschapter 31
13 Attackkeys, secrets, threats and operator accesschapter 27, chapter 28, chapter 34

The company answers questions 1 to 6 once, and every workload inherits them. Each workload then answers 7 to 10 with the decision figures from Part III, and tests its answers against 11 to 13. The three designs that follow each show the result the same way: the design, its decisions, how it fails, what drives its cost, and how it resists attack.

Read more AWS Well-Architected Framework

The Well-Architected pillars#

AWS's Well-Architected Framework names six qualities that designs trade against each other: operational excellence, security, reliability, performance efficiency, cost optimization and sustainability. Its questions start a conversation about decisions; AWS says a review is not an audit.

The Well-Architected Framework grew from four pillars to six, and gained a tool for reviewsThe framework, withfour pillars2015AWS Well-ArchitectedTool, at no charge2018Sustainability, thesixth pillar2021
Figure The Well-Architected Framework grew from four pillars to six, and gained a tool for reviews#
As text
  1. 2015: The framework, with four pillars
  2. 2018: AWS Well-Architected Tool, at no charge
  3. 2021: Sustainability, the sixth pillar

The AWS Well-Architected Tool walks a workload through those questions, records the decisions and recommends improvements. Custom lenses add a company's own questions, such as MegaCorp's rules for card data.

False friend: Well-Architected Framework
On Azure

Microsoft's framework has five pillars: reliability, security, cost optimization, operational excellence and performance efficiency, assessed through the Azure Well-Architected Review.

On AWS

AWS's has six, adding sustainability, and reviews run in the Well-Architected Tool, with custom lenses for a company's own questions.

Read more The pillars of the framework · What is AWS Well-Architected Tool? · Azure Well-Architected Framework pillars

Decision records#

An architecture decision record, or ADR, gives a decision's context, the options, the choice and its consequences, with a status: proposed, accepted, rejected or superseded. MegaCorp keeps its records in the payments repository, beside the code they govern, and reviewers cite them in code reviews.

When a decision earns a recordyesnoyesnoyesnoDoes it shape the structure, such asaccounts, Regions or the data store?Write an ADRDoes it trade one pillar againstanother?Write an ADR that names thetrade-offWould it be hard or costly to undo?Write an ADR that lists therejected optionsDecide in the code review
Figure When a decision earns a record#
As text
  1. Does it shape the structure, such as accounts, Regions or the data store? Yes: Write an ADR. No: the next step.
  2. Does it trade one pillar against another? Yes: Write an ADR that names the trade-off. No: the next step.
  3. Would it be hard or costly to undo? Yes: Write an ADR that lists the rejected options. No: the next step.
  4. Decide in the code review
A record's life: proposed, reviewed and accepted, and later superseded rather than editedAlexPayments teamDecision logADR-007, proposed: the ledger in eu-west-2, restored from backups1accepted after review, with a version andstakeholders2ADR-019 supersedes ADR-007: a pilot light in eu-west-13
Figure A record's life: proposed, reviewed and accepted, and later superseded rather than edited#
As text
  1. Alex to Decision log: ADR-007, proposed: the ledger in eu-west-2, restored from backups
  2. Payments team to Decision log: accepted after review, with a version and stakeholders
  3. Alex to Decision log: ADR-019 supersedes ADR-007: a pilot light in eu-west-1

An accepted record is never edited. When the board set a one-hour recovery time, Alex wrote ADR-019 and marked ADR-007 superseded, so the log shows both what changed and why.

Gotcha

Unrecorded decisions come back. Without the reasons written down, the next team reopens the debate, or quietly reverses a choice that had a good reason.

Read more Architectural decision record process · Maintain an architecture decision record

A web platform#

Part V · Designing · Chapter 36·5 min read

MegaCorp's merchant portal and payments API, designed end to end: the questions answered, the choices recorded, and the design tested against failure, cost and attack.

Merchants sign in to a portal to see payments and refunds, and their systems call the payments API. Questions 1 to 6 were answered once for MegaCorp: the payments-prod account, IAM Identity Center and OIDC, a VPC on the shared hub, SCPs, the organization trail and CodePipeline. This chapter answers the rest for the web platform.

On Azure
The merchant portal and payments API, on each cloud, with a web application firewall at the front door (on Azure)Resource group payments-prodVirtual network 10.20.0.0/16Subnet appsDataMerchantsAzure FrontDoor(portal)Azure ContainerApps(payments)Azure Databasefor PostgreSQL(ledger)Azure Cache forRedis(settings)123
1Merchants reach the portal and the API through Front Door
2Front Door forwards requests to the container app
3The app writes the ledger
On AWS
The merchant portal and payments API, on each cloud, with a web application firewall at the front door (on AWS)Region eu-west-2VPC 10.20.0.0/16DataMerchantsAmazonCloudFront(portal)ApplicationLoad Balancer(payments)AWS Fargate(payments)Amazon Aurora(ledger)AmazonElastiCache(settings)1234
1Merchants reach the portal and the API through CloudFront
2CloudFront forwards dynamic requests to the load balancer
3The load balancer picks a healthy task
4The tasks write the ledger
Figure The merchant portal and payments API, on each cloud, with a web application firewall at the front door#
As text

On Azure:

  • Merchants
  • Resource group payments-prod
    • Azure Front Door (portal)
    • Virtual network 10.20.0.0/16
      • Subnet apps
        • Azure Container Apps (payments)
    • Data
      • Azure Database for PostgreSQL (ledger)
      • Azure Cache for Redis (settings)
  1. Merchants reach the portal and the API through Front Door
  2. Front Door forwards requests to the container app
  3. The app writes the ledger

On AWS:

  • Merchants
  • Amazon CloudFront (portal)
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Application Load Balancer (payments)
      • AWS Fargate (payments)
      • Data
        • Amazon Aurora (ledger)
        • Amazon ElastiCache (settings)
  1. Merchants reach the portal and the API through CloudFront
  2. CloudFront forwards dynamic requests to the load balancer
  3. The load balancer picks a healthy task
  4. The tasks write the ledger

The decisions#

DecisionChosenRejectedReason
Where it runsECS on Fargate, in two zonesEC2 Auto Scaling; LambdaJava containers the team already builds, steady traffic, and no instances to patch
Front doorCloudFront with AWS WAFthe load balancer aloneCaches the portal near merchants and blocks attacks before they reach the Region
LedgerAurora PostgreSQL, a writer and a readerDynamoDB; RDSTransactions across tables, and a reader in the other zone ready to take over
SettingsElastiCache, read before AuroraAurora for every readSettings change rarely and are read on every request
Duplicate webhooksDynamoDB conditional writesa unique index in AuroraAbsorbs webhook bursts without adding load to the ledger
Merchant sign-ina Cognito user poola users table of its ownHosted sign-in and tokens, with no passwords in MegaCorp's database

Each row is an ADR in the payments repository. The rejected options stay in the records, so the next team can see what was weighed.

Read more Architectural decision record process

When a zone fails#

Zone a fails: the load balancer sends requests only to the tasks in zone b, and Aurora promotes the reader there to writerRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPrivate subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24ApplicationLoad Balancer(payments)AWS Fargate(payments)Amazon Aurora(writer, lost)AWS Fargate(payments)Amazon Aurora(reader,promoted)Failed12
1The load balancer sends requests only to healthy tasks, now all in zone b
2The cluster endpoint now names the promoted writer in zone b
Figure Zone a fails: the load balancer sends requests only to the tasks in zone b, and Aurora promotes the reader there to writer#
As text
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Application Load Balancer (payments)
      • Availability Zone a (failed)
        • Private subnet 10.20.10.0/24
          • AWS Fargate (payments)
          • Amazon Aurora (writer, lost)
      • Availability Zone b
        • Private subnet 10.20.11.0/24
          • AWS Fargate (payments)
          • Amazon Aurora (reader, promoted)
  1. The load balancer sends requests only to healthy tasks, now all in zone b
  2. The cluster endpoint now names the promoted writer in zone b

Each zone runs enough tasks to carry the whole load, so nothing has to launch during the failure. Aurora's volume already spans three zones, so the promoted reader has every committed write. The settings cache fails over to its replica, and the tasks in zone b keep using their own NAT gateway.

Gotcha

A cache can turn a failure into an overload. If the settings cache fails and every request falls through to Aurora, the database meets a load it never carried. AWS calls this bimodal behaviour. Size Aurora for it, or keep serving the last settings the task read.

Read more REL11-BP05 Use static stability

What drives the cost#

DriverGrows withLever
Fargate tasksvCPU, memory and storage, per secondSize tasks from real use, and cover the steady base with a Compute Savings Plan
Aurorainstance hours, storage and I/OCompare standard and I/O-Optimized clusters, which bill I/O differently
CloudFrontdata out and requestsCache the portal's static files; transfer from AWS origins is free
Zone to zonegigabytes between zones, billed both waysAccept it as the price of two zones, and keep chatty calls within one
NAT gatewayshours in each zone, and gigabytes processedSend S3 and DynamoDB traffic through gateway endpoints

Read more Understanding data transfer charges

How it resists attack#

The security view: each hop has its own guard, from the web ACL at the edge to the key that encrypts the ledger's credentialsRegion eu-west-2Sign-in and secretsVPC 10.20.0.0/16Security group albSecurity group appSecurity group ledgerMerchantsAWS WAF(web ACL)AmazonCloudFront(portal)Amazon Cognito(merchants)AWS KMS(payments key)AWS SecretsManager(ledger)ApplicationLoad Balancer(payments)AWS Fargate(payments)Amazon Aurora(ledger)1234567
1The web ACL blocks common attacks and floods of requests
2CloudFront serves the requests the web ACL allows
3Dynamic requests go to the load balancer
4The app's group admits only the load balancer's group
5The ledger's group admits only the app's group
6Tasks read the ledger's credentials when they start
7The credentials are encrypted under the payments key
Figure The security view: each hop has its own guard, from the web ACL at the edge to the key that encrypts the ledger's credentials#
As text
  • Merchants
  • AWS WAF (web ACL)
  • Amazon CloudFront (portal)
  • Region eu-west-2
    • Sign-in and secrets
      • Amazon Cognito (merchants)
      • AWS KMS (payments key)
      • AWS Secrets Manager (ledger)
    • VPC 10.20.0.0/16
      • Security group alb
        • Application Load Balancer (payments)
      • Security group app
        • AWS Fargate (payments)
      • Security group ledger
        • Amazon Aurora (ledger)
  1. The web ACL blocks common attacks and floods of requests
  2. CloudFront serves the requests the web ACL allows
  3. Dynamic requests go to the load balancer
  4. The app's group admits only the load balancer's group
  5. The ledger's group admits only the app's group
  6. Tasks read the ledger's credentials when they start
  7. The credentials are encrypted under the payments key

Every hop checks for itself. A request that passes the web ACL still needs a valid token from the Cognito user pool. A task needs the right group to reach the ledger, and the ledger's password lives in Secrets Manager, never in the image. GuardDuty and Security Hub watch the account from the audit account.

Read more Security groups

An event-driven system#

Part V · Designing · Chapter 37·5 min read

Card-scheme webhooks become ledger entries and events that other teams react to, without any team calling another directly. The same five views test the design.

Card schemes call MegaCorp whenever a payment clears, fails or is disputed, sometimes thousands of times a minute and sometimes twice for the same event. Settlement must post each one to the ledger exactly once, and fraud, statements and merchant alerts each need to hear about it.

On Azure
Webhooks to ledger entries and events, on each cloud (on Azure)Resource group payments-prodProcessingCard schemesAzure APIManagement(webhooks)Azure ServiceBus(webhooks)Azure Cosmos DB(idempotencykeys)Azure Functions(settlement)Azure EventGrid(payments)12345
1Card schemes post webhooks to the API
2The API puts each one on a queue
3A function settles each message
4It records the event ID, once
5It publishes a payment event
On AWS
Webhooks to ledger entries and events, on each cloud (on AWS)Region eu-west-2ProcessingCard schemesAmazon APIGateway(webhooks)Amazon SQS(webhooks)Amazon DynamoDB(idempotencykeys)AWS Lambda(settlement)AmazonEventBridge(payments)12345
1Card schemes post webhooks to the API
2API Gateway sends each one straight to SQS
3Lambda hands the function batches of messages
4A conditional write records the event ID, once
5The function puts a payment event on the bus
Figure Webhooks to ledger entries and events, on each cloud#
As text

On Azure:

  • Card schemes
  • Resource group payments-prod
    • Azure API Management (webhooks)
    • Azure Service Bus (webhooks)
    • Processing
      • Azure Cosmos DB (idempotency keys)
      • Azure Functions (settlement)
      • Azure Event Grid (payments)
  1. Card schemes post webhooks to the API
  2. The API puts each one on a queue
  3. A function settles each message
  4. It records the event ID, once
  5. It publishes a payment event

On AWS:

  • Card schemes
  • Region eu-west-2
    • Amazon API Gateway (webhooks)
    • Amazon SQS (webhooks)
    • Processing
      • Amazon DynamoDB (idempotency keys)
      • AWS Lambda (settlement)
      • Amazon EventBridge (payments)
  1. Card schemes post webhooks to the API
  2. API Gateway sends each one straight to SQS
  3. Lambda hands the function batches of messages
  4. A conditional write records the event ID, once
  5. The function puts a payment event on the bus

The decisions#

DecisionChosenRejectedReason
Intakea REST API sending straight to SQS, behind AWS WAFan HTTP API; an intake functionHTTP APIs cannot take a web ACL, and an intake function adds code and cost for nothing
Bufferan SQS standard queue, with a dead-letter queuea FIFO queueCard schemes need no order, and the consumer already ignores duplicates
SettlementLambda, with reserved concurrencyan ECS serviceArrivals come in bursts, and the reserved limit protects the ledger
Duplicatesa DynamoDB conditional write per event ID, with a TTLa unique index in AuroraKeeps the duplicate check off the ledger, and old keys expire by themselves
Fan-outan EventBridge bus, with a rule per consumeran SNS topicMany teams, routing on content, and consumers in other accounts
Card eventsKinesis Data StreamsEventBridge aloneFraud models replay an ordered stream of card events

Read more Choose an API Gateway API integration type

When the ledger is down#

The ledger is down: messages wait in the queue, failed ones return after their visibility timeout, and those that keep failing move to the dead-letter queueRegion eu-west-2Amazon SQS(dead letters)Amazon SQS(webhooks)AWS Lambda(settlement)Amazon Aurora(ledger)123
1Lambda reads batches; a failed message becomes visible again after the timeout
2Every write fails while the ledger is down
3After the maximum number of receives, a message moves to the dead-letter queue
Figure The ledger is down: messages wait in the queue, failed ones return after their visibility timeout, and those that keep failing move to the dead-letter queue#
As text
  • Region eu-west-2
    • Amazon SQS (dead letters)
    • Amazon SQS (webhooks)
    • AWS Lambda (settlement)
    • Amazon Aurora (ledger) (failed)
  1. Lambda reads batches; a failed message becomes visible again after the timeout
  2. Every write fails while the ledger is down
  3. After the maximum number of receives, a message moves to the dead-letter queue

Nothing is lost while the ledger is down: the queue keeps messages for four days by default, and up to fourteen. The function reports only the messages that failed, so the rest are not retried, and its reserved concurrency sets the pace when the ledger returns. An alarm on the dead-letter queue's depth tells the team which webhooks need a look.

Gotcha

Every hop can deliver twice. SQS standard queues, EventBridge and Lambda all deliver at least once, so the idempotency keys are not optional: without them, a retry posts a payment twice.

Read more Using dead-letter queues in Amazon SQS · Handling errors for an SQS event source in Lambda

What drives the cost#

DriverGrows withLever
SQSrequests, each 64 KB counting as oneSend and receive in batches of up to 10 messages
Lambdarequests and GB-secondsTune memory to the work; SnapStart starts Java quickly at no extra cost
EventBridgecustom events, per millionEvents from AWS services on the default bus are free
KinesisGB written and read on demand, or shard-hoursSwitch to provisioned shards once the volume is steady
DynamoDBrequests on demand, and storageLet TTL delete old keys, which uses no write capacity

Read more Amazon SQS pricing · AWS Lambda pricing

How it resists attack#

The security view: a web ACL in front, a role that may only send to one queue, keys for data at rest, and a bus policy for the fraud accountRegion eu-west-2IntakeProcessingAWS Account fraudCard schemesAWS WAF(web ACL)Amazon APIGateway(webhooks)Amazon SQS(webhooks)AWS KMS(payments key)AWS Lambda(settlement)AmazonEventBridge(payments)AmazonEventBridge(fraud)1234567
1The web ACL filters webhooks before API Gateway sees them
2Allowed requests reach the API
3API Gateway's role may send only to this queue
4Messages are encrypted at rest under the payments key
5The function's role may only receive from this queue
6and put events only on this bus
7A rule sends card events to the fraud account, whose bus policy admits them
Figure The security view: a web ACL in front, a role that may only send to one queue, keys for data at rest, and a bus policy for the fraud account#
As text
  • Card schemes
  • AWS WAF (web ACL)
  • Region eu-west-2
    • Intake
      • Amazon API Gateway (webhooks)
      • Amazon SQS (webhooks)
      • AWS KMS (payments key)
    • Processing
      • AWS Lambda (settlement)
      • Amazon EventBridge (payments)
  • AWS Account fraud
    • Amazon EventBridge (fraud)
  1. The web ACL filters webhooks before API Gateway sees them
  2. Allowed requests reach the API
  3. API Gateway's role may send only to this queue
  4. Messages are encrypted at rest under the payments key
  5. The function's role may only receive from this queue
  6. and put events only on this bus
  7. A rule sends card events to the fraud account, whose bus policy admits them

Each service holds only the permissions its step needs, and each resource names who may use it: the queue policy, the key policy and the fraud account's bus policy each admit one caller.

Read more Amazon EventBridge quotas

Hybrid and multi-Region#

Part V · Designing · Chapter 38·4 min read

The payments platform in two Regions, joined to MegaCorp's data centre: one private link that reaches both, a hub in each Region, and a runbook that moves everything to eu-west-1.

The mainframe in MegaCorp's data centre still receives settlement files, and the board's one-hour recovery objective put a pilot light in eu-west-1. The network must reach both Regions from the data centre, keep working when eu-west-2 is lost, and keep development away from production.

On Azure
The data centre joined to two regions, on each cloud (on Azure)UK SouthUK WestAzureExpressRoute(circuit)Virtual networkgateway(hub gateway)Azure Firewall(hub firewall)Virtual networkgateway(hub gateway)Azure Firewall(hub firewall)12
1The circuit connects to the hub in UK South
2and to the hub in UK West
On AWS
The data centre joined to two regions, on each cloud (on AWS)Region eu-west-2Region eu-west-1Customergateway(data centre)AWS DirectConnect(gateway)AWS NetworkFirewall(inspection)AWS TransitGateway(hub)AWS TransitGateway(hub)1234
1Direct Connect, with a VPN as backup, reaches the gateway
2The gateway associates with the hub in eu-west-2
3and with the hub in eu-west-1
4The hubs peer, with static routes
Figure The data centre joined to two regions, on each cloud#
As text

On Azure:

  • Azure ExpressRoute (circuit)
  • UK South
    • Virtual network gateway (hub gateway)
    • Azure Firewall (hub firewall)
  • UK West
    • Virtual network gateway (hub gateway)
    • Azure Firewall (hub firewall)
  1. The circuit connects to the hub in UK South
  2. and to the hub in UK West

On AWS:

  • Customer gateway (data centre)
  • AWS Direct Connect (gateway)
  • Region eu-west-2
    • AWS Network Firewall (inspection)
    • AWS Transit Gateway (hub)
  • Region eu-west-1
    • AWS Transit Gateway (hub)
  1. Direct Connect, with a VPN as backup, reaches the gateway
  2. The gateway associates with the hub in eu-west-2
  3. and with the hub in eu-west-1
  4. The hubs peer, with static routes

The decisions#

DecisionChosenRejectedReason
Link to the data centreDirect Connect, with Site-to-Site VPN as backupa VPN aloneSettlement files are large and steady, and the VPN keeps a path if the link fails
Reaching two Regionsone Direct Connect gateway, associated with each Region's huba connection per RegionThe gateway is global, so one link reaches both Regions
Hubsa transit gateway in each Region, peeredVPC peering between every pairRoute tables keep development from production, and peering joins the Regions
InspectionNetwork Firewall in each hubsecurity groups aloneAuditors want data centre traffic inspected and outbound traffic kept to known domains
Hybrid DNSResolver endpoints and shared forwarding rulescopying data centre names into Route 53Each side answers for its own names, and nothing is copied by hand
Recoverya pilot light in eu-west-1, failed over with ARCa warm standbyMeets the one-hour objective for less, as ADR-019 records

Read more Direct Connect gateways · Transit gateway peering attachments

When eu-west-2 is lost#

eu-west-2 is lost: the data centre still reaches eu-west-1 through the global Direct Connect gateway, and the pilot light takes overRegion eu-west-2Region eu-west-1Customergateway(data centre)AWS DirectConnect(gateway)AWS TransitGateway(hub)Amazon Aurora(ledger writer)AWS TransitGateway(hub)Amazon Aurora(ledger,promoted)Failed123
1The link from the data centre ends on the global gateway, not in a Region
2Routes to eu-west-1 keep working
3Settlement files reach the recovery VPC, where the promoted ledger takes writes
Figure eu-west-2 is lost: the data centre still reaches eu-west-1 through the global Direct Connect gateway, and the pilot light takes over#
As text
  • Customer gateway (data centre)
  • AWS Direct Connect (gateway)
  • Region eu-west-2 (failed)
    • AWS Transit Gateway (hub)
    • Amazon Aurora (ledger writer)
  • Region eu-west-1
    • AWS Transit Gateway (hub)
    • Amazon Aurora (ledger, promoted)
  1. The link from the data centre ends on the global gateway, not in a Region
  2. Routes to eu-west-1 keep working
  3. Settlement files reach the recovery VPC, where the promoted ledger takes writes

A Direct Connect gateway is global and sits outside the data path, so losing eu-west-2 does not take it down. The on-call engineer runs the runbook from chapter 30: promote the ledger, deploy the tasks, and flip the ARC routing control.

Gotcha

A backup in the failed Region fails with it. A Site-to-Site VPN ends on one Region's transit gateway, so the VPN that backs up Direct Connect in eu-west-2 is gone with it. Give eu-west-1's hub a VPN of its own.

Read more What is AWS Site-to-Site VPN? · AWS Direct Connect Resiliency Toolkit

What drives the cost#

DriverGrows withLever
Direct Connectport hours, and data out of AWSBuy the port for the steady volume, and keep the VPN for failures
Transit gatewaysattachments per hour, and gigabytes sent into eachAttach only the VPCs that need the data centre or each other
Between Regionsgigabytes replicated, billed as they leave the source RegionReplicate the ledger and backups, not logs and caches
Network Firewallendpoints per zone per hour, and gigabytes inspectedInspect traffic to and from the data centre, not between trusted VPCs
Pilot lightthe secondary ledger, running all the timeKeep compute undeployed until a failover

Read more AWS Transit Gateway pricing · AWS Network Firewall pricing

How it resists attack#

The security view: the private link is encrypted, the hub inspects what crosses it, and route tables keep development away from productionAWS Account networkRegion eu-west-2VPC inspectionWorkloadsAWS Account payments-prodRegion eu-west-2VPC 10.20.0.0/16AWS Account payments-devRegion eu-west-2VPC 10.30.0.0/16Customergateway(data centre)AWS TransitGateway(hub)AWS NetworkFirewall(firewall)Transit gatewayattachment(attachment)Transit gatewayattachment(attachment)1234
1MACsec on the Direct Connect link, or a VPN over it, encrypts the traffic
2The hub sends data centre traffic through the firewall
3Production's route table reaches the data centre
4Development's route table has no route to production
Figure The security view: the private link is encrypted, the hub inspects what crosses it, and route tables keep development away from production#
As text
  • Customer gateway (data centre)
  • AWS Account network
    • Region eu-west-2
      • AWS Transit Gateway (hub)
      • VPC inspection
        • AWS Network Firewall (firewall)
  • Workloads
    • AWS Account payments-prod
      • Region eu-west-2
        • VPC 10.20.0.0/16
          • Transit gateway attachment (attachment)
    • AWS Account payments-dev
      • Region eu-west-2
        • VPC 10.30.0.0/16
          • Transit gateway attachment (attachment)
  1. MACsec on the Direct Connect link, or a VPN over it, encrypts the traffic
  2. The hub sends data centre traffic through the firewall
  3. Production's route table reaches the data centre
  4. Development's route table has no route to production

Direct Connect is not encrypted by default, while traffic between the two Regions' hubs is encrypted by AWS as it leaves a Region. VPCs on hubs that share one Direct Connect gateway could reach each other, so blackhole routes close that path.

Read more Encryption in AWS Direct Connect · What is AWS Network Firewall?

Design review drills#

Part V · Designing · Chapter 39·3 min read

Four designs, each with a flaw this book has taught you to see. Find it before you read the answer at the end of the chapter.

A review walks a design through the thirteen questions and asks, at every box, what happens when it fails, what it costs, and who could misuse it. Each drill below came from a real kind of mistake, and each fails at least one of those questions.

Read more AWS Well-Architected Framework

Drill 1: the way out#

Drill 1: the payments tasks in both zones reach the internet through one NAT gatewayRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24Private subnet 10.20.10.0/24Availability Zone bPrivate subnet 10.20.11.0/24NAT gatewayAWS Fargate(payments)AWS Fargate(payments)12
1Zone a's tasks send outbound traffic to the NAT gateway
2Zone b's tasks use the same NAT gateway
Figure Drill 1: the payments tasks in both zones reach the internet through one NAT gateway#
As text
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • NAT gateway
        • Private subnet 10.20.10.0/24
          • AWS Fargate (payments)
      • Availability Zone b
        • Private subnet 10.20.11.0/24
          • AWS Fargate (payments)
  1. Zone a's tasks send outbound traffic to the NAT gateway
  2. Zone b's tasks use the same NAT gateway

Read more NAT gateway basics · Regional NAT gateways

Drill 2: the ledger's door#

Drill 2: the ledger sits in a public subnet, and its security group allows port 5432 from anywhereRegion eu-west-2VPC 10.20.0.0/16Availability Zone aPublic subnet 10.20.0.0/24InternetgatewayAmazon Aurora(ledger)1
1Any address on the internet can reach port 5432
Figure Drill 2: the ledger sits in a public subnet, and its security group allows port 5432 from anywhere#
As text
  • Region eu-west-2
    • VPC 10.20.0.0/16
      • Internet gateway
      • Availability Zone a
        • Public subnet 10.20.0.0/24
          • Amazon Aurora (ledger)
  1. Any address on the internet can reach port 5432

Read more Security groups · Subnets for your VPC

Drill 3: the backups#

Drill 3: nightly backups are copied to another Region, in the same account, to a vault without a lockAWS Account payments-prodRegion eu-west-2Region eu-west-1AWS Backup(backup plan)AWS Backup(vault,unlocked)1
1Each night's ledger backup is copied to eu-west-1
Figure Drill 3: nightly backups are copied to another Region, in the same account, to a vault without a lock#
As text
  • AWS Account payments-prod
    • Region eu-west-2
      • AWS Backup (backup plan)
    • Region eu-west-1
      • AWS Backup (vault, unlocked)
  1. Each night's ledger backup is copied to eu-west-1

Read more AWS Backup Vault Lock

Drill 4: the consumer#

Drill 4: the settlement function takes batches from the queue and posts each message to the ledgerRegion eu-west-2Amazon SQS(webhooks)AWS Lambda(settlement)Amazon Aurora(ledger)12
1Batches of ten; any error throws, failing the whole batch
2Each message posts a payment, with no check for repeats
Figure Drill 4: the settlement function takes batches from the queue and posts each message to the ledger#
As text
  • Region eu-west-2
    • Amazon SQS (webhooks)
    • AWS Lambda (settlement)
    • Amazon Aurora (ledger)
  1. Batches of ten; any error throws, failing the whole batch
  2. Each message posts a payment, with no check for repeats
Gotcha

A diagram shows only what someone drew. In a review, ask for the route tables, security groups, key policies and quotas behind each box; the flaws in these drills all hide there.

Read more Using dead-letter queues in Amazon SQS · Handling errors for an SQS event source in Lambda

Designing
  1. Drill 1: what fails, and what does it cost?
    Answer
    This NAT gateway lives in zone a. If zone a fails, zone b loses its way out too, and zone b's traffic crosses zones, billed both ways. Give each zone its own NAT gateway and route table, or use a regional NAT gateway.
  2. Drill 2: what would you change?
    Answer
    Move the ledger to private subnets, and let its security group admit only the payments tasks' group. Nothing on the internet should reach a database port.
  3. Drill 3: why is this not enough?
    Answer
    Anyone who can administer payments-prod can delete both copies. Copy to a vault in another account, locked in compliance mode, so not even the root user can delete a backup early.
  4. Drill 4: what goes wrong under load?
    Answer
    Delivery is at least once and a thrown error returns the whole batch, so payments post twice. Report only the failed messages, check an idempotency key with a DynamoDB conditional write, and add a dead-letter queue.

Service map#

Appendices · Lookup · Chapter 40·4 min read

Every AWS service this book names, with its Azure counterparts, in both directions. Start from a service you know and follow it to the figure that shows its AWS match.

The first table runs from AWS to Azure, grouped by category; the second runs from Azure to AWS. "none" means Azure has no single counterpart, and the service's profile or note says how Azure does the job. A service shown in a figure links to the first one, and every row links to its documentation.

From AWS to Azure

AWS serviceOn AzureCategoryWhere
AWS GlueAzure Data Factory, Microsoft FabricAnalyticsshown here · docs
AWS Lake FormationMicrosoft PurviewAnalyticsshown here · docs
Amazon AthenaMicrosoft Fabric, Azure Synapse AnalyticsAnalyticsshown here · docs
Amazon Kinesis Data StreamsAzure Event HubsAnalyticsshown here · docs
Amazon MSKAzure Event HubsAnalyticsdocs
Amazon RedshiftMicrosoft Fabric, Azure Synapse AnalyticsAnalyticsshown here · docs
AWS Step FunctionsAzure Logic Apps, Durable FunctionsApplication integrationshown here · docs
Amazon EventBridgeAzure Event GridApplication integrationshown here · docs
Amazon MQAzure Service BusApplication integrationdocs
Amazon SNSAzure Service Bus, Azure Event GridApplication integrationshown here · docs
Amazon SQSAzure Service Bus, Azure Queue StorageApplication integrationshown here · docs
Amazon BedrockMicrosoft FoundryArtificial intelligenceshown here · docs
Amazon SESAzure Communication ServicesBusiness applicationsdocs
AWS BudgetsMicrosoft Cost ManagementCloud financial managementshown here · docs
AWS Cost ExplorerMicrosoft Cost ManagementCloud financial managementshown here · docs
Savings PlansSavings plan for compute, Azure ReservationsCloud financial managementshown here · docs
AWS App RunnerAzure App Service, Azure Container AppsComputedocs
AWS Elastic BeanstalkAzure App ServiceComputeshown here · docs
AWS LambdaAzure FunctionsComputeshown here · docs
Amazon EC2Azure Virtual MachinesComputeshown here · docs
Amazon EC2 Auto ScalingAzure Virtual Machine Scale SetsComputeshown here · docs
Amazon LightsailnoneComputedocs
AWS FargateAzure Container Apps, Azure Container InstancesContainersshown here · docs
Amazon ECRAzure Container RegistryContainersshown here · docs
Amazon ECSAzure Container AppsContainersshown here · docs
Amazon EKSAzure Kubernetes ServiceContainersshown here · docs
Amazon AuroraAzure Database for PostgreSQL, Azure SQL Database HyperscaleDatabasesshown here · docs
Amazon DocumentDBAzure Cosmos DB for MongoDBDatabasesdocs
Amazon DynamoDBAzure Cosmos DBDatabasesshown here · docs
Amazon ElastiCacheAzure Managed Redis, Azure Cache for RedisDatabasesshown here · docs
Amazon RDSAzure SQL Database, Azure Database for PostgreSQL, Azure Database for MySQLDatabasesshown here · docs
AWS CDKnoneDeveloper toolsshown here · docs
AWS CloudFormationAzure Resource Manager templates, BicepDeveloper toolsshown here · docs
AWS CodeBuildAzure PipelinesDeveloper toolsshown here · docs
AWS CodeDeployAzure PipelinesDeveloper toolsdocs
AWS CodePipelineAzure PipelinesDeveloper toolsshown here · docs
AWS AmplifyAzure Static Web AppsFront-end web and mobiledocs
AWS AppConfigAzure App ConfigurationManagement and governanceshown here · docs
AWS CloudTrailAzure MonitorManagement and governanceshown here · docs
AWS ConfigAzure Policy, Azure Resource GraphManagement and governanceshown here · docs
AWS Control TowerAzure landing zoneManagement and governanceshown here · docs
AWS Fault Injection ServiceAzure Chaos StudioManagement and governanceshown here · docs
AWS OrganizationsAzure management groupsManagement and governanceshown here · docs
AWS Systems ManagerAzure Automation, Azure Update Manager, Azure ArcManagement and governanceshown here · docs
AWS Well-Architected ToolAzure Well-Architected ReviewManagement and governancedocs
AWS X-RayApplication InsightsManagement and governanceshown here · docs
Amazon CloudWatchAzure MonitorManagement and governanceshown here · docs
Service QuotasnoneManagement and governancedocs
AWS Application Migration ServiceAzure MigrateMigration and modernizationdocs
AWS DataSyncAzure Storage MoverMigration and modernizationdocs
AWS Database Migration ServiceAzure Database Migration ServiceMigration and modernizationdocs
AWS Transfer FamilySFTP support for Azure Blob StorageMigration and modernizationdocs
AWS Cloud WANAzure Virtual WANNetworking and content deliverydocs
AWS Direct ConnectAzure ExpressRouteNetworking and content deliveryshown here · docs
AWS Global AcceleratorAzure Front Door, Azure Load BalancerNetworking and content deliveryshown here · docs
AWS Network FirewallAzure FirewallNetworking and content deliveryshown here · docs
AWS PrivateLinkAzure Private LinkNetworking and content deliveryshown here · docs
AWS Site-to-Site VPNAzure VPN GatewayNetworking and content deliveryshown here · docs
AWS Transit GatewayAzure Virtual WAN, Azure virtual network peeringNetworking and content deliveryshown here · docs
Amazon API GatewayAzure API ManagementNetworking and content deliveryshown here · docs
Amazon Application Recovery ControllernoneNetworking and content deliveryshown here · docs
Amazon CloudFrontAzure Front DoorNetworking and content deliveryshown here · docs
Amazon Route 53Azure DNS, Azure Traffic ManagerNetworking and content deliveryshown here · docs
Amazon VPCAzure Virtual NetworkNetworking and content deliveryshown here · docs
Elastic Load BalancingAzure Load Balancer, Azure Application GatewayNetworking and content deliveryshown here · docs
AWS Certificate ManagerAzure Key VaultSecurity, identity and complianceshown here · docs
AWS IAMAzure role-based access control, Microsoft Entra IDSecurity, identity and complianceshown here · docs
AWS IAM Identity CenterMicrosoft Entra IDSecurity, identity and complianceshown here · docs
AWS KMSAzure Key VaultSecurity, identity and complianceshown here · docs
AWS Resource Access ManagernoneSecurity, identity and complianceshown here · docs
AWS Secrets ManagerAzure Key VaultSecurity, identity and complianceshown here · docs
AWS Security HubMicrosoft Defender for CloudSecurity, identity and complianceshown here · docs
AWS Security Token ServiceMicrosoft Entra IDSecurity, identity and complianceshown here · docs
AWS ShieldAzure DDoS ProtectionSecurity, identity and compliancedocs
AWS WAFAzure Web Application FirewallSecurity, identity and complianceshown here · docs
Amazon CognitoMicrosoft Entra External IDSecurity, identity and complianceshown here · docs
Amazon GuardDutyMicrosoft Defender for CloudSecurity, identity and complianceshown here · docs
Amazon InspectorMicrosoft Defender for CloudSecurity, identity and compliancedocs
AWS BackupAzure BackupStorageshown here · docs
AWS Elastic Disaster RecoveryAzure Site RecoveryStoragedocs
Amazon EBSAzure managed disksStorageshown here · docs
Amazon EFSAzure FilesStorageshown here · docs
Amazon FSxAzure Files, Azure NetApp FilesStorageshown here · docs
Amazon S3Azure Blob StorageStorageshown here · docs

From Azure to AWS

On AzureAWS serviceWhere
Application InsightsAWS X-Rayshown here
Azure API ManagementAmazon API Gatewayshown here
Azure App ConfigurationAWS AppConfigshown here
Azure App ServiceAWS App Runner
Azure App ServiceAWS Elastic Beanstalkshown here
Azure Application GatewayElastic Load Balancingshown here
Azure ArcAWS Systems Managershown here
Azure AutomationAWS Systems Managershown here
Azure BackupAWS Backupshown here
Azure Blob StorageAmazon S3shown here
Azure Cache for RedisAmazon ElastiCacheshown here
Azure Chaos StudioAWS Fault Injection Serviceshown here
Azure Communication ServicesAmazon SES
Azure Container AppsAWS App Runner
Azure Container AppsAWS Fargateshown here
Azure Container AppsAmazon ECSshown here
Azure Container InstancesAWS Fargateshown here
Azure Container RegistryAmazon ECRshown here
Azure Cosmos DBAmazon DynamoDBshown here
Azure Cosmos DB for MongoDBAmazon DocumentDB
Azure Data FactoryAWS Glueshown here
Azure Database for MySQLAmazon RDSshown here
Azure Database for PostgreSQLAmazon Aurorashown here
Azure Database for PostgreSQLAmazon RDSshown here
Azure Database Migration ServiceAWS Database Migration Service
Azure DDoS ProtectionAWS Shield
Azure DNSAmazon Route 53shown here
Azure Event GridAmazon EventBridgeshown here
Azure Event GridAmazon SNSshown here
Azure Event HubsAmazon Kinesis Data Streamsshown here
Azure Event HubsAmazon MSK
Azure ExpressRouteAWS Direct Connectshown here
Azure FilesAmazon EFSshown here
Azure FilesAmazon FSxshown here
Azure FirewallAWS Network Firewallshown here
Azure Front DoorAWS Global Acceleratorshown here
Azure Front DoorAmazon CloudFrontshown here
Azure FunctionsAWS Lambdashown here
Azure Key VaultAWS Certificate Managershown here
Azure Key VaultAWS KMSshown here
Azure Key VaultAWS Secrets Managershown here
Azure Kubernetes ServiceAmazon EKSshown here
Azure landing zoneAWS Control Towershown here
Azure Load BalancerAWS Global Acceleratorshown here
Azure Load BalancerElastic Load Balancingshown here
Azure Logic AppsAWS Step Functionsshown here
Azure managed disksAmazon EBSshown here
Azure Managed RedisAmazon ElastiCacheshown here
Azure management groupsAWS Organizationsshown here
Azure MigrateAWS Application Migration Service
Azure MonitorAWS CloudTrailshown here
Azure MonitorAmazon CloudWatchshown here
Azure NetApp FilesAmazon FSxshown here
Azure PipelinesAWS CodeBuildshown here
Azure PipelinesAWS CodeDeploy
Azure PipelinesAWS CodePipelineshown here
Azure PolicyAWS Configshown here
Azure Private LinkAWS PrivateLinkshown here
Azure Queue StorageAmazon SQSshown here
Azure ReservationsSavings Plansshown here
Azure Resource GraphAWS Configshown here
Azure Resource Manager templatesAWS CloudFormationshown here
Azure role-based access controlAWS IAMshown here
Azure Service BusAmazon MQ
Azure Service BusAmazon SNSshown here
Azure Service BusAmazon SQSshown here
Azure Site RecoveryAWS Elastic Disaster Recovery
Azure SQL DatabaseAmazon RDSshown here
Azure SQL Database HyperscaleAmazon Aurorashown here
Azure Static Web AppsAWS Amplify
Azure Storage MoverAWS DataSync
Azure Synapse AnalyticsAmazon Athenashown here
Azure Synapse AnalyticsAmazon Redshiftshown here
Azure Traffic ManagerAmazon Route 53shown here
Azure Update ManagerAWS Systems Managershown here
Azure Virtual Machine Scale SetsAmazon EC2 Auto Scalingshown here
Azure Virtual MachinesAmazon EC2shown here
Azure Virtual NetworkAmazon VPCshown here
Azure virtual network peeringAWS Transit Gatewayshown here
Azure Virtual WANAWS Cloud WAN
Azure Virtual WANAWS Transit Gatewayshown here
Azure VPN GatewayAWS Site-to-Site VPNshown here
Azure Web Application FirewallAWS WAFshown here
Azure Well-Architected ReviewAWS Well-Architected Tool
BicepAWS CloudFormationshown here
Durable FunctionsAWS Step Functionsshown here
Microsoft Cost ManagementAWS Budgetsshown here
Microsoft Cost ManagementAWS Cost Explorershown here
Microsoft Defender for CloudAWS Security Hubshown here
Microsoft Defender for CloudAmazon GuardDutyshown here
Microsoft Defender for CloudAmazon Inspector
Microsoft Entra External IDAmazon Cognitoshown here
Microsoft Entra IDAWS IAMshown here
Microsoft Entra IDAWS IAM Identity Centershown here
Microsoft Entra IDAWS Security Token Serviceshown here
Microsoft FabricAWS Glueshown here
Microsoft FabricAmazon Athenashown here
Microsoft FabricAmazon Redshiftshown here
Microsoft FoundryAmazon Bedrockshown here
Microsoft PurviewAWS Lake Formationshown here
Savings plan for computeSavings Plansshown here
SFTP support for Azure Blob StorageAWS Transfer Family

CLI side by side#

Appendices · Lookup · Chapter 41·1 min read

The Azure CLI and the AWS CLI do the same daily jobs with different words. Both query their JSON output with JMESPath, so most of what you know about --query carries over.

Signing in, and choosing where you work#

TaskAzure CLIAWS CLI
Set up sign-in once<code>az login</code><code>aws configure sso</code>, which writes a profile to the <code>config</code> file
Sign in<code>az login</code><code>aws sso login --profile my-dev-profile</code>
Who am I?<code>az account show</code><code>aws sts get-caller-identity</code>
What can I choose from?<code>az account list</code><code>aws configure list-profiles</code>
Switch where I work<code>az account set --subscription "My Demos"</code><code>export AWS_PROFILE=my-dev-profile</code>, or <code>--profile</code> on each command
Change a settingper command, such as <code>--subscription</code><code>aws configure set region eu-west-2 --profile my-dev-profile</code>
Sign out<code>az logout</code><code>aws sso logout</code>

An AWS profile names one account, one role and a default Region, so switching accounts means switching profiles. While the Identity Center sign-in lasts, the CLI renews the role's credentials by itself.

$ aws sso login --profile my-dev-profile
SSO authorization page has automatically been opened in your default browser.
Follow the instructions in the browser to complete this authorization request.
Successfully logged into Start URL: https://my-sso-portal.awsapps.com/start

Read more Configuring IAM Identity Center authentication with the AWS CLI · Configuration and credential file settings in the AWS CLI · Manage Azure subscriptions with the Azure CLI

Querying output#

Both CLIs take --query with a JMESPath expression and --output table for people. The AWS CLI also passes server-side filters, such as --filters, which each API defines; they cut what comes back before --query shapes it.

az vm list --resource-group QueryDemo \
  --query "[?storageProfile.osDisk.osType=='Linux'].{Name:name, admin:osProfile.adminUsername}" \
  --output table

aws ec2 describe-volumes --query 'Volumes[?Size < `20`].VolumeId'
OutputAs documented in Filtering output in the AWS CLI
[
  "vol-2e410a47",
  "vol-a1b3c7nd"
]
Gotcha

Text output queries each page. With --output text, the AWS CLI splits the results into pages first and runs --query on each, so a query for the first match returns one per page. Use JSON output when the query must see everything at once.

Read more Filtering output in the AWS CLI · Query Azure CLI command results

IAM policy cookbook#

Appendices · Lookup · Chapter 42·5 min read

Five policies the payments team writes again and again, quoted from the book's Terraform: what each allows, why it is shaped that way, and what to change before you reuse it.

When Alex needs a new permission, the platform team starts from one of these recipes rather than a blank page. Each is an aws_iam_policy_document data source, which Terraform renders as the JSON policy that IAM stores. Its statements have the shape shown in chapter 9: an effect, actions, resources and optional conditions. The account and key IDs are placeholders.

RecipeKind of policyAttached toSee
Read and write one bucketidentity-basedthe settlement job's rolechapter 20
Consume one queueidentity-basedthe settlement function's execution rolechapter 23
Let a pipeline deploytrust policythe deploy rolechapter 10
Keep every account in two Regionsservice control policythe Workloads OUchapter 7
Let tags decideidentity-basedevery team's rolechapter 9

Read and write one bucket#

The settlement job writes each day's statements to megacorp-statements and reads them back to reconcile. The bucket encrypts objects with the payments team's own KMS key, so the role needs the key as well as the bucket.

aws/iam/main.tfs3-read-write
# The settlement job writes statements to one bucket and reads them back.
data "aws_iam_policy_document" "statements_read_write" {
  statement {
    sid       = "ListTheBucket"
    actions   = ["s3:ListBucket"]
    resources = ["arn:aws:s3:::megacorp-statements"]
  }

  statement {
    sid       = "ReadAndWriteObjects"
    actions   = ["s3:GetObject", "s3:PutObject"]
    resources = ["arn:aws:s3:::megacorp-statements/*"]
  }

  statement {
    sid       = "UseTheKeyOnlyThroughS3"
    actions   = ["kms:GenerateDataKey", "kms:Decrypt"]
    resources = ["arn:aws:kms:eu-west-2:111111111111:key/1234abcd-12ab-34cd-56ef-1234567890ab"]

    condition {
      test     = "StringEquals"
      variable = "kms:ViaService"
      values   = ["s3.eu-west-2.amazonaws.com"]
    }
  }
}

Listing is an action on the bucket, so its resource is the bucket's ARN. Reading and writing are actions on objects, so theirs is the bucket's ARN followed by /*. IAM's own example allows s3:*Object, which matches every action whose name ends in Object, including DeleteObject. Naming the two actions keeps a bug in the job from deleting statements.

Writing an object encrypted with a KMS key needs kms:GenerateDataKey, and reading one needs kms:Decrypt. The kms:ViaService condition lets the role use the key only for requests that come through S3 in London, so the job cannot take the key and call KMS itself. To reuse the recipe, change the bucket, the key's ARN and the Region in the condition.

Gotcha

The bucket and its objects are different resources. Put s3:ListBucket on the object ARN, or the object actions on the bucket ARN, and the statement matches nothing: the job is refused although the policy looks right.

Read more IAM example: read and write objects in one bucket · Using server-side encryption with AWS KMS keys (SSE-KMS) · kms:ViaService condition key

Consume one queue#

Card-scheme webhooks land on the webhooks queue, and the settlement function settles each one, as in chapter 19. Lambda polls the queue for the function with the function's execution role, so that role needs the right to receive, delete and inspect messages.

aws/iam/main.tfsqs-consumer
# The settlement function's execution role: Lambda polls one queue on its behalf.
data "aws_iam_policy_document" "settlement_consumer" {
  statement {
    sid = "ConsumeOneQueue"
    actions = [
      "sqs:ReceiveMessage",
      "sqs:DeleteMessage",
      "sqs:GetQueueAttributes",
    ]
    resources = ["arn:aws:sqs:eu-west-2:111111111111:webhooks"]
  }

  statement {
    sid       = "ReadEncryptedMessages"
    actions   = ["kms:Decrypt"]
    resources = ["arn:aws:kms:eu-west-2:111111111111:key/1234abcd-12ab-34cd-56ef-1234567890ab"]
  }
}

These are the three SQS actions in the AWS managed policy AWSLambdaSQSQueueExecutionRole. That policy also allows three logs: actions so the function can write its logs; add them too, limited to the function's log group. A queue encrypted with a customer managed key also needs kms:Decrypt, as here. The function and the queue must be in the same Region.

Gotcha

The managed policy reaches every queue. AWSLambdaSQSQueueExecutionRole allows its actions on every resource, so a function given it can read and delete messages from any queue in the account. Start with it if you must, then replace it with a policy that names the queue.

Read more AWSLambdaSQSQueueExecutionRole managed policy

Let a pipeline deploy#

The payments pipeline runs in GitHub Actions and deploys with no stored keys, as in chapter 10. This is a trust policy: rather than what the role may do, it says who may assume it. Here, that is a workflow on the main branch of one repository, vouched for by the account's OIDC provider for GitHub.

aws/iam/main.tfgithub-oidc-trust
# The deploy role's trust policy: only the main branch of one repository may assume it.
data "aws_iam_policy_document" "deploy_trust" {
  statement {
    actions = ["sts:AssumeRoleWithWebIdentity"]

    principals {
      type        = "Federated"
      identifiers = ["arn:aws:iam::111111111111:oidc-provider/token.actions.githubusercontent.com"]
    }

    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:aud"
      values   = ["sts.amazonaws.com"]
    }

    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:sub"
      values   = ["repo:megacorp/payments:ref:refs/heads/main"]
    }
  }
}

The aud condition checks that the token was issued for AWS STS, and sub names the repository and branch. For GitHub, IAM requires a sub condition that is present and is more than a wildcard. To reuse the recipe, change the account, the repository and the branch.

Gotcha

A broad sub hands the role to others. A sub that names no organization or repository lets workflows in repositories you don't control assume the role. One that names only the organization lets any of its repositories deploy. Name one repository and the branch that deploys.

Read more Create a role for OpenID Connect federation

Keep every account in two Regions#

MegaCorp keeps payment data in London, with its recovery copy in Ireland, as in chapter 30. The platform team makes that a guardrail: a service control policy on the Workloads OU refuses any request made to another Region.

aws/iam/main.tfdeny-outside-regions
# A service control policy: every account works in London and Ireland, except global services.
data "aws_iam_policy_document" "deny_outside_regions" {
  statement {
    sid    = "DenyOutsideLondonAndIreland"
    effect = "Deny"
    not_actions = [
      "cloudfront:*",
      "iam:*",
      "organizations:*",
      "route53:*",
      "support:*",
    ]
    resources = ["*"]

    condition {
      test     = "StringNotEquals"
      variable = "aws:RequestedRegion"
      values   = ["eu-west-2", "eu-west-1"]
    }
  }
}

aws:RequestedRegion names the Region a request was made to. A few popular global services, among them IAM, CloudFront and Route 53, have a single endpoint in us-east-1, so a plain Region deny would block them; not_actions lists them as exceptions. Like every service control policy, this one grants nothing, so each role still needs its own permissions.

Gotcha

A deny with not_actions lists exceptions, not permissions. The statement denies every other action outside the two Regions. When a team starts using another global service, add it to the list, or its calls are refused everywhere. Attach the policy to a test OU first: a wrong guardrail stops every account below it at once.

Read more IAM example: deny access based on the requested Region

Let tags decide#

Each team keeps its database passwords in Secrets Manager. Each team's role carries a team tag, and so does each secret. chapter 9 showed the idea; this is the policy, one document attached to every team's role.

aws/iam/main.tfabac-team-tag
# One policy for every team's role: a role reads the secrets tagged with its own team.
data "aws_iam_policy_document" "own_team_secrets" {
  statement {
    sid       = "ReadOwnTeamSecrets"
    actions   = ["secretsmanager:DescribeSecret", "secretsmanager:GetSecretValue"]
    resources = ["*"]

    condition {
      test     = "StringEquals"
      variable = "aws:ResourceTag/team"
      values   = ["&{aws:PrincipalTag/team}"]
    }
  }

  statement {
    sid       = "KeepTheTeamTag"
    effect    = "Deny"
    actions   = ["secretsmanager:UntagResource"]
    resources = ["*"]

    condition {
      test     = "ForAnyValue:StringEquals"
      variable = "aws:TagKeys"
      values   = ["team"]
    }
  }
}

The first statement lets a role read a secret only when the secret's team tag equals the role's own. A new secret tagged team=ledger is readable by the ledger role at once, with no policy change. The second statement stops a team role from removing that tag, because the tag is now a permission.

Gotcha

Terraform writes IAM's policy variables with &{. IAM spells the caller's tag ${aws:PrincipalTag/team}, which Terraform would take for its own interpolation. Write &{aws:PrincipalTag/team} in aws_iam_policy_document, and IAM receives this condition:

"Condition": {
  "StringEquals": {
    "aws:ResourceTag/team": "${aws:PrincipalTag/team}"
  }
}

A tag carries one value, so a person who works for two teams needs two roles. Actions that don't act on one secret, such as listing secrets, ignore resource tags and need a statement of their own. A broader policy on the same role, such as AdministratorAccess, is not narrowed by these conditions.

Read more IAM tutorial: permissions based on tags · aws_iam_policy_document data source

Before a policy ships#

Validate each new policy with IAM Access Analyzer, which runs more than 100 checks on it, and let an external access analyzer watch for anything shared outside the organization. When a request is still refused, the message often names the kind of policy that refused it: look it up in chapter 43.

Read more Using IAM Access Analyzer · Policy evaluation logic

Error index#

Appendices · Lookup · Chapter 43·3 min read

Messages you will meet in your first months on AWS, each quoted from the documentation, with what it means and how to fix it. Paste a message into search to land on its entry.

Account IDs, names and ARNs in these messages are the documentation's examples; yours will show your own. The rest of each message is as AWS writes it.

Access denied#

An access denied message names the caller, the action and, often, the kind of policy that refused. When several kinds refuse, it names only one, so check the others too.

User: arn:aws:iam::123456789012:role/HR is not authorized to perform: codecommit:ListRepositories
because no identity-based policy allows the codecommit:ListRepositories action
Means
Nothing denied the call, but no policy attached to the role allowed it either: an implicit deny.
Fix
Add an Allow for the action, and the resource, to a policy attached to the role or its permission set.
User: arn:aws:iam::123456789012:user/John is not authorized to perform: codecommit:ListRepositories
with an explicit deny in a service control policy
Means
A service control policy in the organization denies the action for this account. No policy inside the account can override it.
Fix
Ask whoever owns the organization's policies; the guardrail may be deliberate, such as a Region restriction.
User: arn:aws:iam::123456789012:user/John is not authorized to perform: codedeploy:ListDeployments
on resource: arn:aws:codedeploy:us-east-1:123456789012:deploymentgroup:*
because no permissions boundary allows the codedeploy:ListDeployments action
Means
The caller's policies may allow the action, but its permissions boundary, the most it may ever be allowed, does not.
Fix
Allow the action in the boundary as well, or use a role whose boundary already covers it.
User: arn:aws:iam::123456789012:user/John is not authorized to perform: sts:AssumeRole
because no role trust policy allows the sts:AssumeRole action
Means
The role's trust policy does not name this caller, so it cannot assume the role.
Fix
Add the caller to the role's trust policy. From another account, the caller's own policies must also allow sts:AssumeRole on the role.

Credentials and the CLI#

An error occurred (InvalidClientTokenId) when calling the ListBuckets operation: The security token
included in the request is invalid.
Means
AWS does not recognize the credentials the CLI sent: often the wrong profile, old keys, or credentials read from an unexpected place.
Fix
Run aws configure list to see which credentials and profile are in use, then sign in again with aws sso login.
An error occurred (SignatureDoesNotMatch) when calling the ListBuckets operation: The request
signature we calculated does not match the signature you provided. Check your key and signing
method.
Means
The signed request didn't check out, most often because the computer's clock is out of step, or a key was mangled by special characters.
Fix
Synchronize the clock; if that isn't it, create a new key pair.
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed
Means
The CLI does not trust the certificate it was shown, usually a corporate proxy's.
Fix
Point the CLI at the company's CA bundle with the ca_bundle setting, --ca-bundle or AWS_CA_BUNDLE.

Lambda#

Task timed out after 3.00 seconds
Means
The function ran past its configured timeout; with a short timeout, even initialization can use it all.
Fix
Raise the timeout, add memory, which also adds CPU, or make initialization faster.
User: arn:aws:iam::123456789012:user/developer is not authorized to perform: lambda:InvokeFunction
on resource: my-function
Means
The caller may not invoke the function.
Fix
Allow lambda:InvokeFunction on the function. Another service or account needs permission in the function's resource-based policy.
KMSDisabledException: Lambda was unable to decrypt the environment variables because the KMS key
used is disabled. Please check the function's KMS key settings.
Means
The key that encrypts the function's environment variables is disabled, or Lambda's grant on it was revoked.
Fix
Enable the key, or configure the function with another key so that Lambda recreates the grant.

Containers#

API error (500): Get https://111122223333.dkr.ecr.us-east-1.amazonaws.com/v2/: net/http: request
canceled while waiting for connection
Means
The task has no route to the registry.
Fix
In a private subnet, give the task a way out through a NAT gateway, or through VPC endpoints for ECR; in a public subnet, give it a public IP address.
The task can’t pull the image. Check that the role has the permissions to pull images from the
registry.
Means
The image exists, but the role that pulls it may not.
Fix
On Fargate, give the task execution role permission to pull from ECR; on EC2, the container instance role.
CannotPullContainerError: pull image manifest has been retried 5 time(s): failed to resolve
ref
Means
The image named in the task definition isn't in the repository, often a tag that has moved or gone.
Fix
Match the task definition to an image that exists, and pin a version rather than :latest.

EC2#

InsufficientInstanceCapacity
Means
AWS does not have enough On-Demand capacity of that type, in that zone, right now.
Fix
Try again in a few minutes, launch fewer instances at once, leave the zone unspecified, or choose another instance type.
InstanceLimitExceeded
Means
The account has reached its instance quota in this Region.
Fix
Request an increase in Service Quotas, for this Region.
You are not authorized to perform this operation.
Means
A launch is missing a permission, usually ec2:RunInstances or iam:PassRole for the instance's role.
Fix
Add the missing permission; the encoded authorization message in the full error can be decoded to name it.

The AWS you'll inherit#

Appendices · Lookup · Chapter 44·2 min read

An existing estate carries the AWS of the years it was built in. Each row names something you may find, what it tells you about its age, and what AWS offers today, so you can tell history from a mistake.

Old is not broken: most of what follows still works, and some of it will run for years. Replace it when you are changing that part anyway, or when a date below forces your hand.

What you findWhat it tells youTodaySee
EC2-Classic in old runbooks or scriptsan account from before 4 December 2013; EC2-Classic was retired, and marked deprecated on 31 July 2023VPCschapter 12
A default VPC in every Regionan account created after 4 December 2013VPCs you plan yourselfchapter 12
Classic Load Balancersbuilt before the Application Load Balancer arrived, on 11 August 2016Application or Network Load Balancerschapter 24
Auto Scaling launch configurationsbuilt before launch templates; accounts created since 1 October 2024 cannot create themlaunch templateschapter 17
gp2 EBS volumescreated before gp3 arrived, on 1 December 2020gp3, changed in place with Elastic Volumeschapter 20
An Origin Access Identity on a CloudFront distributionset up before Origin Access Control arrived, on 25 August 2022Origin Access Controlchapter 24
S3 buckets with ACLs, or open to the publiccreated before new buckets got Block Public Access and ACLs off, on 28 April 2023Block Public Access, with ACLs offchapter 20
CloudWatch Events ruleswritten after 14 January 2016; the service became Amazon EventBridge in 2019the same rules, in EventBridgechapter 23
AWS Single Sign-Onset up between 7 December 2017 and its rename on 26 July 2022IAM Identity Center, the same servicechapter 10
Reserved Instancesbought before Savings Plans arrived, on 6 November 2019Savings Planschapter 31
The AWS SDK for Java 1.xcode from before its end of support, on 31 December 2025the AWS SDK for Java 2.xchapter 19
AWS CDK v1 appsbuilt before v1's support ended, on 1 June 2023AWS CDK v2chapter 32
Amazon Linux 2images from before its end of life, on 30 June 2026Amazon Linux 2023chapter 17
Terraform state locked with DynamoDBan older S3 backend; DynamoDB locking is deprecatedthe S3 lock file, <code>use_lockfile</code>chapter 32
ECS blue/green through CodeDeploybuilt before ECS added its own, on 17 July 2025ECS built-in blue/greenchapter 33
AWS App Runner servicesa service no longer open to new customersAmazon ECS Express Modechapter 16
AWS Audit Managera service no longer open to new customersSecurity Hub's standards and Config's ruleschapter 28
ARC readiness checksnot offered to new customers from 30 April 2026ARC Region switch planschapter 30
The Systems Manager CloudWatch dashboardunavailable after 30 April 2026CloudWatch dashboardschapter 29

To date something the table doesn't cover, look it up in the service's document history, which lists each change by date, or in AWS What's New. AWS Config shows how a resource's configuration has changed, and CloudTrail's event history shows the last 90 days of calls.

Read more Restrict access to an Amazon S3 origin · Blocking public access to your Amazon S3 storage · Auto Scaling launch configurations · Amazon EBS General Purpose SSD volumes · End of support for the AWS SDK for Java 1.x

Glossary#

Appendices · Lookup · Chapter 45·1 min read

The general ideas this book explains in primers, in alphabetical order. Each term links to its primer, where the idea is explained with an example.

TermMeaning
Availability ZoneAn Availability Zone is one or more data centres with their own power, networking and connectivity. The zones in a Region sit far enough apart, up to about 100 km, not to fail together, and close enough for synchronous replication within a few milliseconds. Every Region has three or more, as of September 2026.
CIDR blockA CIDR block writes an address range as a base address and a prefix length. 10.20.0.0/16 fixes the first 16 of 32 bits and leaves 16 free: 65,536 addresses. Each extra bit of prefix halves the range, so a /24 holds 256. Ranges that share any address overlap.

Credits and sources#

Appendices · Lookup · Chapter 46·1 min read

Where the icons, the facts and the code come from, and whose names they are.

Icons#

The architecture figures use the official AWS Architecture Icons, package Icon-package_07312026, and the Azure architecture icons, set Azure_Public_Service_Icons_V24, each under its vendor's terms. The icons appear as their packages draw them: never recoloured, rotated or cropped. A failed component is shown by an outline and a label, not by changing its icon.

Read more AWS Architecture Icons · Azure architecture icons

Trademarks#

Amazon Web Services, AWS and the names and icons of AWS services are trademarks of Amazon.com, Inc. or its affiliates. Microsoft, Azure and the names and icons of Azure services are trademarks of the Microsoft group of companies. Other names are trademarks of their owners. This book is not affiliated with, or endorsed by, Amazon or Microsoft.

Primary sources#

Every version, date, status, quota, limit and default in this book comes from the research record, where each fact names its source and the date it was fetched. The sources are AWS documentation, including each service's document history; AWS What's New and the AWS News Blog; Microsoft Learn and Azure Updates; the Terraform registry and HashiCorp documentation; and OpenJDK and Maven Central. Quotas and defaults are stated as of as of September 2026, and costs are described without amounts, because prices change.

No AWS account was used to write this book. The Terraform and Java examples come from a specimen project that is formatted, validated and compiled offline, and command output is quoted only from the documentation, labelled with its source, or marked as illustrative.

Cheat card#

Appendices · Lookup · Chapter 47·1 min read

The book on one page: the thirteen design questions with MegaCorp's answers, and the words that mean something else on AWS.

Thirteen questions, and MegaCorp's answers#

Start a design from these answers, and write an ADR wherever yours differ.

QuestionMegaCorp's answerSee
1 Accountsan account per workload and environment, in OUs, with Control Tower's log-archive and audit accountschapter 15
2 Sign-inIAM Identity Center for people, roles for code, OIDC for pipelines; no long-lived keyschapter 10
3 Networkseu-west-2, recovering to eu-west-1; a VPC per account, attached to a transit gateway hubchapter 13
4 GuardrailsControl Tower controls and SCPs on every OU; Security Hub and Config check posturechapter 15
5 Logs and alertsthe organization trail into log-archive, GuardDuty, and CloudWatch alarms that wake someonechapter 28
6 DeliveryTerraform with state in S3, deployed by CodePipeline with GitHub as the sourcechapter 33
7 Where it runsECS on Fargate; Lambda for short work triggered by events; EC2 when the host matterschapter 16
8 DataAurora PostgreSQL for the ledger, DynamoDB for keys and duplicates, S3 for fileschapter 21
9 How parts talkSQS to buffer work, EventBridge to fan out, Kinesis for ordered streamschapter 23
10 Traffic inCloudFront with AWS WAF, then an Application Load Balancer; certificates from ACMchapter 24
11 Failuretwo zones, each able to carry the load; AWS Backup; a pilot light in eu-west-1chapter 30
12 Costtags on every resource, AWS Budgets alerts, and a Savings Plan for the steady basechapter 31
13 Attackcustomer managed keys, Secrets Manager, least-privilege roles, GuardDuty, and Session Manager for operatorschapter 34

Words that mean something else#

WordOn AzureOn AWSSee
rolepermissions assigned to a user, group or identity at a scopean identity that people and services assume for a limited timechapter 5
policyAzure Policy: rules for how resources are configureda JSON document of permissions; AWS Config rules check configurationchapter 5
security groupallow and deny rules in priority order, on a subnet or a network interfaceallow rules only, on a network interface, never on a subnetchapter 5
resource groupa container: deleting it deletes what it holdsa saved query over tags or a stack; the account is the containerchapter 5
Application Gatewaya layer 7 load balancer, with an optional WAFits match is an Application Load Balancer with AWS WAF; API Gateway fronts APIschapter 5
service endpointa subnet setting that sends service traffic over the backbonethe URL of a service's API; the match is a VPC endpointchapter 5
Availability Zonetraffic between zones is not chargedtraffic between zones is billed, leaving one and arriving in the otherchapter 11
private subneta subnet with default outbound access turned offa subnet with no route to an internet gatewaychapter 12
region pairmany regions have a pair, which geo-redundant storage copies tono Region has a partner: you choose where to recover, and replicatechapter 30
Key Vaultkeys, secrets and certificates in one vaultKMS for keys, Secrets Manager for secrets, ACM for certificateschapter 27
tag inheritancesubscription and resource group tags can flow to usage belowonly account tags flow down; tag each resourcechapter 31
App Configurationsettings and flags, read through client librariesAppConfig: each change is a release, rolled back on an alarmchapter 33
runbooka PowerShell or Python scripta Systems Manager document of steps, run across accounts and Regionschapter 34
Well-Architectedfive pillarssix pillars, adding sustainabilitychapter 35

End of the book