AWS for Azure Architects and Developers#
Design enterprise AWS systems end to end, whatever you know of Azure
Read this first#
Your next systems run on AWS. This book teaches AWS from first principles, with Azure as an optional bridge, so that after about 2 hours 55 minutes of reading you can design an enterprise-grade AWS system end to end.
Who it is for#
Architects and developers. You need no Azure knowledge: where Azure helps, it appears beside AWS in a figure pair, a table column, a profile's "On Azure" line or a false-friend box, and every AWS idea stands without it. You need no AWS account to read it either; the one hands-on step, a first sign-in in chapter 1, uses an account your organisation gives you.
How to read it#
- One story. MegaCorp's payments team moves to AWS, chapter by chapter. Each tool arrives when the team needs it, and its profile tells what it does, how it works, when to use it and when not to, its limits, cost, security and gotchas.
- Figures carry facts. Architecture figures use the official AWS and Azure icons, and numbered steps match the legend under each figure. Every figure has an "As text" version. chapter 1 shows how to read them.
- Nothing is folded away. What you see is the whole book. Each section ends with links to the official documentation, where the fine detail lives.
- Code only where it teaches. Java shows the AWS SDK and AWS Lambda. Terraform appears at most once a chapter, Azure beside AWS, where one tool shows how the clouds differ.
- Search with ⌘K or Ctrl+K. Type an Azure name, such as
Event Grid, to reach its AWS counterpart, or paste an error message.
What it covers#
Part 0 gets you started. Part I orients you: AWS on one page, its history, the six shifts in thinking and the false friends. Part II lays the foundations a company decides once: accounts, identity, networks and the landing zone. Part III covers the building blocks each workload chooses from. Part IV asks what every design must answer: security, failure, cost, delivery and operations. Part V puts it together in three worked designs and a set of review drills.
The teaching parts take about 2 hours 55 minutes to read, as the build measures them. The appendices, from the service map to the cheat card, are for lookup and sit outside that time.
Facts were checked on 12 September 2026, and a statement about what is current carries its date, because AWS changes every week. The official icons are used under each vendor's terms, credited in chapter 46.
Before you start#
Three things make the rest of the book easier: the shape of AWS on one page, how to read its figures, and a first sign-in that tells you exactly who and where you are.
The cloud on one page#
On its first day, MegaCorp's payments team asks what every newcomer asks: where does everything actually live? Almost always, the answer has two parts: an account and a Region.
An AWS account holds resources and is their security boundary. Nothing in one account can reach
another unless someone explicitly allows it, and costs are counted per account. A Region is a
geographic area, such as London, eu-west-2, made of several Availability Zones.
What you create in one Region does not exist in any other unless you copy or replicate it.
An Availability Zone is one or more data centres with their own power, networking and connectivity. The zones in a Region sit far enough apart, up to about 100 km, not to fail together, and close enough for synchronous replication within a few milliseconds. Every Region has three or more, as of September 2026.
A few services are global rather than Regional. IAM, which decides who may do what, and Route 53,
which answers DNS queries, serve every Region, although each is managed from one Region,
us-east-1.
As text
- AWS Cloud
- AWS Account payments-prod
- AWS IAM
- Amazon Route 53
- Region eu-west-2
- Amazon S3 (statements)
- VPC 10.20.0.0/16
- Availability Zone a
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Private subnet 10.20.10.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- Amazon EC2 (app)
- Private subnet 10.20.11.0/24
- Availability Zone a
- AWS Account payments-prod
Read more What is an AWS account? · AWS Fault Isolation Boundaries: Regions · AWS Fault Isolation Boundaries: Availability Zones · AWS Fault Isolation Boundaries: global services
How to read the figures#
Most figures in this book are architecture diagrams drawn the way AWS draws them, with the official icons. Boxes are groups: whatever sits inside a box belongs to it. The colour and line of a box say what kind of group it is.
| Box | Line | What it shows |
|---|---|---|
| AWS Cloud | dark grey | the whole of AWS |
| AWS account | pink | a boundary for security, cost and quotas |
| Region | teal, dotted | one geographic area |
| VPC | purple | a private network in one Region |
| Availability Zone | teal, dashed | one or more data centres within a Region |
| Public subnet | green | a subnet with a route to the internet |
| Private subnet | teal | a subnet without one |
| Security group | red | a firewall around resources |
| Auto Scaling group | orange, dashed | instances that grow and shrink in number |
Numbered black circles mark the steps of a flow, and the legend under the figure explains them in order. In the next figure, a customer's request arrives at a load balancer (step 1), which passes it to an instance in a private subnet (step 2).
| 1 | A customer's request reaches the load balancer in a public subnet |
| 2 | The load balancer passes it to an instance in a private subnet |
As text
- Customers
- Region eu-west-2
- VPC 10.20.0.0/16
- Availability Zone a
- Public subnet 10.20.0.0/24
- Application Load Balancer
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Public subnet 10.20.0.0/24
- Availability Zone a
- VPC 10.20.0.0/16
- A customer's request reaches the load balancer in a public subnet
- The load balancer passes it to an instance in a private subnet
Where Azure helps, a figure pair shows Azure on the left and AWS on the right. Every figure also has an "As text" version beneath it, which a screen reader can read and the search can find.
Read more AWS Architecture Icons · Azure architecture icons
A first sign-in#
Alex joins the payments team as a developer. MegaCorp signs its people in to AWS through IAM Identity Center, so Alex never holds a long-term password or access key for AWS. Instead, the AWS CLI asks Identity Center for short-term credentials for a role: an identity, with its own permissions, that Alex can assume for a limited time.
aws configure sso
aws sso login --profile payments-dev
aws sts get-caller-identity --profile payments-devaws configure sso runs once. It asks for MegaCorp's Identity Center start URL and
Region, opens a browser to sign in, lists the accounts and roles Alex may use, and saves the choice
as a profile. Later, aws sso login signs in again when the session has expired.
As text
- Alex to AWS CLI: aws sso login
- AWS CLI to IAM Identity Center: sign in through the browser
- IAM Identity Center to AWS CLI: a session token, cached on disk
- AWS CLI to IAM Identity Center: credentials for the profile's role
- IAM Identity Center to AWS CLI: short-term credentials
- Alex to AWS CLI: aws sts get-caller-identity
- AWS CLI to AWS STS: GetCallerIdentity
- AWS STS to AWS CLI: the account and the role's ARN
{
"UserId": "AIDASAMPLEUSERID",
"Account": "123456789012",
"Arn": "arn:aws:iam::123456789012:user/DevAdmin"
}The documentation's example comes from an IAM user. Signed in through Identity Center, Alex's
Arn names an assumed role instead, one that Identity Center created:
arn:aws:sts::123456789012:assumed-role/AWSReservedSSO_PowerUserAccess_…/alex.
Account says which account the command reached; check it before changing anything.
The profile also stores a default Region. Commands go to that Region, and resources in any other
Region do not appear in their results.
get-caller-identity always succeeds. It needs no permissions, and
works even when a policy denies it, so it proves who you are, never what you may do.
Read more Installing or updating the AWS CLI · Configuring IAM Identity Center authentication with the AWS CLI · AWS CLI: aws sts get-caller-identity
- Alex launches an instance in
eu-west-2, then lists instances with a profile whose default Region iseu-west-1. The new instance is missing. Why?Answer
The command went toeu-west-1, and resources in one Region do not exist in another. Run it with--region eu-west-2, or use a profile set to that Region. - In a figure, a green box sits inside a teal dashed box, which sits inside a purple box. What are the three boxes?
Answer
A public subnet, inside an Availability Zone, inside a VPC. - Why would the payments team run a copy of each tier in a second Availability Zone?
Answer
A zone is one or more data centres that can fail on its own, and the zones in a Region are built not to fail together. A copy in a second zone keeps running when the first fails. aws sts get-caller-identitysucceeds for Alex. Can Alex now list the team's instances?Answer
Not necessarily. The call needs no permissions. It shows who Alex is and which account the command reached, not what Alex may do.
AWS on one page#
AWS is a long list of separate services, but they fall into a handful of categories, and each category has its own icon colour. Learn them, and a figure tells you what each box does before you read the label.
A handful of categories#
When MegaCorp's payments team first opens the AWS console, the list of services is long enough to discourage anyone. It becomes manageable once the team sees that every service belongs to one category, and that a design needs only a few services from each.
Every AWS service icon sits on a square tile in its category's colour. The first figure shows where workloads run and where data lives; the second shows how the parts connect, and how they are kept safe and visible.
As text
- Compute
- AWS Lambda
- Containers
- AWS Fargate
- Storage
- Amazon S3
- Databases
- Amazon DynamoDB
As text
- Networking
- Amazon VPC
- Integration
- Amazon SQS
- Security
- AWS KMS
- Management
- Amazon CloudWatch
The table lists every category this book teaches, with the services it covers and where. Each service is its own product, with its own API, console pages, quotas and pricing; chapter 4 explains why that matters.
| Category | Icon colour | Services this book teaches | Chapters |
|---|---|---|---|
| Compute | orange | Amazon EC2, Amazon EC2 Auto Scaling, AWS Lambda | chapter 17, chapter 19 |
| Containers | orange | Amazon ECR, Amazon ECS, AWS Fargate, Amazon EKS | chapter 18 |
| Storage | green | Amazon S3, Amazon EBS, Amazon EFS, Amazon FSx, AWS Backup | chapter 20, chapter 30 |
| Databases | magenta | Amazon RDS, Amazon Aurora, Amazon DynamoDB, Amazon ElastiCache | chapter 21, chapter 22 |
| Networking and content delivery | purple | Amazon VPC, Amazon Route 53, Amazon CloudFront, Elastic Load Balancing, Amazon API Gateway | chapter 12, chapter 13, chapter 14, chapter 24 |
| Application integration | pink | Amazon SQS, Amazon SNS, Amazon EventBridge, AWS Step Functions | chapter 23, chapter 19 |
| Analytics | purple | Amazon Athena, AWS Glue, Amazon Redshift, Amazon Kinesis Data Streams, Amazon MSK | chapter 25, chapter 23 |
| Artificial intelligence | teal | Amazon Bedrock | chapter 26 |
| Security, identity and compliance | red | AWS IAM, AWS IAM Identity Center, AWS KMS, Amazon GuardDuty, AWS WAF | chapter 8, chapter 27, chapter 28 |
| Management and governance | pink | AWS Organizations, AWS Control Tower, Amazon CloudWatch, AWS CloudTrail, AWS Systems Manager | chapter 7, chapter 29, chapter 34 |
| Developer tools | magenta | AWS CloudFormation, AWS CDK, AWS CodePipeline, AWS CodeBuild | chapter 32, chapter 33 |
| Cloud financial management | green | AWS Cost Explorer, AWS Budgets, Savings Plans | chapter 31 |
A colour names a category, not a service, and two colours are shared. Pink marks both management and application integration, and purple both networking and analytics, so read the label before you assume.
Read more Overview of Amazon Web Services · AWS Architecture Icons
Colours in a design#
Once the colours are familiar, a design reads at a glance. A customer's request to MegaCorp enters through CloudFront (step 1), purple because it moves traffic, and passes a load balancer (step 2). It reaches the orange compute that runs the payments code (step 3) and ends at a magenta database (step 4).
| 1 | A customer's request reaches CloudFront at the edge |
| 2 | CloudFront forwards it to the load balancer |
| 3 | The load balancer passes it to a payments task |
| 4 | The task reads and writes the database |
As text
- Customers
- Amazon CloudFront
- Region eu-west-2
- VPC 10.20.0.0/16
- Application Load Balancer
- AWS Fargate (payments)
- Amazon Aurora
- VPC 10.20.0.0/16
- A customer's request reaches CloudFront at the edge
- CloudFront forwards it to the load balancer
- The load balancer passes it to a payments task
- The task reads and writes the database
Read more AWS Architecture Icons
How AWS got here#
AWS began in 2006 as a handful of separate services and grew one service at a time; Azure followed from 2008. That history explains why AWS feels like many products, and why older accounts look different from new ones.
Two clouds, two decades#
AWS launched in the spring of 2006 with Amazon S3, a service for storing objects, and added Amazon EC2, for renting servers, a few months later. Each was its own product with its own API. That pattern held: AWS grew by adding separate services, each launched when it was ready.
As text
- 2006. AWS: Amazon S3, then Amazon EC2.
- 2008. Azure: Windows Azure announced.
- 2009. AWS: Amazon VPC.
- 2010. Azure: Windows Azure generally available.
- 2011. AWS: AWS CloudFormation.
- 2014. AWS: AWS Lambda. Azure: Renamed Microsoft Azure.
- 2016. Azure: Azure Functions generally available.
- 2017. AWS: AWS Fargate.
- 2018. AWS: Amazon EKS generally available. Azure: Azure Kubernetes Service generally available.
- 2023. AWS: Amazon Bedrock generally available. Azure: Azure OpenAI Service generally available; Azure AD renamed Microsoft Entra ID.
The strip below follows one thread through that history: how much of the machine you run yourself. Each step hands more of it to AWS, and each is still on sale, so designs choose among them.
As text
- 2006: Amazon EC2: servers you manage
- 2014: AWS Lambda: code, no servers
- 2017: AWS Fargate: containers, no servers
- 2018: Amazon EKS: managed Kubernetes
Read more Overview of Amazon Web Services · Amazon VPC User Guide: document history
Eras you will meet#
History matters because accounts keep what they were given. The payments team inherits an AWS account from a company MegaCorp bought, and its age shows in three places.
- No default VPC. Accounts created after 4 December 2013 got a default VPC in every Region. An older account may have none, and its first instances ran in EC2-Classic, a flat network AWS has since retired.
- NAT instances. Before NAT gateways arrived in December 2015, private subnets reached the internet through NAT instances: ordinary instances the team had to run, patch and scale itself.
- No servers. From 2014, AWS Lambda, and later AWS Fargate, ran code and containers with no servers to manage. Newer parts of an estate lean on them.
As text
- Does every Region have a default VPC? No: Created before 4 December 2013, or someone deleted them. Yes: the next step.
- Do private subnets reach the internet through NAT instances? Yes: Built before December 2015, or never updated. No: the next step.
- Do any subnets use a regional NAT gateway? Yes: Networking updated since November 2025. No: the next step.
- Date the rest with the appendix on the AWS you inherit
An account's age sets its defaults. A script that assumes a default VPC fails in an account created before December 2013, or in one where the default VPC was deleted.
Read more Default VPCs · NAT instances
Why AWS feels like many products#
Because each service launched on its own, each still has its own API, console pages, quotas, pricing and documentation. Services overlap: there are several ways to run a container or send a message, each from a different year and for a different need. Old features retire slowly; EC2-Classic took from 2021 to 2023 to go. chapter 44 dates what an estate contains.
Six shifts in thinking#
Six ideas change how you design on AWS: accounts as boundaries, roles as identities, policies evaluated together, subnets tied to zones, services as separate products, and moving data as a cost. Each gets a chapter later; here is the map.
1. The account is the boundary#
MegaCorp gives payments three AWS accounts: development, test and production. An account holds resources, and it is also the boundary for security, cost and quotas. Nothing crosses it unless someone allows it, costs are counted per account, and quotas apply per account and Region. Many small accounts, grouped with AWS Organizations, take the place of a few large ones. chapter 7 develops this.
As text
On Azure:
- Microsoft Entra tenant MegaCorp
- Management group Workloads
- Subscription payments-prod
- Resource group rg-payments
- Storage accounts (statements)
- Resource group rg-payments
- Subscription payments-prod
- Management group Workloads
On AWS:
- Organization
- OU Workloads
- AWS Account payments-prod
- Region eu-west-2
- Amazon S3 (statements)
- Region eu-west-2
- AWS Account payments-prod
- OU Workloads
Read more What is an AWS account? · Terminology and concepts for AWS Organizations · Azure management groups
2. A role is an identity you assume#
Alex, from chapter 1, has no AWS password at all. Alex assumes a role, an identity with its own permissions, and receives credentials that last hours, not years. Code does the same: an instance or a Lambda function assumes a role, so no key sits in a configuration file. chapter 8 develops this.
As text
- Settlement job to AWS STS: AssumeRole for the settlement role
- AWS STS: checks the role's trust policy
- AWS STS to Settlement job: short-term credentials
- Settlement job to payments-prod: calls AWS with the role's permissions
Read more IAM roles
3. Policies are evaluated together#
Every request starts denied. It succeeds only if some policy allows it and no policy denies it: an explicit deny anywhere, in the caller's policies, the resource's policy or an organization's guardrail, always wins. A request from one account to another must be allowed on both sides. chapter 9 develops this.
Read more Policy evaluation logic
4. A subnet lives in one zone#
A subnet sits in exactly one Availability Zone, and nothing on it says public or private: its route table decides. Resilience is therefore a matter of subnet layout, with one subnet per zone for each tier. chapter 12 develops this.
Read more Subnets for your VPC
5. Every service is its own product#
Each AWS service has its own API, console pages, quotas and pricing, and nothing makes two
services name or tag things alike. Tags are optional and case sensitive, so Env and
env are two different keys. A naming and tagging scheme is the designer's job,
decided early. chapter 31 develops this.
Read more Tagging your AWS resources · Service Quotas documentation
6. Moving data costs money#
Data moving between zones is billed as it leaves and again as it arrives. Data through a NAT gateway is billed per gigabyte processed, and data leaving AWS for the internet is billed too, while data arriving from the internet is free. The figure marks the billed paths of one payments service. chapter 31 develops this.
| 1 | Zone a to zone b: billed as it leaves and as it arrives |
| 2 | Through the NAT gateway: billed per gigabyte processed |
| 3 | Out to the internet: billed as data transfer out |
As text
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- Amazon Aurora (primary)
- Private subnet 10.20.11.0/24
- VPC 10.20.0.0/16
- Zone a to zone b: billed as it leaves and as it arrives
- Through the NAT gateway: billed per gigabyte processed
- Out to the internet: billed as data transfer out
Resilience has a running cost. Spreading tiers across zones, as shift 4 asks, creates the zone-to-zone traffic that shift 6 bills. Design for both at once.
Read more Understanding data transfer charges · Amazon VPC pricing
False friends#
Azure and AWS use some of the same words for different things. Six cause most of the confusion in design reviews: role, policy, security group, resource group, gateway and endpoint.
Identity words#
In the payments team's first design review, an architect who knows Azure asks who holds "the Contributor role" on the payments account. On AWS the question has no answer, because the word means something else.
A set of permissions, a role definition, assigned to a user, group or managed identity at a scope. Assignments add up.
An identity with its own permissions and no password or keys. People and services assume it for a limited time.
Azure Policy checks how resources are configured, such as allowed regions or required tags, whoever makes the change.
A JSON document of permissions. IAM, resource-based and organization policies are evaluated together. The nearest match to Azure Policy is AWS Config rules.
As text
- Does it say which API actions a user or role may call? Yes: An IAM policy. No: the next step.
- Is it attached to a resource, such as a bucket or a queue? Yes: A resource-based policy. No: the next step.
- Does it cap what every identity in an account may do? Yes: A service control policy, set in AWS Organizations. No: the next step.
- To check how resources are configured, use AWS Config rules
Read more IAM roles · Policy evaluation logic · Evaluating resources with AWS Config rules · What is Azure role-based access control? · Overview of Azure Policy
Network words#
Three network words trip people up next.
A network security group: allow and deny rules in priority order, attached to a subnet, a network interface, or both.
Allow rules only, attached to a resource's network interface, never to a subnet. A rule can name another security group.
Application Gateway is a layer 7 load balancer for web traffic, routing on URL path and host, with an optional web application firewall.
API Gateway is a front door for REST, HTTP and WebSocket APIs. For Application Gateway's job, AWS uses an Application Load Balancer with AWS WAF.
As text
On Azure:
- Virtual network 10.20.0.0/16
- Subnet 10.20.1.0/24
- Application gateways
- Subnet 10.20.10.0/24
- Azure Virtual Machines (app)
- Subnet 10.20.1.0/24
On AWS:
- Region eu-west-2
- AWS WAF
- VPC 10.20.0.0/16
- Availability Zone a
- Public subnet 10.20.0.0/24
- Application Load Balancer
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Public subnet 10.20.0.0/24
- Availability Zone a
A subnet setting that sends traffic for services such as Azure Storage over the Azure backbone, and lets those services admit the subnet.
The URL of a service's API, such as https://dynamodb.us-west-2.amazonaws.com.
The match for the Azure feature is a VPC endpoint.
Read more Security groups · Azure network security groups · What is Azure Application Gateway? · AWS service endpoints
Organising words#
The last false friend matters most for estate design. On Azure the team would keep the payments resources in one resource group and delete them together. AWS has resource groups too, but they hold nothing.
A container for resources with a shared lifecycle. Each resource belongs to exactly one, and deleting the group deletes everything in it.
A saved query, over tags or a CloudFormation stack, listing resources in one Region. Resources do not live in it; the account is the container.
As text
On Azure:
- Subscription payments-prod
- Resource group rg-payments
- Azure Virtual Machines (app)
- Storage accounts (statements)
- Resource group rg-payments
On AWS:
- AWS Account payments-prod
- Region eu-west-2
- Resource group: app=payments
- AWS Lambda (settle)
- Amazon S3 (statements)
- Resource group: app=payments
- Region eu-west-2
Read more What are resource groups? · What is Azure Resource Manager?
Your first week#
Joining an AWS estate, you need four answers fast: who and where you are, what you may do, what is already there, and who changed what. Each takes one command, and AWS's refusals are more helpful than they look.
Who and where you are#
Alex's first real task is the account MegaCorp inherited when it bought a smaller payments company. Before touching anything, Alex confirms the account and Region the profile reaches, and whether the account belongs to MegaCorp's organization.
aws sts get-caller-identity --profile acquired
aws organizations describe-organization --profile acquired
aws ec2 describe-vpcs --profile acquired --query "Vpcs[].{id:VpcId,cidr:CidrBlock,default:IsDefault}"The commands name the account and role, the organization and its management account, and the VPCs in the profile's Region, flagging any default VPC.
Read more AWS CLI: aws sts get-caller-identity · Terminology and concepts for AWS Organizations · Default VPCs
What you may do#
Nobody hands Alex a list of permissions. Alex tries the task and reads the refusal: most access denied messages name the missing action and the kind of policy that refused it.
User: arn:aws:iam::123456789012:role/HR is not authorized to perform: codecommit:ListRepositories
because no identity-based policy allows the codecommit:ListRepositories actionAs text
- Does the message say "with an explicit deny"? Yes: A Deny statement blocks you: adding an allow will not help. No: the next step.
- Does it say "no service control policy allows"? Yes: The organization's guardrail stops the action: its owners decide. No: the next step.
- Does it say "no identity-based policy allows"? Yes: Your role lacks the permission: ask for it. No: the next step.
- Another policy lacks an allow, such as the resource's own policy
Read more Troubleshoot access denied error messages
What is already there#
For an inventory, AWS Resource Explorer, on by default since October 2025, searches the current Region by name, tag or ID, at no charge.
aws resource-explorer-2 search --profile acquired --query-string "tag:app=payments"Read more What is AWS Resource Explorer?
Who changed what#
On Wednesday the settlement job stops reaching its database. Someone changed a security group on Tuesday, and CloudTrail knows who. Its event history keeps 90 days of management calls in each Region, free and with no trail needed.
aws cloudtrail lookup-events --profile acquired --region eu-west-2 \
--lookup-attributes AttributeKey=ResourceName,AttributeValue=sg-0123456789abcdef0As text
- Alex to AWS CLI: lookup-events for the security group
- AWS CLI to AWS CloudTrail: LookupEvents in eu-west-2
- AWS CloudTrail to AWS CLI: AuthorizeSecurityGroupIngress, by a role session, on Tuesday
- AWS CLI to Alex: who made the call, and when
As text
- AWS Account acquired-payments
- Region eu-west-2
- AWS CloudTrail (event history)
- Region eu-west-1
- AWS CloudTrail (event history)
- Region eu-west-2
Event history searches one Region at a time. A change made in
eu-west-1 never appears in a search of eu-west-2. To keep events longer
than 90 days, or to search them together, the estate needs a trail or an event data store.
Read more Working with CloudTrail event history
The same questions on each cloud#
| Question | AWS CLI | Azure CLI |
|---|---|---|
| Who am I, and where? | aws sts get-caller-identity | az account show |
| Which organization? | aws organizations describe-organization | az account management-group list |
| Which networks? | aws ec2 describe-vpcs | az network vnet list |
| What is here? | aws resource-explorer-2 search | az resource list |
| Who changed what? | aws cloudtrail lookup-events | az monitor activity-log list |
Read more AWS CLI Command Reference
- An architect asks who holds the Contributor role on the payments account. What do you answer?
Answer
That role is an Azure idea. On AWS, ask which roles exist, who may assume each one under its trust policy, and what its policies allow. - A request fails "with an explicit deny in a service control policy". Will adding a permission to your role fix it?
Answer
No. An explicit deny always wins, and this one is an organization guardrail, so only its owners can change it. - The payments service in zone a calls its database in zone b all day. What does that cost?
Answer
Data between zones is billed as it leaves and again as it arrives, so every call is charged in both directions.
Accounts and Organizations#
On AWS the account is the boundary for security, cost and quotas, so an enterprise runs many small accounts. AWS Organizations groups them, pays for them together and caps what anyone in them may do.
The account is the unit#
MegaCorp's platform team starts with a question that shapes everything else: how many AWS accounts? The answer is more than anyone expects. Each workload gets an account per environment, and shared jobs such as logging and security get accounts of their own.
What it does. An account holds MegaCorp's resources and is their boundary for security, cost and quotas. Payments gets three: development, test and production.
How it works. Every resource's ARN carries its account ID, and nothing crosses to another account unless policies on both sides allow it. Each account also has a root user with complete access, which no IAM policy in the account can restrict.
When to use it. For each workload and environment, and for each shared job such as logging, security tooling and networking, so that a mistake or a breach stays inside one account.
When not to. For two teams that must share data all day, since every crossing needs policies on both sides. Nothing should run in the management account, which service control policies cannot restrict.
Limits. An organization starts with a quota of 10 accounts as of September 2026, raised on request up to 50,000. Most other quotas, such as 5 VPCs per Region, apply per account, so splitting accounts splits the quotas too.
Cost. Resources are billed to the account that holds them, and the organization's management account pays the whole bill.
Security. Never use the root user for daily work, and protect it with MFA. With centralized root access, the management account can remove root credentials from member accounts altogether.
Gotchas. A closed account can be reopened for 90 days. After that it is gone for good: its ID is never reused, its email address can never register another account, and until then it still counts against the organization's quota.
Read more What is an AWS account? · AWS account root user · Close an AWS account
An organization of accounts#
Twenty accounts need somewhere to live. MegaCorp creates an organization, keeps its management account empty, and groups the others by how they are governed: security accounts in one organizational unit, workloads in another, experiments in a third.
As text
- Organization root
- AWS Account management
- AWS Organizations
- Security OU
- AWS Account log-archive
- Amazon S3 (logs)
- AWS Account security-tooling
- Amazon GuardDuty
- AWS Account log-archive
- Workloads OU
- AWS Account payments-dev
- AWS Fargate (payments)
- AWS Account payments-prod
- AWS Fargate (payments)
- AWS Account payments-dev
- Sandbox OU
- AWS Account sandbox-alex
- Amazon EC2 (experiments)
- AWS Account sandbox-alex
- AWS Account management
What it does. AWS Organizations groups MegaCorp's accounts under one management account, pays their bills together, and applies guardrails to whole groups of accounts at once.
How it works. Accounts sit under a single root, in organizational units nested up to five levels deep. A service control policy attached to the root or an OU caps what every user and role below it may do. It never grants anything.
When to use it. From the second account onwards. Group accounts by how they are governed rather than by the org chart; AWS suggests starting with Security, Infrastructure, Workloads and Sandbox OUs.
When not to. As a substitute for IAM. Service control policies only cap permissions: people and code still need IAM policies before they can do anything at all.
Limits. Default quotas as of September 2026: one root, 2,000 OUs, 10 service control policies attached to each root, OU or account, and 10,240 characters in each policy.
Cost. No additional charge; each account's resources are billed as usual, on one bill.
Security. Service control policies do not apply to the management account, so it holds nothing but the organization. Each account it creates gets OrganizationAccountAccessRole, which gives the management account full administrative control.
Gotchas. Every OU and account must keep at least one service control policy. Removing the default
FullAWSAccess without a replacement makes every action in the member accounts
fail.
Read more AWS Organizations documentation · Terminology and concepts for AWS Organizations · Quotas and service limits for AWS Organizations · Recommended OUs and accounts
Guardrails with service control policies#
Guardrails are where Organizations earns its place. The platform team attaches a policy to the
Workloads OU that denies every Region except eu-west-1 and eu-west-2.
A developer with full administrator rights in payments-dev still cannot start an instance in
us-east-1, because every policy from the root down must allow an action, and none
may deny it.
As text
- Does any service control policy between the root and payments-dev deny the action? Yes: Denied, even for the account's root user. No: the next step.
- Does every service control policy on that path allow it? No: Denied by a guardrail above the account. Yes: the next step.
- Does an IAM policy grant it to the role? No: Denied: service control policies never grant. Yes: the next step.
- Allowed
Test a guardrail before it reaches the root. A policy that denies too much at the root locks out every member account at once, the security team's included. Try it first on an OU holding a few accounts.
Read more Service control policies (SCPs) · Creating a member account in an organization
MegaCorp's accounts#
The payments-prod account, in the Workloads OU, becomes the outer boundary of MegaCorp's running design. Every later chapter adds something inside it.
As text
- AWS Account payments-prod
- Region eu-west-2
Read more What is an AWS account?
IAM identities and roles#
Every request to AWS is signed by an identity, and the identities that matter are roles: people and code assume them and receive credentials that expire. A long-term key becomes the exception you have to justify.
Who is calling#
MegaCorp's settlement job has to write statements to Amazon S3. On a first attempt, a developer pastes an access key into the job's configuration. The platform team rejects the change: on AWS, code gets its access from a role, and a key in a file can leak and never expires.
What it does. AWS IAM decides who may do what in an account. The settlement job may write statements to one bucket, and do nothing else.
How it works. IAM holds identities, mainly roles, and the policies attached to them. Every request is signed with an identity's credentials, and AWS checks every applicable policy before the service acts.
When to use it. Always: every account uses it. Give roles to code and to people signing in through IAM Identity Center, and keep IAM users for the rare tool that can use neither.
When not to. Managing people account by account; IAM Identity Center does that centrally, as chapter 10 shows. The customers of an application belong in Amazon Cognito, not in IAM.
Limits. Defaults as of September 2026: 1,000 roles per account, raised up to 10,000, and 20 managed policies per role, up to 25. A role's inline policies total at most 10,240 characters.
Cost. No additional charge. IAM Access Analyzer's unused-access analysis and custom policy checks are billed.
Security. Grant least privilege, and let IAM Access Analyzer propose policies from what a role actually used. Any human IAM user left needs MFA, ideally a passkey or security key.
Gotchas. IAM is eventually consistent, so a role created a moment ago may not work yet. Never create IAM resources on a critical path, such as during a failover.
Read more AWS IAM documentation · Security best practices in IAM · IAM and AWS STS quotas
As text
- Is it a person in MegaCorp's workforce? Yes: Sign in through IAM Identity Center and assume a role. No: the next step.
- Is it code running on an AWS service, such as Amazon EC2 or AWS Lambda? Yes: A role the service assumes for it. No: the next step.
- Is it code outside AWS with an identity provider or certificates? Yes: Federation with OpenID Connect or SAML, or IAM Roles Anywhere. No: the next step.
- Is it another company's account, such as a monitoring vendor? Yes: A cross-account role with an external ID. No: the next step.
- Only then an IAM user with access keys, rotated and watched
Roles for code#
The settlement job gets a role called settlement. Its trust policy lets the compute
service assume it, and its permissions policy allows one action on one bucket. The code carries no
key: the AWS SDK finds the role's credentials, and they renew before they expire.
As text
- Settlement job to AWS SDK for Java: putObject
- AWS SDK for Java to Instance metadata: credentials for the instance's role
- Instance metadata to AWS SDK for Java: short-term credentials
- AWS SDK for Java to Amazon S3: PutObject, signed with them
- Instance metadata: renews them before they expire
What it does. AWS STS issues the short-term credentials behind every role. When the settlement role is assumed, STS returns a key, a secret and a session token, all of which expire.
How it works. A caller asks STS to assume a role. STS checks the role's trust policy and returns credentials valid for one hour unless the caller asks for longer, up to the role's maximum of at most 12 hours. Compute services make that call for your code.
When to use it. Whenever code or people need access: assuming a role in the same account or another, or exchanging a token from an identity provider for AWS credentials.
When not to. Never trade it for long-term keys to save effort. A vendor that monitors your accounts gets a cross-account role, not an IAM user of its own.
Limits. as of September 2026, 600 requests a second per account and Region, shared by AssumeRole, GetCallerIdentity and others. Calls that AWS services make for you do not count.
Cost. No additional charge.
Security. For a third party, require an external ID in the trust policy. It stops another customer of the same vendor from borrowing your role: the confused deputy problem.
Gotchas. Role chaining, one role assuming another, caps the session at one hour. Use Regional endpoints; the AWS SDK for Java 2.x and AWS CLI v2 already do.
Read more AWS STS API reference · AWS STS Regional endpoints · The confused deputy problem
Launching code with a role needs iam:PassRole. Without that check,
anyone allowed to start an instance could hand it a more powerful role and borrow its rights.
Grant iam:PassRole only for the roles a team should use.
Read more Use an IAM role for applications on Amazon EC2 · IAM roles
Crossing accounts#
The settlement job must also copy each statement into the log-archive account, which it cannot reach by default. Log-archive creates a role that trusts the settlement role: the job assumes it (step 1) and writes with the new credentials (step 2).
| 1 | The settlement role assumes statement-writer, whose trust policy names it |
| 2 | It writes the statement with statement-writer's permissions |
As text
- AWS Account payments-prod
- IAM role (settlement)
- AWS Account log-archive
- IAM role (statement-writer)
- Amazon S3 (statements)
- The settlement role assumes statement-writer, whose trust policy names it
- It writes the statement with statement-writer's permissions
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::111111111111:role/settlement" },
"Action": "sts:AssumeRole"
}]
}That is statement-writer's trust policy. The settlement role also needs permission of its own
to call sts:AssumeRole on statement-writer: one account trusts, and the other
allows.
Read more Cross-account policy evaluation logic
IAM policies#
Permissions on AWS come from policies: JSON documents attached to identities, resources and whole accounts. Several kinds are evaluated together on every request, and knowing their order is how you design access and fix a refusal.
Nine kinds of policy#
The settlement role from chapter 8 needs one permission: to write objects to the statements bucket. The platform team writes it as an identity-based policy, the commonest kind, attached to the role.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": "s3:PutObject",
"Resource": "arn:aws:s3:::megacorp-statements/*"
}]
}That is one of nine kinds of policy AWS evaluates. Most are JSON with the same shape: an effect, actions, resources and optional conditions. The kinds an architect meets most often:
| Kind | Attached to | What it does |
|---|---|---|
| Identity-based | a user, group or role | grants the identity actions on resources |
| Resource-based | a resource, such as a bucket, queue or role | grants named principals, even in other accounts, actions on that resource |
| Service and resource control policies | an organization root, OU or account | caps what identities and resources below may do; never grants |
| Permissions boundary | a user or role | caps what that identity's own policies can grant; never grants |
| Session policy | one role session | narrows a session below the role's permissions |
| VPC endpoint policy | a VPC endpoint | limits what can pass through that endpoint |
Prefer customer managed policies, which many roles can share and the team controls. AWS managed policies are quick to start with, but AWS changes them as services grow, and they rarely match least privilege.
Read more Policies and permissions in IAM · Managed policies and inline policies
How AWS decides#
When the settlement job calls PutObject, AWS gathers every policy that applies and
starts from a refusal. Within one account, the request succeeds only if nothing denies it, the
guardrails allow it, and the role's policy or the bucket's policy allows it.
As text
- Does any policy that applies say Deny? Yes: Denied: an explicit deny always wins. No: the next step.
- Do the organization's service and resource control policies allow it? No: Denied by a guardrail. Yes: the next step.
- Does the caller's identity-based policy, or the resource's own policy, allow it? No: Denied: nothing granted it. Yes: the next step.
- Do the caller's permissions boundary and session policy, where set, allow it? No: Denied by a cap on the caller. Yes: the next step.
- Allowed
| 1 | PutObject: the guardrails, the role's policy and the bucket's policy all agree |
As text
- Workloads OU: service control policies
- AWS Account payments-prod
- Region eu-west-2
- IAM role (settlement)
- Amazon S3 (megacorp-statements)
- Region eu-west-2
- AWS Account payments-prod
- PutObject: the guardrails, the role's policy and the bucket's policy all agree
Across accounts, as when the job writes to log-archive, AWS evaluates the request in both accounts, and each must allow it on its own.
Read more Policy evaluation logic · Cross-account policy evaluation logic
Tags as permissions#
MegaCorp soon has queues for payments, ledger and fraud, each with a developer role. Writing a
policy per team per queue does not scale. Instead the team tags every role and resource with
team, and one policy allows an action whenever the resource's team tag
matches the caller's. This is attribute-based access control.
| 1 | Tags match: the payments role reads the payments queue |
| 2 | Tags match: the ledger role reads the ledger queue |
As text
- AWS Account payments-dev
- Region eu-west-2
- IAM role (team=payments)
- Amazon SQS (team=payments)
- IAM role (team=ledger)
- Amazon SQS (team=ledger)
- Region eu-west-2
- Tags match: the payments role reads the payments queue
- Tags match: the ledger role reads the ledger queue
A new queue tagged team=payments is reachable at once, with no policy change,
which is the point: permissions scale as resources are added. The risk moves to the tags, so
who may set them must itself be controlled.
Read more Define permissions based on attributes with ABAC
Delegating safely#
Developers in payments-dev want to create roles for their own Lambda functions without waiting
for the platform team. A permissions boundary makes that safe: developers may create roles only if
they attach the developer-boundary policy, so no role they create can do more than
that boundary allows, whatever its own policy says.
A boundary caps, and never grants. A role with a boundary and no permissions policy can do nothing at all. It needs both, and it gets only what both allow.
Read more Permissions boundaries for IAM entities
Checking policies#
Policies drift: a bucket gets shared for one migration and stays shared. MegaCorp turns on an analyzer for the whole organization before the first workload goes live.
What it does. IAM Access Analyzer finds access MegaCorp did not mean to grant: a bucket shared outside the organization, a role nobody uses, a policy broader than its job.
How it works. An analyzer takes the organization or one account as its zone of trust, and reasons over resource-based policies to flag any that let a principal outside it in. Other analyzers find unused roles and permissions, and it can draft a policy from CloudTrail activity.
When to use it. From the first account: run an organization-wide external access analyzer in every Region in use, and validate each new policy before it ships.
When not to. As proof that a policy suits its job. It catches over-sharing and mistakes, not a permission the job is missing.
Limits. External access analysis is Regional, so each Region needs its own analyzer. A changed policy is analysed within about 30 minutes.
Cost. External access analysis is free. Unused access analysis is billed for each role and user analysed each month, and custom policy checks for each request.
Security. Each finding shows who outside the zone of trust can reach what. Archive it if the access is intended; otherwise, remove the access.
Gotchas. A policy generated from CloudTrail holds only the actions used in the chosen period, so a quarterly job that has not yet run will be missing its permissions.
Read more Using IAM Access Analyzer
Identity for people and code#
Three kinds of caller sign in to MegaCorp's systems: its people, its pipelines and its customers. Each has its own door: IAM Identity Center, OpenID Connect federation and Amazon Cognito.
People: one sign-in for every account#
Alex's first day on the payments team starts with a login MegaCorp already has: a Microsoft Entra ID account, used for email. Nobody creates an IAM user for Alex in each account. The platform team connected Entra ID to IAM Identity Center once, and Alex's group membership does the rest.
In 2013 IAM learned to trust a corporate identity provider over SAML, but account by account.
From 2017, AWS Single Sign-On did it once for the whole organization. Renamed IAM Identity Center
in 2022, it is why the CLI still says aws sso login.
As text
- 2013: IAM: SAML federation, account by account
- 2014: Amazon Cognito: identities for app users
- 2017: AWS Single Sign-On: one sign-in across accounts
- 2022: IAM Identity Center: the same service, renamed
| 1 | SCIM copies the group and its members; SAML signs them in |
| 2 | The PaymentsDeveloper permission set becomes a role in payments-dev |
| 3 | A read-only permission set becomes a role in payments-prod |
As text
- Microsoft Entra ID
- payments-developers
- AWS Account management
- AWS IAM Identity Center
- Workloads OU
- AWS Account payments-dev
- IAM role (PaymentsDeveloper)
- AWS Account payments-prod
- IAM role (PaymentsReadOnly)
- AWS Account payments-dev
- SCIM copies the group and its members; SAML signs them in
- The PaymentsDeveloper permission set becomes a role in payments-dev
- A read-only permission set becomes a role in payments-prod
What it does. IAM Identity Center gives MegaCorp's people one sign-in for every account. Alex signs in with Entra ID and picks payments-dev or payments-prod in the AWS access portal.
How it works. It trusts an identity source, here Entra ID, which signs people in over SAML and copies users and groups over SCIM. A permission set is a template of policies. Assign it to a group in chosen accounts, and Identity Center creates and maintains a matching role in each.
When to use it. For everyone who works in AWS, from the first account. It needs an organization instance, which always lives in the management account; administer it from a delegated member account.
When not to. For code, which uses roles, or for the customers of an application, who belong in Amazon Cognito.
Limits. Defaults as of September 2026: one instance per account; 3,500 permission sets, 500 of them in any one account; one inline policy per permission set; 100 groups per permission set in each account.
Cost. No extra charge.
Security. Assign groups, not people, so someone removed in Entra ID can no longer sign in. In the management account, assign people directly, or whoever edits the group decides who reaches it. Sessions last an hour by default, at most 12.
Gotchas. Entra ID provisions only the direct members of an assigned group, not members of nested groups. Removing a user's attribute in Entra ID leaves it in Identity Center, which matters when attributes drive access.
Read more AWS IAM Identity Center documentation · Manage AWS accounts with permission sets · Configure SAML and SCIM with Microsoft Entra ID and IAM Identity Center · Quotas and limits in IAM Identity Center
Pipelines: nothing to steal#
MegaCorp deploys the payments service from GitHub Actions. The first pipeline kept an access key in a repository secret, which works until the key leaks. Now each run gets a signed OpenID Connect token from GitHub and trades it with AWS STS for credentials that expire.
As text
- Deploy job to GitHub: a token for this run
- GitHub to Deploy job: signed: repository, branch, audience
- Deploy job to AWS STS: AssumeRoleWithWebIdentity with the token
- AWS STS: checks the signature and the deploy role's trust policy
- AWS STS to Deploy job: credentials for the deploy role
The account registers GitHub as an identity provider once. The deploy role's trust policy then accepts only runs on the payments repository's main branch.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::111111111111:oidc-provider/token.actions.githubusercontent.com" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:megacorp/payments:ref:refs/heads/main"
}
}
}]
}The sub condition is the lock. Without it, workflows in
repositories MegaCorp does not control could assume the role. IAM refuses a GitHub trust policy
with no sub condition or a bare wildcard, but repo:megacorp/* still admits
every repository in the organization.
Read more Create a role for OpenID Connect federation
Customers: a directory of their own#
Next comes a merchant portal where shops download their settlement statements. Merchants are not staff, so they belong in neither Entra ID nor IAM, but in a directory of their own.
What it does. Amazon Cognito signs in an application's own users. Merchants sign up for the portal and sign in with multi-factor authentication, and the portal receives standard OpenID Connect tokens.
How it works. A user pool is the directory: it stores users, signs them in and issues tokens. It can also federate, so a large merchant signs in through its own SAML or OIDC provider. An optional identity pool exchanges a token for short-term AWS credentials, for apps that call AWS directly.
When to use it. Sign-up and sign-in for the customers or partners of a web or mobile app, when you want a managed directory rather than running one.
When not to. For MegaCorp's staff, who use IAM Identity Center. For machine-to-machine calls at volume, check the bill first: each token is charged.
Limits. Defaults as of September 2026, per account and Region: 1,000 user pools, 1,000 app clients per pool and 40 million users per pool, all adjustable. API request rates have quotas by category.
Cost. User pools are billed per monthly active user at the pool's feature plan: Lite, Essentials or Plus. Lite and Essentials have a free tier, and federated users have their own rate. Machine-to-machine tokens are billed per token. Identity pools are free.
Security. Require multi-factor authentication, and consider the Plus plan's threat protection when accounts hold money. Give each app client access only to the attributes it needs.
Gotchas. Some choices are fixed when the pool is created: username or email sign-in, which
attributes are required, and every custom attribute. A wrong guess means a new pool and a
migration. Identify users by sub, never by email.
Read more Amazon Cognito documentation · What is Amazon Cognito? · Amazon Cognito pricing · Quotas in Amazon Cognito · Working with user attributes
MegaCorp's design gains its identity layer: merchants sign in through a user pool in eu-west-2, and the pipeline deploys through the deploy role.
| 1 | Merchants sign in through the user pool, which issues their tokens |
As text
- Customers
- AWS Account payments-prod
- Region eu-west-2
- Amazon Cognito (merchants)
- IAM, global to the account
- IAM role (deploy)
- Region eu-west-2
- Merchants sign in through the user pool, which issues their tokens
Regions and zones#
Every resource lives somewhere: in one zone, one Region, or everywhere. Which of the three decides what a failure takes down, where data may sit, and why a Region in Virginia matters to all.
Choosing a Region#
MegaCorp's merchant data must stay in the UK, so the payments team builds in eu-west-2, London, after checking the Regional Services List for every service the design needs. Resources and data stay in their Region unless you copy them, and the console shows one Region at a time.
Regions launched after 20 March 2019, such as Europe (Spain), are opt-in: nobody can use one until it is enabled, and enabling copies the account's IAM data there, which can take hours.
Disabling a Region deletes nothing. The resources in a disabled opt-in Region remain and keep billing; only access to them is lost. Remove them first.
Read more Enable or disable AWS Regions in your account · AWS Regional Services List
Zone names differ between accounts#
The settlement job in payments-prod reads the ledger database in another account. Both teams chose eu-west-2a to keep traffic in one zone, yet the bill shows traffic between zones: AWS maps zone names to physical zones at random for each account.
| 1 | Both sit in eu-west-2a, yet in different zones: each read is billed as it leaves one zone and enters the other |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Availability Zone a (euw2-az2)
- Amazon EC2 (settlement job)
- Availability Zone a (euw2-az2)
- VPC 10.20.0.0/16
- Region eu-west-2
- AWS Account ledger
- Region eu-west-2
- VPC 10.40.0.0/16
- Availability Zone a (euw2-az1)
- Amazon RDS (ledger)
- Availability Zone a (euw2-az1)
- VPC 10.40.0.0/16
- Region eu-west-2
- Both sit in eu-west-2a, yet in different zones: each read is billed as it leaves one zone and enters the other
A zone ID, such as euw2-az1, names the same physical zone in every account. When accounts must agree on a zone, plan and automate with zone IDs.
Zone numbers also map differently in each subscription, but data moving between zones in a region is not charged.
Data moving between zones in a Region is billed, as it leaves one zone and as it arrives in the other.
Read more Availability Zone IDs for your AWS resources · What are Azure availability zones?
Global services, and us-east-1#
IAM, AWS Organizations, Amazon Route 53 and Amazon CloudFront are global: they answer everywhere, but their control planes sit in us-east-1, in northern Virginia.
As text
- Alex to IAM, us-east-1: CreateRole
- IAM, us-east-1: copies the role to every Region, a moment later
- Settlement job to Amazon S3, eu-west-2: PutObject, signed with the role
- Amazon S3, eu-west-2 to Settlement job: checked in the Region
So keep changes to global services out of recovery plans, as chapter 8 said of IAM, and manage CloudFront certificates in us-east-1.
As text
- AWS IAM (global)
- Region eu-west-2
- Amazon S3 (regional)
- VPC 10.20.0.0/16
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway (zonal)
- Public subnet 10.20.0.0/24
- Availability Zone a
| Scope | Examples | What it means for a design |
|---|---|---|
| Zonal | a subnet, an EC2 instance, a NAT gateway | fails with its zone, so run one in each of at least two zones |
| Regional | a VPC, an S3 bucket, a Cognito user pool | spans the Region's zones; another Region needs its own copy |
| Global | IAM, Organizations, Route 53, CloudFront | answers everywhere; changes go through one Region |
Read more AWS Fault Isolation Boundaries: global services
Closer than a Region#
| Location | Where it is | Use it for |
|---|---|---|
| Local Zone | an extension of a Region, near a city | low latency, or data that must stay local; enable it, then add a subnet |
| Wavelength Zone | inside a carrier's 5G network | very low latency to mobile devices |
| AWS Outposts | AWS racks or servers in your own data centre | AWS services on premises, managed as part of a Region |
Read more Regions and Zones · How AWS Local Zones work
VPC fundamentals#
A VPC is your private network in one AWS Region. Three decisions shape every VPC: how to divide its addresses, which zone each subnet lives in, and how traffic leaves for the internet and for AWS services.
A network in one Region#
MegaCorp's payments team is moving its platform to AWS, into the London Region,
eu-west-2. First it needs a network to put the platform in. MegaCorp's address plan,
drawn up for the whole estate so that no two networks overlap, gives payments the range
10.20.0.0/16.
A CIDR block writes an address range as a base address and a prefix length.
10.20.0.0/16 fixes the first 16 of 32 bits and leaves 16 free: 65,536 addresses.
Each extra bit of prefix halves the range, so a /24 holds 256. Ranges that share any
address overlap.
What it does. An Amazon VPC (virtual private cloud) is the team's own network in one Region, isolated from every other network in AWS. Everything the platform runs takes its private address from it.
How it works. The team creates it with 10.20.0.0/16; any range from /16 to /28 is allowed.
Subnets divide the range, one zone each, and route tables steer their traffic.
When to use it. Anything with a private address. Each environment gets its own VPC, so test can never touch production.
When not to. The default VPC that AWS created in the Region has only public subnets: fine for an experiment, not for payments.
Limits. By default as of September 2026: 5 VPCs per Region, and 5 IPv4 ranges and 200 subnets per VPC, all adjustable.
Cost. The VPC is free; traffic is not. Public IPv4 addresses cost by the hour, even when idle, and data between zones is billed as it leaves and again as it arrives.
Security. Without a route to an internet gateway, nothing on the internet can reach a subnet. VPC Flow Logs record who talked to whom; VPC Block Public Access can forbid internet access across the account.
Gotchas. A range cannot be resized, only added to, so the team took a /16 from the plan
instead of guessing. It avoids 172.17.0.0/16, which some AWS services use.
Read more Amazon VPC documentation · Amazon VPC quotas · Amazon VPC pricing
A subnet lives in one zone#
Next come the subnets, and the first rule that shapes every AWS network: a subnet sits in exactly one Availability Zone and cannot span zones. To survive the loss of a zone, the team creates matching subnets in two zones and spreads each tier across them.
As text
On Azure:
- UK South
- Virtual network 10.20.0.0/16
- Subnet 10.20.10.0/24
- Azure Virtual Machines (zone 1)
- Azure Virtual Machines (zone 2)
- Subnet 10.20.10.0/24
- Virtual network 10.20.0.0/16
On AWS:
- Region eu-west-2
- VPC 10.20.0.0/16
- Availability Zone a
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Private subnet 10.20.10.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- Amazon EC2 (app)
- Private subnet 10.20.11.0/24
- Availability Zone a
- VPC 10.20.0.0/16
| Question | AWS | Azure |
|---|---|---|
| Where a subnet lives | one Availability Zone | the whole region, across its zones |
| IPv4 subnet sizes | /16 to /28 | /29 to /2 |
| Reserved addresses in each subnet | the first four and the last | the first four and the last |
Read more Subnets for your VPC · Azure Virtual Network FAQ
Routing makes a subnet public or private#
The team wants public subnets for the load balancer and the NAT gateways, and private ones for everything else. Nothing on a subnet says "public", though: a public subnet is one whose route table has a direct route to an internet gateway.
As text
- Does its route table have a route to an internet gateway? Yes: Public subnet. No: the next step.
- Does any route lead out of the VPC? No: Isolated subnet, reachable only inside the VPC. Yes: the next step.
- Does it route to a Site-to-Site VPN through a virtual private gateway? Yes: VPN-only subnet. No: the next step.
- Private subnet: it reaches the internet only through a NAT device
Read more Configure route tables · Internet gateways
Getting out: NAT gateways#
The settlement job is the first to need the internet. The figure shows the team's layout.
| 1 | Internet traffic from zone a goes to the NAT gateway in zone a |
| 2 | The NAT gateway sends it out through the internet gateway |
| 3 | Zone b routes to its own NAT gateway, so a failure stays in one zone |
| 4 | Traffic for Amazon S3 takes the gateway endpoint's route, not the NAT gateway |
As text
- Region eu-west-2
- Amazon S3
- VPC 10.20.0.0/16
- Internet gateway
- VPC endpoints (gateway)
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Amazon EC2 (app)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Amazon EC2 (app)
- Public subnet 10.20.1.0/24
- Internet traffic from zone a goes to the NAT gateway in zone a
- The NAT gateway sends it out through the internet gateway
- Zone b routes to its own NAT gateway, so a failure stays in one zone
- Traffic for Amazon S3 takes the gateway endpoint's route, not the NAT gateway
What it does. The settlement job runs in a private subnet but must call the card scheme's API. A NAT gateway lets it connect out, while nothing on the internet can connect in.
How it works. The job's internet traffic goes to the NAT gateway in its zone (step 1), which sends it out from its Elastic IP address (step 2). Regional NAT gateways, added in November 2025, span the zones themselves.
When to use it. Private workloads that call the internet: the card scheme, package repositories, partners.
When not to. Traffic for AWS services. Sending the nightly statements to Amazon S3 this way bills every gigabyte as processed; the next section avoids that.
Limits. It scales to 100 Gbps and ten million packets a second, then drops packets. One address holds 55,000 open connections to one destination, so the month-end batch gets a second address. Default quota as of September 2026: 5 per zone.
Cost. Billed per hour and per gigabyte processed, plus data transfer. A gateway in each of two zones doubles the hourly charge: the price of resilience.
Security. It accepts no connection from outside. It cannot have a security group; the resources behind it, and its subnet's network ACL, do the filtering.
Gotchas. A gateway shared by two zones fails with its own zone and takes the other's internet access with it, hence one per zone (step 3). It also ignores traffic arriving over VPC peering.
Read more NAT gateway basics · Regional NAT gateways
Reaching Amazon S3 without the internet#
The ledger service has a different need. It never talks to the internet, only to Amazon S3, where it writes the day's statements every night.
What it does. A gateway endpoint gives private subnets their own route to Amazon S3 or DynamoDB in the same Region, with no NAT gateway in the path.
How it works. It adds a route whose destination is S3's prefix list, the address ranges S3 uses in
eu-west-2. Being more specific than 0.0.0.0/0, that route wins
(step 4).
When to use it. Any VPC whose private subnets use S3 or DynamoDB. It is free, so the team adds one to every VPC.
When not to. Callers outside the VPC: the office network, peered VPCs, VPN and transit gateway traffic. They need an interface endpoint, which chapter 13 covers.
Limits. Its own Region only; a bucket elsewhere is still reached through the NAT gateway. Default quota as of September 2026: 20 per Region.
Cost. Nothing. Moving the nightly export off the NAT gateway removes the per-gigabyte charge on every statement.
Security. An endpoint policy says what may pass; the default allows everything, so the team limits
it to MegaCorp's buckets. Bucket policies can insist on the endpoint with
aws:sourceVpce.
Gotchas. Adding it drops open connections to S3, so the team adds it before go-live. Requests now
come from private addresses, so aws:SourceIp conditions stop matching; they need
aws:VpcSourceIp.
Read more Gateway endpoints · Gateway endpoints for Amazon S3
Security groups and network ACLs#
Two firewalls filter traffic inside a VPC. Security groups are the primary control; network ACLs add a coarse, stateless layer across a whole subnet.
| Characteristic | Security group | Network ACL |
|---|---|---|
| Attached to | a resource, such as an instance | a subnet |
| Rules | allow rules only | allow and deny rules |
| Evaluation | every rule, then one decision | in ascending order, until one matches |
| Return traffic | allowed automatically: stateful | allowed only by a rule: stateless |
| Rules can name | IP ranges, prefix lists, other security groups | IP ranges |
What it does. A firewall on each resource's network interface. The load balancer, the application instances and the database each get their own group.
How it works. Rules only allow. The application's group admits port 8080 from the load balancer's group rather than from addresses, so new instances are admitted with no rule change. Replies return automatically: the group is stateful.
When to use it. Always. It is the main control between tiers.
When not to. To deny something, such as a hostile address range. Security groups cannot deny; a network ACL on the subnet can, so the team uses one for that, sparingly.
Limits. By default as of September 2026: 60 inbound and 60 outbound rules per group, and 5 groups per network interface, adjustable to 16. Rules times groups may not exceed 1,000.
Cost. No charge.
Security. Only a few IAM principals may change groups, and no group opens port 22 or 3389 to the internet. Traffic to Amazon DNS and to instance metadata is never filtered.
Gotchas. A new group allows all outbound traffic until the team narrows it. A group belongs to one VPC, unless it is associated with other VPCs in the same Region.
Read more Security groups · Compare security groups and network ACLs · Azure network security groups
The same subnets in Terraform#
One Terraform comparison shows the zone rule in code. Every aws_subnet names its
Availability Zone, so resilient AWS subnets come in sets; an azurerm_subnet has no zone
to name.
resource "azurerm_subnet" "gateway" {
name = "snet-gateway"
resource_group_name = azurerm_resource_group.payments.name
virtual_network_name = azurerm_virtual_network.payments.name
address_prefixes = ["10.20.0.0/24"]
}
resource "azurerm_subnet" "app" {
name = "snet-app"
resource_group_name = azurerm_resource_group.payments.name
virtual_network_name = azurerm_virtual_network.payments.name
address_prefixes = ["10.20.10.0/24"]
default_outbound_access_enabled = false
}data "aws_availability_zones" "available" {
state = "available"
}
# 10.20.0.0/24 and 10.20.1.0/24: one public subnet in each of two zones
resource "aws_subnet" "public" {
count = 2
vpc_id = aws_vpc.payments.id
cidr_block = cidrsubnet(aws_vpc.payments.cidr_block, 8, count.index)
availability_zone = data.aws_availability_zones.available.names[count.index]
tags = { Name = "payments-public-${count.index}" }
}
# 10.20.10.0/24 and 10.20.11.0/24: one private subnet in each of the same zones
resource "aws_subnet" "private" {
count = 2
vpc_id = aws_vpc.payments.id
cidr_block = cidrsubnet(aws_vpc.payments.cidr_block, 8, count.index + 10)
availability_zone = data.aws_availability_zones.available.names[count.index]
tags = { Name = "payments-private-${count.index}" }
}A subnet with default outbound access turned off. The azurerm provider still defaults
default_outbound_access_enabled to true, so the specimen sets it.
A subnet with no route to an internet gateway.
Read more Terraform: aws_subnet · Terraform: azurerm_subnet · Default outbound access in Azure
How the VPC got here#
Amazon VPC began in 2009 as a VPN bridge to company networks. The flat network that came before it, EC2-Classic, shared with other customers, is now retired.
As text
- 2009: Amazon VPC, reached over a VPN
- 2011: VPCs across zones, several per account
- 2015: NAT gateways
- 2021: EC2-Classic retirement begins (retired)
- 2025: Regional NAT gateways
Read more Amazon VPC User Guide: document history · EC2-Classic is retiring: here's how to prepare
MegaCorp's network#
The figure adds the payments network to MegaCorp's running design.
| 1 | The NAT gateway sends traffic out through the internet gateway |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Public subnet 10.20.1.0/24
- VPC 10.20.0.0/16
- Region eu-west-2
- The NAT gateway sends traffic out through the internet gateway
Read more What is IPAM?
Connecting networks#
MegaCorp's VPCs multiply with its accounts. Each must reach a few others, the data centre and shared services, but never everything.
| Need | Use | Why |
|---|---|---|
| One service that many VPCs call | AWS PrivateLink | consumers reach the service, never each other's networks |
| Two VPCs that must talk | VPC peering | one to one, with nothing in between |
| Teams in one trust boundary sharing a network | a shared VPC, through AWS RAM | one network, one owner of routes |
| Many VPCs, or the data centre | a transit gateway | one attachment each, segmented by route tables |
Two VPCs: peering#
The settlement job in payments-prod must reach the ledger database in another account. A peering connection joins the two VPCs once both owners agree and add routes. Ranges must not overlap, and peering is not transitive: ledger cannot reach a third VPC through it, nor use payments-prod's NAT gateway or VPN.
Peering can make a hub's VPN or ExpressRoute gateway a spoke's way to the data centre, a setting called gateway transit, and user-defined routes can send a spoke's traffic through an appliance in the hub.
Peering joins exactly two VPCs. Neither can use the other's VPN or NAT gateway, or reach a third network through it; a hub on AWS is a transit gateway.
Read more How VPC peering connections work · Azure virtual network peering
Many VPCs: a transit gateway#
With fourteen VPCs, every pair would need 91 peering connections. The network team builds one transit gateway in the network account, shares it through AWS RAM, and each VPC attaches once.
| 1 | The network account shares the hub with the Workloads OU through AWS RAM |
| 2 | payments-prod sends traffic for other networks to the hub |
| 3 | payments-dev attaches too, but its route table has no route to production |
| 4 | Traffic for the data centre leaves the hub through the VPN |
| 5 | Two IPsec tunnels cross the internet to the data centre's router |
As text
- Customer gateway (data centre)
- AWS Site-to-Site VPN
- Workloads OU
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Transit gateway attachment (attachment)
- VPC 10.20.0.0/16
- Region eu-west-2
- AWS Account payments-dev
- Region eu-west-2
- VPC 10.30.0.0/16
- Transit gateway attachment (attachment)
- VPC 10.30.0.0/16
- Region eu-west-2
- AWS Account payments-prod
- AWS Account network
- Region eu-west-2
- AWS Transit Gateway (hub)
- AWS Resource Access Manager (share)
- Region eu-west-2
- The network account shares the hub with the Workloads OU through AWS RAM
- payments-prod sends traffic for other networks to the hub
- payments-dev attaches too, but its route table has no route to production
- Traffic for the data centre leaves the hub through the VPN
- Two IPsec tunnels cross the internet to the data centre's router
What it does. AWS Transit Gateway is a Regional router that VPCs, VPNs and Direct Connect attach to: one connection per network, not one per peer.
How it works. Each VPC attaches through one subnet per zone, and route tables decide who reaches whom.
When to use it. More than a handful of VPCs, or one way in for the data centre.
When not to. Two VPCs on their own: peering is simpler.
Limits. Defaults as of September 2026: 5 per account per Region and 5,000 attachments each, both adjustable; up to 100 Gbps per VPC attachment per zone.
Cost. Billed per attachment per hour, and per gigabyte sent into it, charged to the sender.
Security. Segment with route tables and blackhole routes, and share one hub through AWS RAM.
Gotchas. Only zones with an attachment subnet reach the hub. A stateful appliance behind it needs appliance mode, or replies are dropped.
Read more AWS Transit Gateway documentation · How AWS Transit Gateway works · AWS Transit Gateway quotas · AWS Transit Gateway pricing
One VPC, many accounts#
For a dozen small tools, the network team shares one VPC's subnets with their accounts through AWS RAM, and alone changes the routes.
What it does. AWS RAM shares a resource one account owns with other accounts or OUs: subnets, transit gateways, DNS rules and more.
How it works. The owner creates a resource share. Inside the organization no invitation is needed, and the resource appears in the other account as if it were its own.
When to use it. One network resource serving many accounts.
When not to. When the other account must fully control the resource: the owner keeps it.
Limits. Regional resources are shared within their Region, and global ones only from us-east-1.
Cost. No additional charge.
Security. The share's permission caps what other accounts can do, and their own policies and SCPs still apply.
Gotchas. Participants in a shared VPC cannot create NAT gateways, change routes or use the owner's default security group.
Read more AWS RAM documentation · Share your VPC subnets with other accounts · Responsibilities and permissions for owners and participants
Offering one service privately#
Twenty VPCs call the fraud team's scoring API. Instead of routing them all into its VPC, fraud publishes an endpoint service, and each consumer adds an interface endpoint.
| 1 | Requests go one way, from the endpoint to the service, on the AWS network |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- VPC endpoints (interface endpoint)
- VPC 10.20.0.0/16
- Region eu-west-2
- AWS Account fraud
- Region eu-west-2
- VPC 10.50.0.0/16
- AWS PrivateLink (scoring API)
- VPC 10.50.0.0/16
- Region eu-west-2
- Requests go one way, from the endpoint to the service, on the AWS network
What it does. AWS PrivateLink lets a VPC reach one service in another VPC, or an AWS service API, through private addresses in its own subnets.
How it works. The provider fronts the service with a load balancer and allows chosen accounts. A consumer's interface endpoint puts a network interface in each chosen subnet.
When to use it. One service for many VPCs, accounts or companies, and AWS APIs without the internet.
When not to. Networks that must talk both ways. For Amazon S3, the gateway endpoint in chapter 12 often does.
Limits. One network interface per chosen subnet, so plan one subnet per zone.
Cost. Billed per endpoint per zone per hour, and per gigabyte processed.
Security. The default endpoint policy allows everything: narrow it.
Gotchas. An endpoint in only one zone makes clients elsewhere cross zones, and pay for it.
Read more AWS PrivateLink documentation · AWS PrivateLink concepts · AWS PrivateLink pricing
Reaching the data centre#
Settlement files still go to a mainframe in MegaCorp's data centre. Site-to-Site VPN links the hub to it first; Direct Connect follows when volumes grow, with the VPN as backup.
What it does. AWS Site-to-Site VPN joins a VPC or transit gateway to an on-premises network through IPsec tunnels over the internet.
How it works. Each connection has two tunnels, for high availability, and ends on a virtual private gateway or a transit gateway.
When to use it. A first link, small sites, and the backup for Direct Connect.
When not to. Large, steady volumes or strict latency: internet paths vary.
Limits. as of September 2026: 1.25 Gbps per tunnel, or up to 5 Gbps with large bandwidth tunnels on a transit gateway.
Cost. Billed per connection per hour, plus data transfer out.
Security. IPsec encrypts traffic between AWS and the customer's router.
Gotchas. Overlapping address ranges between the VPCs and the data centre cannot be routed apart, so keep them distinct.
Read more AWS Site-to-Site VPN documentation · What is AWS Site-to-Site VPN?
What it does. AWS Direct Connect is a private fibre link from MegaCorp's network to AWS at a Direct Connect location.
How it works. A dedicated connection is a port for one customer; a hosted one comes through a partner. Virtual interfaces over it reach VPCs, transit gateways or public AWS services.
When to use it. Large, steady volumes, or latency that must be predictable.
When not to. As the only link for anything critical.
Limits. The 99.99% model needs separate devices in more than one location.
Cost. Billed per port hour, plus data transfer out, charged to the account that sends it.
Security. Not encrypted by default: use MACsec where supported, or VPN over it.
Gotchas. A private virtual interface reaches one VPC; to reach many, use a Direct Connect gateway.
Read more AWS Direct Connect documentation · What is Direct Connect? · AWS Direct Connect Resiliency Toolkit · Encryption in AWS Direct Connect
Inspecting traffic#
Payment data brings auditors, who want traffic from the data centre inspected and outbound traffic limited to known domains. Security groups cannot filter by domain, so MegaCorp adds AWS Network Firewall to the hub.
What it does. AWS Network Firewall is a managed, stateful firewall with intrusion prevention and domain allow-lists for VPC traffic.
How it works. An endpoint in a dedicated subnet per zone inspects what route tables send it; or the firewall attaches straight to a transit gateway.
When to use it. Central inspection between networks, to the internet, or from on premises.
When not to. As a replacement for security groups: it adds to them.
Limits. An endpoint cannot filter its own subnet, so firewall subnets hold nothing else.
Cost. Billed per endpoint per zone per hour and per gigabyte; a NAT gateway on the same path has its charges waived.
Security. Manage firewalls across accounts with AWS Firewall Manager.
Gotchas. Behind a transit gateway, appliance mode keeps both directions of a flow on one endpoint.
Read more AWS Network Firewall documentation · What is AWS Network Firewall? · AWS Network Firewall pricing
MegaCorp's design gains its connectivity layer.
| 1 | The NAT gateway sends traffic out through the internet gateway |
| 2 | Traffic for MegaCorp's other networks leaves through the attachment to the shared hub |
| 3 | The hub, owned by the network account, reaches the data centre over Site-to-Site VPN |
As text
- Customer gateway (data centre)
- AWS Account payments-prod
- Region eu-west-2
- AWS Transit Gateway (shared hub)
- VPC 10.20.0.0/16
- Internet gateway
- Transit gateway attachment (attachment)
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Public subnet 10.20.1.0/24
- Region eu-west-2
- The NAT gateway sends traffic out through the internet gateway
- Traffic for MegaCorp's other networks leaves through the attachment to the shared hub
- The hub, owned by the network account, reaches the data centre over Site-to-Site VPN
DNS#
Names on AWS come from Amazon Route 53: public names for the internet, private names inside VPCs, and a resolver in every VPC that forwards to and from the data centre.
Public and private names#
The merchant portal needs portal.megacorp.com on the internet; the settlement job needs ledger.internal.megacorp.com inside MegaCorp's VPCs only. The first lives in a public hosted zone, the second in a private hosted zone associated with those VPCs.
What it does. Amazon Route 53 is AWS's DNS: public and private zones, routing policies, and health checks.
How it works. A hosted zone holds a domain's records; a private zone answers only in its associated VPCs. Alias records point a name, even the apex, at an AWS resource and follow its address changes.
When to use it. Every domain on AWS, and routing by weight, latency, location or health, such as failover between Regions.
When not to. Moving a corporate domain that another provider must keep: delegate a subdomain instead.
Limits. Defaults as of September 2026: 500 hosted zones per account and 10,000 records per zone, both adjustable.
Cost. Billed per hosted zone per month and per query; alias queries to AWS resources are free.
Security. VPC Resolver can validate DNSSEC, and DNS Firewall blocks listed domains.
Gotchas. A private zone needs DNS hostnames and DNS support turned on in each VPC.
Read more Amazon Route 53 documentation · Considerations when working with a private hosted zone · Choosing a routing policy · Route 53 quotas · Amazon Route 53 pricing
As text
- Does a forwarding rule match the name? Yes: Forward the query, such as to the data centre. No: the next step.
- Does a private hosted zone associated with the VPC match? Yes: Answer from that zone, or "no such domain" if it has no record. No: the next step.
- Resolve the name on the internet
A private zone hides the public names it overlaps. Had the team called the private zone megacorp.com, VPC Resolver would answer every megacorp.com query inside the VPCs from it. A name only the public zone holds, such as portal.megacorp.com, would get "no such domain". A separate internal subdomain avoids the clash.
Names across the data centre#
The mainframe's name lives in the data centre's DNS, which must also resolve internal.megacorp.com. The network team adds an outbound endpoint with a forwarding rule for corp.megacorp.com, and an inbound endpoint the data centre forwards to, and shares the rule through AWS RAM.
As text
- Settlement job to VPC Resolver: mainframe.corp.megacorp.com?
- VPC Resolver: the rule for corp.megacorp.com matches
- VPC Resolver to Outbound endpoint: forward the query
- Outbound endpoint to Data centre DNS: over the VPN or Direct Connect
- Data centre DNS to Settlement job: the mainframe's address
Read more What is Route 53 VPC Resolver? · Choosing between alias and non-alias records
MegaCorp's design gains its DNS layer.
| 1 | The NAT gateway sends traffic out through the internet gateway |
| 2 | VPC Resolver answers internal.megacorp.com from the private hosted zone |
As text
- AWS Account payments-prod
- Route 53 hosted zone (internal.megacorp.com)
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Route 53 VPC Resolver (VPC Resolver)
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Public subnet 10.20.1.0/24
- VPC 10.20.0.0/16
- The NAT gateway sends traffic out through the internet gateway
- VPC Resolver answers internal.megacorp.com from the private hosted zone
Landing zones#
MegaCorp's accounts were set up one at a time. A landing zone gives every account the same logging, access and controls from birth, and keeps them there.
From accounts to a landing zone#
By its second year MegaCorp has thirty hand-built accounts, and nobody can promise the auditors that every one logs every API call where its own administrators cannot change it. The platform team adopts AWS Control Tower.
| 1 | The Workloads OU's controls govern payments-prod from the day it is made |
| 2 | payments-prod's API calls are delivered to the log-archive account |
| 3 | Drift and compliance alerts gather in the audit account |
As text
- AWS Control Tower
- Security OU
- AWS Account audit
- Amazon SNS (drift alerts)
- AWS Account log-archive
- Amazon S3 (every account's logs)
- AWS Account audit
- Workloads OU
- AWS Account payments-prod
- AWS CloudTrail (trail)
- AWS Account payments-prod
- The Workloads OU's controls govern payments-prod from the day it is made
- payments-prod's API calls are delivered to the log-archive account
- Drift and compliance alerts gather in the audit account
The Cloud Adoption Framework's Azure landing zone is a platform landing zone plus one application landing zone per workload: that workload's development, test and production subscriptions.
The landing zone is the whole multi-account environment, which Control Tower sets up and governs. A workload gets accounts in it, not a landing zone of its own.
Read more What is AWS Control Tower? · What is an Azure landing zone?
What it does. AWS Control Tower sets up and governs a landing zone: shared log-archive and audit accounts, controls on every OU, and a factory for new accounts.
How it works. It orchestrates Organizations, IAM Identity Center, Service Catalog, CloudTrail and Config. Controls are preventive (SCPs and RCPs), detective (Config rules) or proactive (CloudFormation hooks).
When to use it. Starting a multi-account estate, or bringing an existing organization under one baseline.
When not to. When another tool edits the same SCPs: Control Tower treats that as drift.
Limits. as of September 2026: up to 10,000 accounts, ten SCPs per OU, and a home Region that cannot be changed.
Cost. No additional charge; AWS Config, CloudTrail, Amazon S3 and the other services it uses are billed.
Security. Preventive controls do not bind the management account, so keep it nearly empty.
Gotchas. Moving an account between OUs or editing a managed SCP outside Control Tower is drift, and enrolling accounts stops until it is repaired.
Read more AWS Control Tower documentation · Control behavior and guidance · Detect and resolve drift in AWS Control Tower · How AWS Regions work with AWS Control Tower · Limitations and quotas in AWS Control Tower · AWS Control Tower pricing
Vending accounts#
New accounts come from Account Factory: a staging account requested in the Workloads OU arrives with the baseline and the OU's controls in force. Terraform teams use Account Factory for Terraform instead.
As text
- Payments team to Account Factory: a new account in the Workloads OU
- Account Factory to AWS Organizations: create the account in that OU
- AWS Organizations: the OU's controls apply at once
- Account Factory to payments-staging: the landing zone's baseline
- payments-staging to Payments team: ready
Read more Provision and manage accounts with Account Factory · Overview of Account Factory for Terraform (AFT)
MegaCorp's design gains its governance layer.
| 1 | Every API call in payments-prod is delivered to the log-archive account |
As text
- AWS Account payments-prod
- AWS CloudTrail (organization trail)
- Region eu-west-2
- AWS Account log-archive
- Amazon S3 (trail logs)
- Every API call in payments-prod is delivered to the log-archive account
- A developer's new role can do nothing, although its permissions boundary allows everything. Why?
Answer
A boundary only caps. The role also needs a permissions policy, and gets what both allow. - Two accounts both put their services in eu-west-2a to keep traffic inside one zone. Can the bill still show traffic between zones?
Answer
Yes: zone names map differently in each account. Compare zone IDs, such as euw2-az1. - Fourteen VPCs and the data centre must reach each other, but development must never reach production. What do you build?
Answer
A transit gateway shared through AWS RAM, with separate route tables for production and development.
Choosing compute#
AWS runs code on servers you manage, in containers it schedules, or as functions it runs per event. The less you manage, the less you control: choose the most managed option each workload can live with.
Three ways to run code#
MegaCorp's payments team has three workloads. The payments API is a Java service that runs all day, the nightly settlement job takes about 40 minutes, and a handler takes bursts of small webhook calls from the card schemes. Each could run in any of three models, which differ in who does the work of running it.
| Model | You manage | AWS manages | Services |
|---|---|---|---|
| Instances | the operating system, patching, scaling and security of each server | the hardware | Amazon EC2 with EC2 Auto Scaling |
| Containers | the image and the size of each task | the servers, with AWS Fargate, and the scheduling | Amazon ECS or Amazon EKS, on Fargate or EC2 |
| Functions | the code and its memory | servers, operating system, capacity, scaling and logging | AWS Lambda |
Read more Choosing an AWS compute service
Deciding, workload by workload#
The team asks the same questions of each workload, in order, and stops at the first yes.
As text
- Does it need a particular operating system, GPU, licence or agent on the host? Yes: Amazon EC2, in an Auto Scaling group. No: the next step.
- Is it short work triggered by events, each piece done within 15 minutes and keeping no state? Yes: AWS Lambda. No: the next step.
- Does the organization already run Kubernetes well? Yes: Amazon EKS. No: the next step.
- Amazon ECS on AWS Fargate
The webhook handler is short, event-driven work, so it goes to Lambda. The API runs all day and the settlement job runs for 40 minutes, too long for one Lambda invocation as of September 2026, so both go to ECS on Fargate. Only the fraud team's scoring model, which needs GPUs, lands on EC2.
As text
- AWS Account payments-prod
- Region eu-west-2
- AWS Fargate (payments API)
- AWS Fargate (settlement job)
- AWS Lambda (webhooks)
- Region eu-west-2
- AWS Account fraud
- Region eu-west-2
- Amazon EC2 (GPU scoring)
- Region eu-west-2
Kubernetes is a commitment, not a default. It releases three versions a year and retires old ones, so every cluster needs regular upgrades. Choose Amazon EKS when the skills and the need are already there.
Read more Choosing an AWS container service · Lambda quotas
How AWS got here#
Each step in AWS's compute history hid more of the server. Simpler platforms remain: AWS Elastic Beanstalk runs your code on EC2 that it manages, and Amazon Lightsail bundles servers, databases and networking at a predictable monthly price. AWS App Runner, the simplest path for web apps, is closed to new customers, and AWS points new applications to Amazon ECS Express Mode instead.
As text
- 2006: Amazon EC2: a server you run
- 2011: AWS Elastic Beanstalk: your code on managed EC2
- 2014: AWS Lambda: code, no servers
- 2015: Amazon ECS: containers on a managed cluster
- 2017: AWS Fargate: containers, no servers
Read more AWS App Runner availability change · What is Amazon Lightsail?
EC2 and Auto Scaling#
Amazon EC2 rents virtual servers by the second, and EC2 Auto Scaling keeps a healthy number of them running. Use them when a workload needs the server itself, and put even one inside a group.
One instance#
The fraud team's scoring model needs GPUs and a driver installed on the host, so it runs on EC2. An instance boots from an Amazon Machine Image, which holds its operating system and software, and its instance type fixes the CPU, memory and accelerators. In c7gn.xlarge, c means compute optimized, 7 the generation, g Graviton, n extra networking, and xlarge the size.
What it does. Amazon EC2 rents virtual servers, called instances, by the second. The fraud team's model runs on GPU instances with the drivers it needs.
How it works. An instance boots from an AMI into a subnet, with an instance type that sets its hardware and an IAM role for its credentials. Everything above the hypervisor is yours to run.
When to use it. A particular operating system, GPU, licence or host agent, or software that expects a whole server.
When not to. Work that a container on Fargate or a Lambda function can do: every instance is yours to patch, scale and secure.
Limits. as of September 2026: On-Demand and Spot capacity is capped per Region in vCPUs, by instance family; request increases well before a launch.
Cost. On-Demand is billed per second. Savings Plans lower the rate for a one- or three-year hourly commitment, and Spot sells spare capacity at a discount. Volumes and data transfer are billed apart.
Security. Require IMDSv2, which is optional by default, and reach instances through Session Manager: no inbound ports, bastion hosts or SSH keys.
Gotchas. An AMI belongs to one Region, so copy it before launching in another. A Spot Instance gets two minutes' notice before it is taken back.
Read more Amazon EC2 documentation · Amazon EC2 instance type naming conventions · Amazon Machine Images in Amazon EC2 · Amazon EC2 service quotas
The metadata service hands out the role's credentials. Applications on an instance get its role's temporary credentials from instance metadata. IMDSv2 guards that door with session tokens, AWS's defence against server-side request forgery bugs, but by default an instance also accepts IMDSv1. Require IMDSv2 on every instance.
A fleet: Auto Scaling groups#
One instance is one failure away from an outage, and scoring load triples on sale days. The team runs the instances in an Auto Scaling group across two zones, behind a load balancer: at least two, at most eight, and replaced when a health check fails.
| 1 | The load balancer sends requests to healthy instances in both zones |
| 2 | The group replaces an instance that fails its health check |
As text
- AWS Account fraud
- Region eu-west-2
- VPC 10.60.0.0/16
- Application Load Balancer (scoring)
- Amazon EC2 Auto Scaling (scoring group)
- Availability Zone a
- Private subnet 10.60.10.0/24
- Amazon EC2 (GPU instance)
- Private subnet 10.60.10.0/24
- Availability Zone b
- Private subnet 10.60.11.0/24
- Amazon EC2 (GPU instance)
- Private subnet 10.60.11.0/24
- VPC 10.60.0.0/16
- Region eu-west-2
- The load balancer sends requests to healthy instances in both zones
- The group replaces an instance that fails its health check
What it does. EC2 Auto Scaling keeps the right number of healthy instances running: the scoring group never drops below two, and grows to eight under load.
How it works. A group launches instances from a launch template across chosen zones, balancing them evenly, registers them with a load balancer, and replaces any that fail a health check. Policies and schedules change the desired count.
When to use it. Every EC2 workload, even a single instance: a group of one still replaces it when it fails.
When not to. Containers on Fargate or functions on Lambda, where AWS runs the capacity.
Limits. as of September 2026: 500 groups per Region, and 50 scaling policies and 125 scheduled actions per group.
Cost. No additional fees; you pay for the instances, volumes and alarms it uses.
Security. Put the IAM role and the IMDSv2 requirement in the launch template, so every new instance starts with them.
Gotchas. Instances come and go, so keep nothing on them that must survive: logs, sessions and files belong in services. A lifecycle hook gives a terminating instance time to finish.
Read more Amazon EC2 Auto Scaling documentation · What is Amazon EC2 Auto Scaling? · Quotas for Auto Scaling resources and groups
Paying for instances#
The scoring fleet runs all year, while the nightly retraining job can stop and restart. The team pays for each differently.
As text
- Can the work stop at two minutes' notice and pick up again? Yes: Spot Instances. No: the next step.
- Will this much compute run steadily for a year or more? Yes: A Savings Plan for the steady part. No: the next step.
- Must capacity be certain in one zone, such as for a failover? Yes: A Capacity Reservation. No: the next step.
- On-Demand, by the second
Retraining goes to Spot. A Savings Plan covers the fleet's two steady instances, and the sale-day bursts run On-Demand.
Read more Amazon EC2 billing and purchasing options · Spot Instance interruption notices · Use the Instance Metadata Service to access instance metadata · AWS Systems Manager Session Manager
How EC2 grew#
As text
- 2006: Amazon EC2: a server you rent
- 2009: Auto Scaling: fleets that follow demand
- 2009: Spot Instances: spare capacity, cheaper
- 2018: AWS Graviton: Arm-based instances
- 2019: Savings Plans: commit to spend, not to a type
Read more Amazon EC2 documentation
Containers#
A container image is built once and runs anywhere. On AWS, Amazon ECR stores the images, Amazon ECS or Amazon EKS runs them, and AWS Fargate supplies the capacity.
Images, and where they live#
The pipeline builds the payments API into a container image on every merge and pushes it to a private Amazon ECR repository; the same image then runs in development and in production.
What it does. Amazon ECR stores the payments team's images privately, beside the services that run them.
How it works. Repositories hold images and a resource policy, and pushes and pulls use IAM. ECR can scan on push, replicate across Regions and accounts, cache public images and expire old ones.
When to use it. Every image that ECS, EKS or Lambda runs.
When not to. Pulling public images from the internet in production: cache them in ECR.
Limits. Repositories are Regional, so replicate images to each Region that runs them.
Cost. Storage, data transfer for pushes and pulls, and opted-in actions such as replication.
Security. Scan on push, make tags immutable, and grant other accounts pull access by policy.
Gotchas. Private tasks pull images through the NAT gateway, and pay per gigabyte, unless the VPC has ECR endpoints.
Read more Amazon ECR documentation · What is Amazon Elastic Container Registry?
| 1 | The pipeline pushes the image once, tagged with its commit |
| 2 | Development pulls it, as the repository's policy allows |
| 3 | Production pulls the same image once the tests pass |
As text
- AWS Account tooling
- Region eu-west-2
- AWS CodePipeline (payments)
- Amazon ECR (payments-api)
- Region eu-west-2
- Workloads OU
- AWS Account payments-dev
- Region eu-west-2
- AWS Fargate (API tasks)
- Region eu-west-2
- AWS Account payments-prod
- Region eu-west-2
- AWS Fargate (API tasks)
- Region eu-west-2
- AWS Account payments-dev
- The pipeline pushes the image once, tagged with its commit
- Development pulls it, as the repository's policy allows
- Production pulls the same image once the tests pass
Running containers with ECS#
The API runs as an ECS service of at least two tasks on Fargate behind a load balancer; the settlement job is a task that runs nightly and stops.
What it does. Amazon ECS runs and scales containers with no control plane to operate.
How it works. A task definition names the containers, their size and two roles: the task role the code calls AWS with, and the execution role ECS uses to pull images and write logs. A service keeps tasks running behind a load balancer.
When to use it. Most container workloads, when nobody needs Kubernetes itself.
When not to. An organization committed to Kubernetes tooling: that is EKS.
Limits. as of September 2026: 5,000 services per cluster and 5,000 tasks per service.
Cost. Only the capacity the tasks use, on Fargate or EC2.
Security. Give each service its own task role, with least privilege.
Gotchas. ECS Express Mode, which builds a web service from one call, puts an internet-facing load balancer in the default VPC unless you give it subnets.
Read more Amazon ECS documentation · What is Amazon Elastic Container Service? · Amazon ECS task IAM role · Create your first Express Mode service using the AWS CLI · Amazon ECS endpoints and quotas
A task has two roles, and they are easily swapped. The task role is what the payments code calls AWS with; the execution role is what ECS uses to pull the image and write logs. Grant S3 to the execution role and the code is still refused, and a failed image pull points to the execution role, not the task role.
What it does. AWS Fargate runs ECS tasks and EKS pods on capacity AWS manages: no instances to patch.
How it works. You set CPU and memory per task. Each task gets its own kernel and network interface, on the latest patched platform.
When to use it. The default capacity for containers, and wherever tasks must be isolated.
When not to. GPUs, privileged containers or tasks above 32 vCPU.
Limits. as of September 2026: 0.25 to 32 vCPU per task; new accounts start with low vCPU quotas.
Cost. Per second for the vCPU, memory and storage requested. Fargate Spot is cheaper but can be interrupted.
Security. No task can reach another task's credentials.
Gotchas. AWS retires tasks on a platform revision with a security issue, so a service must tolerate any task being replaced.
Read more AWS Fargate for Amazon ECS · Architect for AWS Fargate for Amazon ECS · Amazon ECS task definition differences for Fargate · AWS Fargate pricing
ECS or EKS#
The risk team already runs Kubernetes in the data centre, so its model platform moves to Amazon EKS; the payments team, new to containers, stays on ECS.
As text
- Does the organization already run Kubernetes, or need its tools and portability? Yes: Amazon EKS. No: the next step.
- Does it need GPUs or particular instance types? Yes: ECS on ECS Managed Instances. No: the next step.
- Amazon ECS on AWS Fargate
What it does. Amazon EKS runs the Kubernetes control plane, so the risk team keeps its tooling without running masters.
How it works. Nodes come from EKS Auto Mode, which AWS also runs, from managed node groups on EC2, or from Fargate. Pods get AWS credentials through EKS Pod Identity.
When to use it. Kubernetes skills, tooling or portability needs already in place.
When not to. A team new to containers: ECS does the job with far less to learn and upgrade.
Limits. as of September 2026: each version has 14 months of standard support and 12 of extended support, then an automatic upgrade.
Cost. Per cluster per hour, higher in extended support, plus the nodes and any Auto Mode fee.
Security. Give each service account its own role through Pod Identity, and restrict instance metadata on nodes.
Gotchas. A control plane upgrade leaves managed and self-managed nodes behind: upgrade them too.
Read more Amazon EKS documentation · What is Amazon EKS? · Understand the Kubernetes version lifecycle on EKS · Amazon EKS pricing · Learn how EKS Pod Identity grants pods access to AWS services
Read more Choosing an AWS container service
How AWS got here#
As text
- 2015: Amazon ECS: containers on your EC2 cluster
- 2017: AWS Fargate: no cluster to run
- 2018: Amazon EKS: managed Kubernetes
- 2024: EKS Auto Mode: AWS runs the nodes too
- 2025: ECS Express Mode: a service from one call
MegaCorp's design gains its compute layer.
| 1 | Tasks reach the internet through the NAT gateway in their own zone |
| 2 | The NAT gateway sends traffic out through the internet gateway |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- AWS Fargate (payments)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- AWS Fargate (payments)
- Public subnet 10.20.1.0/24
- VPC 10.20.0.0/16
- Region eu-west-2
- Tasks reach the internet through the NAT gateway in their own zone
- The NAT gateway sends traffic out through the internet gateway
Read more Amazon ECS documentation
Serverless#
AWS Lambda runs a function for each event and bills by the millisecond; AWS Step Functions joins functions and services into workflows that retry, wait and remember. Both remove servers, not the limits you design around.
Functions that run per event#
Card-scheme webhooks land on an Amazon SQS queue, and a Java function settles each message. Lambda reads the queue, hands the function batches of messages, and runs as many copies as the queue needs, up to the account's concurrency.
public class SettlementHandler implements RequestHandler<SQSEvent, SQSBatchResponse> {
private final Settlements settlements;
public SettlementHandler() {
this(new Settlements());
}
SettlementHandler(Settlements settlements) {
this.settlements = settlements;
}
@Override
public SQSBatchResponse handleRequest(SQSEvent event, Context context) {
List<SQSBatchResponse.BatchItemFailure> failures = new ArrayList<>();
for (SQSEvent.SQSMessage message : event.getRecords()) {
try {
settlements.settle(message.getMessageId(), message.getBody());
} catch (IllegalArgumentException e) {
failures.add(new SQSBatchResponse.BatchItemFailure(message.getMessageId()));
}
}
return new SQSBatchResponse(failures);
}
}The handler reports only the messages that failed, so only those are retried. The team reserves concurrency for it, so no other function can starve it and it cannot flood the ledger database.
What it does. AWS Lambda runs a function for each event, from a queue, a bucket, an API or a schedule, and bills only while it runs.
How it works. Each concurrent request gets its own execution environment, created on demand and then reused. The first request to a new one waits for initialization: the cold start. Functions call AWS with an execution role.
When to use it. Short, event-driven work: webhooks, queue consumers, file processing, glue between services.
When not to. Work over 15 minutes, or steady heavy load that suits containers or Lambda Managed Instances.
Limits. as of September 2026: 15 minutes per invocation, up to 10,240 MB of memory, 6 MB synchronous payloads, and 1,000 concurrent executions per Region, shared by every function.
Cost. Per request and per GB-second, with a monthly free tier. Provisioned concurrency is billed while configured; SnapStart for Java costs nothing extra.
Security. One execution role per function, with least privilege. A function in a VPC reaches the internet only through a NAT gateway, even in a public subnet.
Gotchas. Events can arrive more than once, and a failed asynchronous event is retried twice: make handlers idempotent, and catch what still fails in a dead-letter queue.
Read more AWS Lambda documentation · Lambda quotas · Understanding Lambda function scaling · AWS Lambda pricing · Giving Lambda functions access to resources in an Amazon VPC · Handling errors for an SQS event source in Lambda
When a statement file lands in Amazon S3, S3 invokes an indexing function asynchronously, and Lambda owns the retries.
As text
- Amazon S3 to Lambda: a new file, queued
- Lambda to Indexing function: invoke
- Indexing function to Lambda: error
- Lambda: two more attempts, one minute then two minutes apart
- Lambda to Dead-letter queue: the event that still failed
Read more How Lambda handles errors and retries with asynchronous invocation
Java and cold starts#
A Java function initializes slowly: the JVM starts and frameworks load before the first request. SnapStart snapshots the initialized function when a version is published and resumes new environments from it, at no extra cost for Java. Provisioned concurrency keeps environments warm, for a fee.
SnapStart copies one initialized state into many environments. Anything unique made during initialization, such as a random seed or an ID, is duplicated: create it after restore.
Read more Improving startup performance with Lambda SnapStart · Building Lambda functions with Java
Workflows#
Settling a merchant's day takes several steps: validate the file, post to the ledger, notify the merchant, and undo if a step fails. The team models it as a Step Functions state machine instead of functions that call each other.
| 1 | An EventBridge rule starts an execution for each new file |
| 2 | The file is validated first, with retries |
| 3 | Posting runs exactly once, in a Standard workflow |
As text
- Region eu-west-2
- Amazon S3 (settlement files)
- AWS Step Functions (settlement)
- steps
- AWS Lambda (validate)
- AWS Lambda (post to ledger)
- An EventBridge rule starts an execution for each new file
- The file is validated first, with retries
- Posting runs exactly once, in a Standard workflow
What it does. AWS Step Functions runs workflows as state machines: steps, choices, retries, waits and parallel branches, with each execution's progress recorded.
How it works. A workflow is written in Amazon States Language and calls Lambda or over 220 other services directly. Standard workflows run up to a year, exactly once; Express workflows run up to five minutes, at least once.
When to use it. Multi-step processes that retry, wait for people or roll back, and must be auditable afterwards.
When not to. A single step, or logic the team would rather keep in Java: Lambda durable functions do that.
Limits. as of September 2026: 25,000 history events per Standard execution and 256 KiB per input or output. The workflow type cannot be changed later.
Cost. Standard is billed per state transition, retries included; Express per execution, duration and memory.
Security. Give each state machine its own role, allowed to call only its steps, and log executions to CloudWatch Logs.
Gotchas. Express workflows may run a step twice, so keep them to idempotent work; payments belong in Standard.
Read more AWS Step Functions documentation · Choosing workflow type in Step Functions · Step Functions service quotas · AWS Step Functions pricing
As text
- Does the work finish in one function call within 15 minutes? Yes: One Lambda function. No: the next step.
- Would the team rather write the workflow in Java, inside Lambda? Yes: Lambda durable functions. No: the next step.
- Is it high volume, under five minutes, and safe to repeat? Yes: A Step Functions Express workflow. No: the next step.
- A Step Functions Standard workflow
Read more Lambda durable functions · Lambda Managed Instances
Object and file storage#
Amazon S3 keeps objects, Amazon EBS gives one instance a disk, and Amazon EFS and Amazon FSx share files. Choose by how data is reached, not how much there is.
As text
- Is it a disk for one instance, such as a boot volume or a database? Yes: Amazon EBS. No: the next step.
- Must many Linux clients share files over NFS? Yes: Amazon EFS. No: the next step.
- Must Windows clients share files over SMB, or does it need Lustre, ONTAP or ZFS? Yes: Amazon FSx. No: the next step.
- Amazon S3, over its API
Read more Choosing an AWS storage service
Objects: Amazon S3#
Every merchant statement is a PDF in an S3 bucket, kept for seven years because the regulator says so.
What it does. Amazon S3 stores objects, written whole and read by key, with strong read-after-write consistency.
How it works. Each object has a storage class; lifecycle rules move it colder or delete it, and versioning keeps old copies.
When to use it. Documents, backups, logs, data lakes and static content.
When not to. Files edited in place: use a file system, or mount the bucket with S3 Files.
Limits. as of September 2026: 10,000 buckets per account by default; a bucket's name and Region never change.
Cost. Storage by class, requests, retrievals and data transfer out.
Security. New buckets are private, block public access and encrypt every object by default.
Gotchas. Colder classes bill a minimum of 30 to 180 days, even if you delete sooner.
Read more Amazon S3 documentation · Understanding and managing Amazon S3 storage classes · Locking objects with Object Lock
As text
- S3 Standard to Instant Retrieval: after 90 days
- Instant Retrieval to Deep Archive: after a year
- Deep Archive: deleted after seven years
Object Lock can lock you out too. It makes objects write-once, and in compliance mode no one, the root user included, can delete one before its date. Test in governance mode first.
Disks: Amazon EBS#
The fraud team's GPU instances boot from EBS volumes and keep model files on a second gp3 volume.
What it does. Amazon EBS gives EC2 instances durable network disks.
How it works. A volume lives in one zone, replicated within it. gp3 sets size, IOPS and throughput separately; snapshots copy data to other zones and Regions.
When to use it. An instance's own disk: boot, database or working data.
When not to. Data that several instances must share.
Limits. as of September 2026: gp3 up to 64 TiB and 80,000 IOPS; io2 Block Express up to 256,000 IOPS.
Cost. What you provision, plus snapshot storage.
Security. Encryption by default is set per Region and skips existing volumes: turn it on early.
Gotchas. A volume cannot move zone: restore a snapshot in the new one.
Read more Amazon EBS documentation · Amazon EBS volume types · Enable Amazon EBS encryption by default
Shared files: EFS and FSx#
Settlement tasks in both zones share a directory of bank files on EFS.
| 1 | Zone a mounts the file system |
| 2 | Zone b mounts the same one |
As text
- Region eu-west-2
- VPC 10.20.0.0/16
- Amazon EFS (bank files)
- Availability Zone a
- Private subnet 10.20.10.0/24
- AWS Fargate (settlement)
- Private subnet 10.20.10.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- AWS Fargate (settlement)
- Private subnet 10.20.11.0/24
- VPC 10.20.0.0/16
- Zone a mounts the file system
- Zone b mounts the same one
What it does. Amazon EFS is shared NFS storage that grows and shrinks with its files.
How it works. A Regional file system spans zones, so clients in any zone mount it; lifecycle policies move cooling files to Infrequent Access and Archive classes.
When to use it. Linux instances, containers and functions that share files.
When not to. Windows clients, which EFS does not support.
Limits. A One Zone file system can be lost with its zone.
Cost. Storage by class, plus reads, writes and tiering.
Security. Encrypt in transit when mounting; control access with IAM, security groups and POSIX permissions.
Gotchas. Encryption at rest can only be chosen when the file system is created.
Read more Amazon EFS documentation · Amazon EFS pricing
The risk team's Windows reports read an SMB share on Amazon FSx for Windows File Server. FSx also runs NetApp ONTAP, OpenZFS and Lustre file systems. You provision storage, IOPS and throughput; a Windows file system is Single-AZ or Multi-AZ, and joins Active Directory when you create it.
Read more What is FSx for Windows File Server?
Relational databases#
Amazon RDS runs six familiar database engines for you; Amazon Aurora rebuilds MySQL and PostgreSQL on storage shared across three zones. Either way, the schema, the queries and their tuning stay yours.
The ledger is the payments team's most precious data: every posting, in PostgreSQL. It moves to Aurora PostgreSQL. The risk team's reports run on SQL Server, which Aurora does not offer, so that database moves to RDS for SQL Server.
As text
- Must it run SQL Server, Oracle, Db2 or MariaDB, or community PostgreSQL or MySQL? Yes: Amazon RDS. No: the next step.
- Is the load spiky, or idle for hours at a time? Yes: Aurora Serverless v2. No: the next step.
- Amazon Aurora with provisioned instances
Read more Choosing an AWS database service
Amazon RDS#
What it does. Amazon RDS runs Db2, MariaDB, SQL Server, MySQL, Oracle or PostgreSQL, and handles backups, patching and failover.
How it works. Pick an engine, instance class and storage. A Multi-AZ standby in another zone takes over on failure but serves no reads; read replicas do, copied asynchronously.
When to use it. Commercial engines, and community PostgreSQL or MySQL.
When not to. Spiky or mostly idle load: an instance bills for every hour it runs.
Limits. as of September 2026: 40 instances per Region, shared with Aurora; 15 read replicas each; backups kept up to 35 days.
Cost. Instance hours, storage, IOPS, backups and data transfer; each replica bills as an instance.
Security. Encryption can only be chosen at creation. Let RDS keep the master password in Secrets Manager, rotated every seven days.
Gotchas. A version past its standard support moves to paid Extended Support: plan upgrades.
Read more Amazon RDS documentation · Configuring and managing a Multi-AZ deployment for Amazon RDS · Amazon RDS Extended Support with Amazon RDS
Amazon Aurora#
What it does. Amazon Aurora is a MySQL- and PostgreSQL-compatible engine whose instances share one storage volume.
How it works. A writer and up to 15 readers share a volume kept on six storage nodes across three zones. The cluster endpoint always names the writer.
When to use it. PostgreSQL or MySQL systems needing fast failover, read scaling or a copy in another Region.
When not to. Other engines, which need RDS; writes in several Regions at once, which suit Aurora DSQL.
Limits. as of September 2026: a 256 TiB volume; a global database copies to up to 10 more Regions, typically under a second behind.
Cost. Instance hours or Serverless v2 capacity, storage used, and I/O unless the cluster is I/O-Optimized.
Security. Use IAM database authentication: tokens that last 15 minutes instead of stored passwords.
Gotchas. With no reader, a failed writer is recreated, typically within 10 minutes: keep a reader in another zone.
Read more Amazon Aurora User Guide · High availability for Amazon Aurora · Using Aurora serverless
As text
- Payments API to Writer (zone a): writes, through the cluster endpoint
- Writer (zone a): fails
- Payments API to Reader (zone b): reconnects to the same endpoint, now this instance
Failover changes DNS, and clients cache DNS. A client that keeps the old address keeps calling the failed instance. RDS Proxy bypasses DNS caches, cutting Aurora failover time by up to 66%, and pools connections so a surge of clients cannot overwhelm the database.
MegaCorp's design gains its data layer: the ledger's writer and reader sit in the private subnets of both zones.
| 1 | Tasks reach the internet through the NAT gateway in their own zone |
| 2 | The NAT gateway sends traffic out through the internet gateway |
| 3 | Payments tasks write the ledger through the cluster endpoint |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- AWS Fargate (payments)
- Amazon Aurora (ledger writer)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- AWS Fargate (payments)
- Amazon Aurora (ledger reader)
- Public subnet 10.20.1.0/24
- Amazon S3 (statements)
- VPC 10.20.0.0/16
- Region eu-west-2
- Tasks reach the internet through the NAT gateway in their own zone
- The NAT gateway sends traffic out through the internet gateway
- Payments tasks write the ledger through the cluster endpoint
Read more Using Amazon Aurora Global Database · Amazon RDS Proxy
NoSQL and caching#
Amazon DynamoDB trades joins for predictable speed at any size, and Amazon ElastiCache keeps hot data in memory. Both reward knowing how data will be read before you store it.
In the 2004 holiday season, outages at Amazon were traced to commercial technology pushed past its limits. Amazon built Dynamo, described it in a 2007 paper, and found it still hard to run. DynamoDB, launched in 2012, put Dynamo's scaling behind a service AWS operates.
As text
- Is it a copy kept only to make reads faster? Yes: Amazon ElastiCache. No: the next step.
- Is every way of reading it known, and done by key? Yes: Amazon DynamoDB. No: the next step.
- A relational database: Aurora or RDS
Read more Choosing an AWS database service
Amazon DynamoDB#
Card schemes sometimes send a webhook twice. The settlement function writes each event ID to a DynamoDB table only if it is absent, so a repeat fails the condition and is skipped. Time to Live removes old IDs.
As text
- Card scheme to Settlement function: webhook for event 81
- Settlement function to Idempotency table: put 81, only if absent
- Card scheme to Settlement function: the same webhook again
- Settlement function to Idempotency table: put 81: condition fails, skip
What it does. Amazon DynamoDB stores items by key and answers in single-digit milliseconds at any scale, with no servers or versions to manage.
How it works. A partition key is hashed to choose a partition, and a sort key orders items within it. Indexes add other keys. Data is copied across three zones.
When to use it. Known access by key at high or unpredictable volume: sessions, idempotency keys, carts.
When not to. Ad hoc queries and joins, which DynamoDB does not do.
Limits. as of September 2026: 400 KB per item; each partition serves 3,000 reads and 1,000 writes per second.
Cost. On-demand per request, the default, or provisioned per hour; plus storage and backups.
Security. Every call is authorized by IAM, with no passwords; data is encrypted at rest by default.
Gotchas. Reads are eventually consistent unless you ask for strong ones, which global secondary indexes cannot give.
Read more Amazon DynamoDB documentation · Core components of Amazon DynamoDB · DynamoDB read consistency
On-demand still has a ceiling on day one. A new on-demand table sustains 4,000 writes and 12,000 reads per second, and absorbs double its previous peak; grow faster within 30 minutes and requests are throttled. Pre-warm before a big launch.
Read more DynamoDB on-demand capacity mode
Amazon ElastiCache#
The merchant portal reads each merchant's settings on every page. The team caches them in Valkey, the open-source fork of Redis made in 2024: read the cache first, fall back to Aurora on a miss, and store the answer for five minutes.
What it does. Amazon ElastiCache runs Valkey, Memcached or Redis OSS as a managed in-memory cache.
How it works. Serverless caches scale themselves; node-based clusters are sized by node type and count, with replicas in other zones that take over in seconds.
When to use it. Hot reads in front of a database, session stores and leaderboards.
When not to. The only copy of important data: replication is asynchronous, so failover can lose recent writes.
Limits. as of September 2026: 40 serverless caches and 300 nodes per Region; Memcached has no replication.
Cost. Serverless by data stored and processing units, nodes by the hour; Valkey costs less than Redis OSS.
Security. Keep caches in private subnets, encrypt in transit, and authenticate with role-based access control.
Gotchas. Cached data goes stale when the database changes: give every key a time to live.
Read more Amazon ElastiCache documentation · Caching strategies
MegaCorp's design gains a NoSQL layer.
| 1 | Tasks reach the internet through the NAT gateway in their own zone |
| 2 | The NAT gateway sends traffic out through the internet gateway |
| 3 | Tasks read merchant settings from the cache before Aurora |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Amazon ElastiCache (settings)
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- AWS Fargate (payments)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- AWS Fargate (payments)
- Public subnet 10.20.1.0/24
- Amazon DynamoDB (idempotency keys)
- VPC 10.20.0.0/16
- Region eu-west-2
- Tasks reach the internet through the NAT gateway in their own zone
- The NAT gateway sends traffic out through the internet gateway
- Tasks read merchant settings from the cache before Aurora
Messaging and events#
Queues hold work, topics copy messages, event buses route events, and streams keep an ordered record to replay. Choose by who needs a message, and for how long.
When a payment settles, the merchant wants a notice, statements need a line, and fraud wants every card event. The payments service publishes what happened, and each consumer subscribes.
As text
- Must readers replay an ordered record, at high volume? Yes: Amazon Kinesis Data Streams. No: the next step.
- Do many services react to events, chosen by content? Yes: Amazon EventBridge. No: the next step.
- Must every subscriber get a copy pushed to it, people included? Yes: Amazon SNS. No: the next step.
- Amazon SQS, for work one consumer does at its own pace
Read more Amazon SQS, Amazon SNS, or Amazon EventBridge?
Queues: Amazon SQS#
What it does. Amazon SQS holds messages until a consumer takes and deletes them: the buffer between producer and worker.
How it works. A received message hides for the visibility timeout and returns if it is not deleted. After several failed receives it moves to a dead-letter queue.
When to use it. Work one consumer does at its own pace, absorbing spikes.
When not to. Many consumers that each need a copy: put SNS or EventBridge in front.
Limits. as of September 2026: 1 MiB messages, kept up to 14 days; FIFO queues keep order within a message group.
Cost. Per request, each 64 KB counting as one; the first million each month are free.
Security. IAM and a queue policy decide who may send and receive; encryption uses KMS.
Gotchas. Standard queues deliver at least once, sometimes out of order: make consumers idempotent.
Read more Amazon SQS documentation · Using dead-letter queues in Amazon SQS
Dead letters keep their birthday. A message moved to a dead-letter queue keeps its original enqueue time, so it expires early unless the dead-letter queue keeps messages longer than the source queue.
Topics and buses: SNS and EventBridge#
What it does. Amazon SNS pushes each message published to a topic to every subscriber: queues, functions, HTTPS endpoints, email or SMS.
How it works. Subscribers can filter on attributes or body; FIFO topics keep order into SQS FIFO queues.
When to use it. Fan-out to a few known consumers, and notifications to people.
When not to. Routing on many rules across many services: EventBridge filters more richly.
Limits. as of September 2026: 256 KiB messages; in eu-west-2, 300 publishes a second per account by default.
Cost. Per publish and per delivery; delivery to SQS and Lambda is not charged per message.
Security. IAM and topic policies decide who publishes and who subscribes.
Gotchas. Messages are not kept: subscribe a queue for anything that must not be lost.
Read more Amazon SNS documentation · Amazon SNS message filtering
What it does. Amazon EventBridge routes events from AWS services, your applications and SaaS partners to targets, by rules that match their content.
How it works. Rules on an event bus match event patterns and send each match to up to five targets. Pipes join one source to one target; Scheduler runs timed tasks.
When to use it. Events that many teams react to, across services and accounts.
When not to. Processing that depends on order, which EventBridge does not guarantee.
Limits. as of September 2026, in eu-west-2: 1,200 PutEvents requests a second and 300 rules per bus.
Cost. Events from AWS services on the default bus are free; custom events are charged per million.
Security. A resource policy on a bus lets other accounts send to it.
Gotchas. Delivery is at least once, retried for up to 24 hours: consumers must be idempotent.
Read more Amazon EventBridge documentation · Amazon EventBridge quotas
Streams: Kinesis Data Streams#
What it does. Amazon Kinesis Data Streams keeps an ordered, replayable stream of records that many applications read independently.
How it works. A partition key picks a shard, and order holds within it; each shard takes 1 MB or 1,000 records a second.
When to use it. Clickstreams, telemetry and card events read by several consumers in near real time.
When not to. Work items that each need one consumer: that is a queue.
Limits. as of September 2026: records kept 24 hours by default, up to 365 days.
Cost. On-demand per GB written and read, or provisioned per shard-hour; longer retention costs more.
Security. IAM for producers and consumers, and server-side encryption with KMS.
Gotchas. A hot partition key is capped at one shard's limit, even on demand: choose keys that spread.
Read more Amazon Kinesis Data Streams documentation · What is Amazon Kinesis Data Streams?
Teams already running Apache Kafka can use Amazon MSK instead: the same Kafka APIs and tools, with AWS running the brokers.
As text
- 2006: Amazon SQS: queues, after a 2004 beta
- 2010: Amazon SNS: publish once, deliver to many
- 2013: Amazon Kinesis: streams of data in real time
- 2019: EventBridge, from CloudWatch Events
- 2022: EventBridge Scheduler and Pipes
MegaCorp's design gains its messaging layer.
| 1 | The NAT gateway sends traffic out through the internet gateway |
| 2 | A rule sends settled payments to the merchant alerts topic |
As text
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Public subnet 10.20.1.0/24
- Amazon SQS (webhooks)
- Amazon EventBridge (payments)
- Amazon SNS (merchant alerts)
- Amazon Kinesis Data Streams (card events)
- VPC 10.20.0.0/16
- Region eu-west-2
- The NAT gateway sends traffic out through the internet gateway
- A rule sends settled payments to the merchant alerts topic
APIs and the edge#
Every outside request meets a front door: a load balancer, an API gateway or a content delivery network, with a firewall in front. Choose by what comes through.
As text
- Is it web content or an API used far from its Region? Yes: Amazon CloudFront, in front of the rest. No: the next step.
- Is it an API that needs keys, quotas or validation? Yes: Amazon API Gateway. No: the next step.
- A load balancer: ALB for HTTP; NLB for TCP, UDP or fixed IPs
Read more What is Elastic Load Balancing?
Load balancers#
What it does. Elastic Load Balancing spreads traffic across healthy targets in several zones, and scales itself.
How it works. An ALB routes HTTP by host, path or header; an NLB passes TCP and UDP at millions of requests a second.
When to use it. Web services and containers (ALB); other protocols or fixed IPs (NLB).
When not to. Classic Load Balancers, the previous generation: migrate them.
Limits. An NLB keeps each zone's traffic in that zone unless cross-zone balancing is on.
Cost. Hours running, plus capacity units for connections, bytes and rule evaluations.
Security. TLS ends at the balancer with ACM certificates; an ALB can take AWS WAF.
Gotchas. An NLB's targets see the client's IP address only for some target types.
Read more Elastic Load Balancing documentation · What is an Application Load Balancer?
APIs: Amazon API Gateway#
What it does. Amazon API Gateway fronts Lambda, AWS services or private load balancers, throttling and authorizing each call.
How it works. REST APIs add keys, usage plans, validation, caching and WAF; HTTP APIs drop them for a lower price.
When to use it. Partner APIs such as the card-scheme webhooks, which need keys and quotas, and Lambda backends.
When not to. Plain web traffic to containers: an ALB is enough.
Limits. as of September 2026: 10 MB payloads; 10,000 requests a second per account and Region, shared by every API.
Cost. Per million calls, HTTP APIs costing less; a cache bills by the hour.
Security. IAM, Cognito or Lambda authorizers; a REST API can be private to your VPCs.
Gotchas. One busy API can use up the shared throttle: give each client a usage plan.
Read more Amazon API Gateway documentation · Choose between REST APIs and HTTP APIs
An integration gets 29 seconds. After that, API Gateway gives up even if the backend finishes. Regional REST APIs can wait longer at a cost to the account throttle; long work belongs on a queue.
The edge: CloudFront and AWS WAF#
What it does. Amazon CloudFront serves content from edge locations near users, caching what it can.
How it works. A distribution maps paths to origins such as S3, load balancers or API Gateway; objects stay cached 24 hours by default.
When to use it. Websites, downloads and APIs used far from their Region.
When not to. Non-HTTP traffic or fixed IPs: that is AWS Global Accelerator.
Limits. Edge code: CloudFront Functions for sub-millisecond work, Lambda@Edge for longer work.
Cost. Data out and requests, or a flat-rate plan; transfer from AWS origins is free.
Security. Keep S3 origins private with origin access control, and attach AWS WAF.
Gotchas. An S3 website endpoint cannot use origin access control.
Read more Amazon CloudFront documentation · Restrict access to an Amazon S3 origin
What it does. AWS WAF inspects HTTP requests to CloudFront, ALBs and API Gateway REST APIs, and allows, blocks or counts them.
How it works. A web ACL holds rules on IPs, countries, headers, SQL injection or request rates, plus managed rule groups.
When to use it. Every public web entry point.
When not to. Non-HTTP traffic: use Network Firewall and security groups.
Limits. HTTP APIs cannot take a web ACL: put CloudFront in front of them.
Cost. Per web ACL and rule each month, and per million requests; bot rules cost extra.
Security. AWS Firewall Manager applies the same web ACLs across accounts.
Gotchas. Managed rules can block real customers: run them in Count mode first.
Read more AWS WAF documentation · What is AWS WAF?
AWS Shield Standard guards every AWS customer against common DDoS attacks at no extra charge; Shield Advanced is paid.
As text
- 2009: Elastic Load Balancing
- 2015: Amazon API Gateway
- 2015: AWS WAF, first for CloudFront
- 2016: Application Load Balancer
- 2018: AWS Global Accelerator
MegaCorp's design gains its front door.
| 1 | The NAT gateway sends traffic out through the internet gateway |
| 2 | Requests meet the web ACL first |
| 3 | CloudFront serves the requests it allows |
| 4 | Dynamic requests go to the load balancer |
As text
- Customers
- AWS WAF (web ACL)
- Amazon CloudFront (portal)
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Application Load Balancer (payments)
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- Public subnet 10.20.0.0/24
- Availability Zone b
- Public subnet 10.20.1.0/24
- NAT gateway
- Private subnet 10.20.11.0/24
- Public subnet 10.20.1.0/24
- VPC 10.20.0.0/16
- Region eu-west-2
- The NAT gateway sends traffic out through the internet gateway
- Requests meet the web ACL first
- CloudFront serves the requests it allows
- Dynamic requests go to the load balancer
Data and analytics#
Analytics on AWS starts with files in S3: AWS Glue reshapes and catalogues them, AWS Lake Formation grants access, and Amazon Athena or Amazon Redshift answers the questions.
As text
- Do dashboards run the same queries over curated data all day? Yes: Amazon Redshift. No: the next step.
- Is it SQL over files in S3, run now and then? Yes: Amazon Athena. No: the next step.
- Reshape the data first with an AWS Glue job
Read more Choosing an AWS analytics service
| 1 | A Glue job reads the raw files |
| 2 | It writes partitioned Parquet tables to the catalog |
| 3 | Athena and Redshift query them within Lake Formation's grants |
As text
- Region eu-west-2
- AWS Glue (jobs and catalog)
- AWS Lake Formation (grants)
- S3 data lake
- Amazon S3 (raw files)
- Amazon S3 (curated tables)
- Queries
- Amazon Athena (ad hoc)
- Amazon Redshift (dashboards)
- A Glue job reads the raw files
- It writes partitioned Parquet tables to the catalog
- Athena and Redshift query them within Lake Formation's grants
AWS Glue runs serverless Spark jobs and keeps the Data Catalog that Athena, Redshift and Amazon EMR read. AWS Lake Formation adds grants to that catalog, down to columns, rows and cells.
Read more What is AWS Glue? · What is AWS Lake Formation?
Athena and Redshift#
What it does. Amazon Athena runs standard SQL, or Spark, on data in S3, with nothing to provision.
How it works. Tables live in the Glue Data Catalog; Athena scans the files in parallel and returns results in seconds.
When to use it. Ad hoc questions, log analysis and exploring the lake.
When not to. Dashboards repeating the same queries all day: that is a warehouse's job.
Limits. as of September 2026: concurrent queries are capped per Region; a scan reads at most 1 million partitions.
Cost. Per terabyte scanned, 10 MB minimum per query.
Security. IAM and Lake Formation decide what each analyst reads.
Gotchas. Results land in an S3 bucket: whoever reads it sees every answer.
Read more Amazon Athena documentation · Amazon Athena pricing
What it does. Amazon Redshift is a fully managed, petabyte-scale data warehouse that BI tools query with SQL.
How it works. Serverless scales in seconds and bills only while queries run; provisioned clusters run RA3 nodes. Both can query S3 too.
When to use it. Curated data that many dashboards and analysts query all day.
When not to. Occasional questions over raw files, where Athena needs no warehouse.
Limits. as of September 2026: Python user-defined functions lose support after June 30, 2026.
Cost. Serverless per RPU-hour, by the second with a 60-second minimum; or node hours; plus storage.
Security. Lake Formation can govern its shared data down to rows and columns.
Gotchas. Every Serverless query bills at least 60 seconds, so floods of tiny queries add up.
Read more Amazon Redshift documentation · What is Amazon Redshift Serverless?
SELECT * is a spending decision. Athena bills the bytes it scans: store tables as compressed Parquet, partition them by date, and name only the columns you need.
As text
- 2012: Amazon Redshift
- 2016: Amazon Athena: SQL on S3
- 2017: AWS Glue
- 2022: Redshift Serverless
Generative AI#
Amazon Bedrock serves foundation models from many providers behind one API, with retrieval, guardrails and agents around them. The design questions stay familiar: data, identity, cost and limits.
MegaCorp's support assistant answers merchants' payout questions from the merchant handbook, with no investment advice and no card numbers.
As text
- Must answers draw on your own documents? Yes: A Bedrock knowledge base, for retrieval augmented generation. No: the next step.
- Must the model behave in ways prompting cannot reach? Yes: Customize: Bedrock fine-tuning, or SageMaker AI for full control. No: the next step.
- A Bedrock model and a good prompt
Read more Amazon Bedrock or Amazon SageMaker AI?
What it does. Amazon Bedrock, generally available since 2023, runs foundation models from Amazon, Anthropic, OpenAI and others behind one serverless API.
How it works. Applications call a model by ID through the Converse or OpenAI-compatible APIs; knowledge bases add retrieval, guardrails filter, and AgentCore runs agents.
When to use it. Chat, summaries, extraction and agents built on existing models.
When not to. Training your own models: that is SageMaker AI.
Limits. as of September 2026: each model has token quotas per Region, and some models use tokens faster.
Cost. Tokens in and out on demand; batch at half price for some models; provisioned throughput by the hour.
Security. Model providers never see your prompts or completions; reach Bedrock privately through PrivateLink.
Gotchas. A global inference profile may run a request in any commercial Region: use a geographic one where data residency matters.
Read more Amazon Bedrock documentation · Amazon Bedrock pricing · Data protection in Amazon Bedrock
Knowledge bases and guardrails#
As text
- Support app to Knowledge base: the merchant's question
- Knowledge base to Model: the question and the passages that match it
- Model to Support app: an answer with citations, checked by the guardrail
The team's guardrail denies investment advice as a topic, masks card numbers, and blocks answers not grounded in the handbook. It checks prompts and answers alike, and the ApplyGuardrail API applies it to models outside Bedrock.
Retrieval ignores who is asking, unless you tell it. Only managed knowledge bases filter documents by each user's permissions; a vector store you manage returns any matching passage to anyone.
Read more Amazon Bedrock Knowledge Bases · Amazon Bedrock Guardrails
| 1 | The knowledge base indexes the handbook |
| 2 | The assistant calls Bedrock through the endpoint |
| 3 | which carries the request on the AWS network |
As text
- Region eu-west-2
- Amazon Bedrock (models and knowledge base)
- Amazon S3 (merchant handbook)
- VPC 10.20.0.0/16
- Availability Zone a
- Private subnet 10.20.10.0/24
- AWS Fargate (support assistant)
- VPC endpoints (interface endpoint)
- Private subnet 10.20.10.0/24
- Availability Zone a
- The knowledge base indexes the handbook
- The assistant calls Bedrock through the endpoint
- which carries the request on the AWS network
- The nightly settlement job runs for 40 minutes. Lambda, or something else?
Answer
Not Lambda, which stops at 15 minutes: run an ECS task on Fargate, or split the work into a Step Functions workflow. - Card-scheme webhooks sometimes arrive twice. How does the settlement function avoid posting twice?
Answer
It writes each event ID to DynamoDB only if the ID is absent; a repeat fails the condition and is skipped. - Finance's Athena bill doubles every month. What do you check first?
Answer
How much each query scans: store tables as compressed Parquet, partition them by date, and select only the columns needed.
Keys, secrets and certificates#
AWS splits what Azure keeps in Key Vault: AWS KMS holds keys, AWS Secrets Manager holds secrets, and AWS Certificate Manager issues and renews TLS certificates.
One key vault stores keys, secrets and certificates, with access granted through Azure RBAC or vault access policies.
KMS keeps only keys, which never leave it unencrypted; passwords go to Secrets Manager and certificates to ACM, each with its own permissions.
The ledger's password used to sit in a configuration file. Alex moves it to Secrets Manager, where RDS rotates it every seven days, and encrypts the settlement files under a customer managed key, so the audit trail shows every decrypt.
As text
- Is it a key that encrypts or signs data? Yes: AWS KMS. No: the next step.
- Is it a TLS certificate for a load balancer, CloudFront or API Gateway? Yes: AWS Certificate Manager. No: the next step.
- Is it a password, API key or token that should rotate? Yes: AWS Secrets Manager. No: the next step.
- Plain configuration: Parameter Store
Read more AWS Systems Manager Parameter Store
Keys: AWS KMS#
As text
- Settlement job to AWS KMS: GenerateDataKey under the payments key
- AWS KMS to Settlement job: a data key, in plaintext and encrypted
- Settlement job to Amazon S3: the file encrypted with the data key, plus the encrypted key
- Settlement job: forgets the plaintext key
What it does. AWS KMS creates and controls the keys that encrypt and sign your data, in hardware security modules they never leave unencrypted.
How it works. Services such as S3, EBS and RDS ask KMS for data keys and encrypt the data themselves; KMS only wraps and unwraps the data keys.
When to use it. Customer managed keys where you must control policy, rotation and audit; AWS owned keys where convenience matters most.
When not to. Storing passwords or certificates: those belong in Secrets Manager and ACM.
Limits. as of September 2026: in eu-west-2, 20,000 symmetric requests a second, shared across the account.
Cost. A monthly fee per customer managed key, plus requests beyond a free tier; AWS owned keys are free.
Security. A key policy decides who may use each key, and CloudTrail records each use of your keys.
Gotchas. Deleting a key makes everything encrypted under it unrecoverable: disable it first, and the deletion itself waits 7 to 30 days.
Read more AWS KMS documentation · AWS KMS keys · AWS KMS request quotas
Your buckets spend your KMS quota. Every upload or download of an S3 object encrypted with SSE-KMS is a KMS request made on your behalf, counted against the same account quota as your own calls. A busy bucket can throttle an unrelated service.
Secrets: AWS Secrets Manager#
What it does. AWS Secrets Manager stores database credentials, API keys and tokens, and gives them to code at run time.
How it works. Each secret is encrypted with KMS; rotation replaces it on a schedule, managed for RDS or through a Lambda function.
When to use it. Any credential that would otherwise sit in code or configuration.
When not to. Plain configuration values, which Parameter Store's standard tier holds at no extra charge.
Limits. A secret lives in one Region unless you replicate it.
Cost. Per secret per month, replicas included, and per 10,000 API calls; rotation functions bill as Lambda.
Security. Let each role read only its own secrets.
Gotchas. Rotation breaks code that caches a password forever: fetch the secret again when a login fails.
Read more AWS Secrets Manager documentation · What is AWS Secrets Manager?
Certificates: AWS Certificate Manager#
What it does. AWS Certificate Manager issues, stores and renews TLS certificates for load balancers, CloudFront and API Gateway.
How it works. Request a certificate for your domains, or import one; ACM renews the certificates it issues.
When to use it. HTTPS on AWS's own front doors, where public certificates cost nothing.
When not to. Certificates on your own servers: use ACME automation, or pay for exportable ones.
Limits. Certificates are Regional and cannot be copied; CloudFront uses only us-east-1.
Cost. Free for public certificates on integrated services; exportable ones cost per domain.
Security. A wildcard certificate covers every subdomain, so issue it sparingly.
Gotchas. Every Region that serves the domain needs its own certificate, validated there.
Read more AWS Certificate Manager documentation · What is AWS Certificate Manager?
As text
- 2014: AWS KMS
- 2016: AWS Certificate Manager, free certificates
- 2018: AWS Secrets Manager
MegaCorp's design gains its keys: the payments key and the ledger's credentials in eu-west-2, and the portal's certificate in us-east-1.
| 1 | CloudFront serves the requests it allows |
| 2 | CloudFront uses a certificate from us-east-1 |
| 3 | The ledger's credentials are encrypted under the payments key |
As text
- AWS WAF (web ACL)
- Amazon CloudFront (portal)
- AWS Account payments-prod
- Region eu-west-2
- AWS KMS (payments key)
- AWS Secrets Manager (ledger)
- Region us-east-1
- AWS Certificate Manager (portal)
- Region eu-west-2
- CloudFront serves the requests it allows
- CloudFront uses a certificate from us-east-1
- The ledger's credentials are encrypted under the payments key
Read more Delete an AWS KMS key · AWS Certificate Manager pricing
Detection and posture#
Four questions keep an estate safe: who did what, how is it configured, is anyone attacking, and where do we stand. CloudTrail, Config, GuardDuty and Security Hub answer them.
As text
- Who did what, and when? Yes: AWS CloudTrail. No: the next step.
- How is a resource configured, and has it drifted? Yes: AWS Config. No: the next step.
- Is someone attacking us now? Yes: Amazon GuardDuty. No: the next step.
- Where do we stand, across accounts? AWS Security Hub
Read more What is AWS CloudTrail?
Records: CloudTrail and Config#
What it does. AWS CloudTrail records API calls made from the console, CLI, SDKs and AWS services.
How it works. Event history keeps 90 days of management events per Region; a trail delivers events to S3, for one account or the whole organization, for as long as you keep them.
When to use it. Always, with one organization trail into a separate log-archive account.
When not to. Application logs: those belong in CloudWatch Logs.
Limits. as of September 2026: event history covers 90 days, one Region at a time.
Cost. The first copy of management events is free; data events and extra copies are charged.
Security. Deliver to a bucket in another account that the workload's admins cannot change.
Gotchas. Event history is per Region: activity in a Region you never use shows up only there.
Read more AWS CloudTrail documentation · AWS CloudTrail pricing
What it does. AWS Config records how each resource is configured, and every change, over time.
How it works. Rules check resources as they change and flag noncompliant ones; conformance packs bundle rules, and an aggregator gathers every account and Region into one view.
When to use it. Proving how something was configured on a given day, and catching drift.
When not to. Recording every resource type everywhere without a plan: each change is billed.
Limits. It records only the resource types you choose.
Cost. Per configuration item recorded and per rule evaluation.
Security. Remediation can fix a noncompliant resource automatically.
Gotchas. Rules detect after the fact; to stop a change happening at all, use a service control policy.
Read more AWS Config documentation · AWS Config pricing
Threats and posture: GuardDuty and Security Hub#
What it does. Amazon GuardDuty detects threats such as stolen credentials, cryptomining and data exfiltration.
How it works. Once enabled, it analyses CloudTrail management events, VPC flow logs and DNS logs with no agents; protection plans add S3, EKS, RDS, Lambda and runtime monitoring.
When to use it. Every account and Region, managed from one delegated administrator.
When not to. Checking configuration against a standard: that is Security Hub.
Limits. It watches only the Regions where you enable it.
Cost. Per million events and per GB of logs analysed, after a 30-day free trial.
Security. Route high-severity findings through EventBridge to whoever is on call.
Gotchas. Protection plans launched after you enabled GuardDuty stay off until you turn them on.
Read more Amazon GuardDuty documentation · Amazon GuardDuty pricing
What it does. AWS Security Hub scores the estate against security standards and gathers findings from GuardDuty, Inspector, Macie and partners in one place.
How it works. Controls from the AWS Foundational Security Best Practices, CIS, PCI DSS and NIST standards run as checks; findings share one format, and automation rules and EventBridge act on them.
When to use it. One view of security across every account, for security teams and auditors.
When not to. Detecting threats itself: it relies on GuardDuty and others for that.
Limits. It sees only Regions where it is enabled, and only findings from after that.
Cost. Per security check and per finding ingested, after a 30-day free trial.
Security. Run it from a delegated administrator account for the whole organization.
Gotchas. Most controls need AWS Config recording, which is billed separately.
Read more AWS Security Hub documentation · Introduction to AWS Security Hub CSPM
Amazon Inspector continually scans EC2 instances, ECR images and Lambda functions for known vulnerabilities, and reports to Security Hub too.
Watch the Regions you don't use. AWS recommends enabling GuardDuty in every supported Region, even where you run nothing, so it can report unusual activity there.
As text
- Amazon GuardDuty to Amazon EventBridge: a finding: keys used from an unfamiliar network
- Amazon EventBridge to On-call: a rule matches high severity and publishes to SNS
- On-call: contain first, revoke the keys, then investigate
AWS Security Incident Response can triage findings for you, escalating fewer than 1% of them, and engages engineers within 15 minutes on cases you raise.
MegaCorp's design gains its detection layer, run from the audit account.
| 1 | Config records each change, and Security Hub checks it against its standards |
| 2 | GuardDuty findings from every account land in Security Hub |
As text
- AWS Account payments-prod
- Region eu-west-2
- AWS Config (recorder)
- Region eu-west-2
- AWS Account audit
- delegated administrator
- AWS Security Hub (all findings)
- Amazon GuardDuty (organization)
- delegated administrator
- Config records each change, and Security Hub checks it against its standards
- GuardDuty findings from every account land in Security Hub
Read more What is Amazon Inspector? · What is AWS Security Incident Response?
Observability#
Metrics say something is wrong, logs say what happened, and traces say where. Amazon CloudWatch holds the metrics, logs and alarms, and AWS X-Ray follows a request from service to service.
At two in the morning, settlement slows. An alarm on the function's 99th-percentile latency pages Alex; the trace shows the ledger query taking eight seconds; a Logs Insights query finds the lock timeouts behind it.
As text
- Must you know within minutes that something is wrong? Yes: A CloudWatch metric, with an alarm. No: the next step.
- Do you need what happened, line by line? Yes: CloudWatch Logs, queried with Logs Insights. No: the next step.
- Do you need where the time went, across services? Yes: A trace in X-Ray or Application Signals. No: the next step.
- A dashboard that puts them side by side
Read more Metrics concepts
Metrics, logs and alarms: Amazon CloudWatch#
What it does. Amazon CloudWatch collects metrics and logs from AWS services and your code, alarms on them, and draws dashboards.
How it works. Services publish their metrics free; your code adds custom metrics through OpenTelemetry or PutMetricData, and writes logs to log groups that Logs Insights queries.
When to use it. Monitoring everything on AWS, from one team's dashboard to a cross-account monitoring account.
When not to. Keeping every debug line forever: set a retention, or route old logs to S3.
Limits. as of September 2026: metrics stay in their Region, with 1-minute detail for 15 days and hourly data for 15 months.
Cost. Custom metrics per month, log ingestion and storage per GB, queries per GB scanned, and alarms.
Security. Mask sensitive data, such as card numbers, with a data protection policy on the log group.
Gotchas. Each unique combination of dimensions is a separate custom metric, billed on its own.
Read more Amazon CloudWatch documentation · What is Amazon CloudWatch Logs? · Amazon CloudWatch pricing
As text
- Settlement function to CloudWatch alarm: p99 latency above two seconds, three minutes running
- CloudWatch alarm to On-call: the state changes to ALARM, and SNS pages
- On-call: open the trace, then query the logs
Logs never expire unless you say so. A new log group keeps its logs forever by default, and storage is billed every month. Set a retention on every log group when you create it.
Traces: AWS X-Ray#
What it does. AWS X-Ray follows each request through your services and the AWS resources, databases and APIs they call.
How it works. Instrumented code sends segments, and integrated services such as Lambda add their own; X-Ray joins them into traces and a map of services.
When to use it. Finding which hop in a chain of services is slow or failing.
When not to. Counting things or keeping audit records: those are metrics and logs.
Limits. A trace covers only instrumented code and integrated services.
Cost. Trace volume drives the bill, and sampling rules control the volume.
Security. Keep secrets and card numbers out of trace annotations.
Gotchas. By default only the first request each second and 5% of the rest are traced, so a rare failure may be missed.
Read more AWS X-Ray documentation · What is AWS X-Ray? · Configuring sampling rules
CloudWatch Application Signals instruments Java, Python, Node.js and .NET services on EKS, ECS, EC2 and Lambda through OpenTelemetry, and shows latency, errors and service level objectives without a dashboard to build.
Read more Application Signals
MegaCorp's design gains its observability layer: payments-prod shares its signals with a monitoring account.
| 1 | payments-prod shares its metrics and logs with the monitoring account |
| 2 | and its traces |
As text
- AWS Account payments-prod
- Region eu-west-2
- Amazon CloudWatch (payments)
- AWS X-Ray (traces)
- Region eu-west-2
- AWS Account monitoring
- Amazon CloudWatch (monitoring)
- payments-prod shares its metrics and logs with the monitoring account
- and its traces
Read more Using Amazon CloudWatch alarms
Resilience and disaster recovery#
A zone fails, a Region fails, or someone deletes the data. Zones answer the first, a disaster recovery strategy the second, and backups kept out of reach the third.
MegaCorp's payments board sets two numbers. Payments must flow again within an hour: the recovery time objective, or RTO. At most a minute of payments may be lost: the recovery point objective, or RPO. Lower numbers cost more.
Read more Recovery objectives (RTO and RPO)
Losing a zone#
Most disasters hit only one zone, so running in several already covers much of the risk. Make the design statically stable: the capacity to survive a lost zone runs before the failure, because launching it during one relies on control planes and on spare capacity. In two zones, each must carry the whole load; in three, each carries half.
What it does. Amazon Application Recovery Controller (ARC) moves traffic away from a failing zone or Region.
How it works. A zonal shift takes a load balancer's or Auto Scaling group's traffic out of one zone; with zonal autoshift, AWS starts the shift when a zone is impaired. Routing controls fail over between Regions.
When to use it. A bad deployment or an impaired zone; a rehearsed Region failover.
When not to. Automatic Region failover on one alarm: a false alarm costs availability and data.
Limits. as of September 2026: a zonal shift lasts up to three days, extendable, for ALBs, NLBs, Auto Scaling groups and EKS.
Cost. Zonal shift is free; routing control clusters are billed.
Security. Safety rules can keep only one Region switched on at a time.
Gotchas. A load balancer that is failing open ignores a zonal shift.
Read more Amazon Application Recovery Controller documentation · Zonal shift in ARC · ARC pricing
Shift only onto capacity that exists. ARC moves traffic, not servers: before a zonal shift, the other zones must already be able to carry the extra load.
What it does. AWS Fault Injection Service (FIS) breaks things on purpose, so you learn how the system copes before a real failure teaches you.
How it works. An experiment template names actions, such as stopping tasks or failing over a database; targets, chosen by tag; and stop conditions, CloudWatch alarms that end it early. Ready-made scenarios include AZ Availability: Power Interruption.
When to use it. Proving that failover, alarms and runbooks work.
When not to. Production, before the experiment has passed in pre-production.
Limits. A single-account experiment reaches only its own account.
Cost. Per action-minute, more for each extra target account.
Security. It acts through an IAM role you give it; limit that role to the targets.
Gotchas. The actions are real: without a stop condition, a test can become an outage.
Read more AWS Fault Injection Service documentation · AWS FIS scenarios reference · AWS FIS pricing
Read more REL11-BP05 Use static stability
Losing a Region: four strategies#
Guarding against the loss of a Region costs more. AWS names four strategies, in rising order of cost and falling RTO and RPO.
As text
- Can the business wait up to a day, and lose hours of data? Yes: Backup and restore. No: the next step.
- Can it wait tens of minutes, and lose minutes of data? Yes: Pilot light. No: the next step.
- Can it wait minutes, and lose seconds of data? Yes: Warm standby. No: the next step.
- Multi-site active/active, near zero
As text
- In the recovery Region
- No compute running
- Backup and restore
- AWS Backup (copied backups)
- Pilot light
- Amazon Aurora (live copy)
- Backup and restore
- Compute running
- Warm standby
- Amazon Aurora (live copy)
- AWS Fargate (a few tasks)
- Multi-site active/active
- Amazon Aurora (live copy)
- AWS Fargate (full size, serving)
- Warm standby
- No compute running
Backup and restore rebuilds from backups and infrastructure as code; pilot light keeps data live but deploys compute only when needed; warm standby only has to scale up; multi-site active/active already serves from every Region.
MegaCorp picks pilot light, in eu-west-1 and a separate account, payments-dr. An Aurora global database keeps the ledger there, typically under a second behind.
As text
- On-call to Aurora global database: promote eu-west-1 to take writes, in under a minute
- On-call: deploy the payments tasks in eu-west-1, and scale them out
- On-call to ARC routing control: switch traffic from eu-west-2 to eu-west-1
A person starts the failover, since a false alarm costs availability and data. The steps are scripted, and ARC's data plane is designed for higher availability than control planes.
Microsoft pairs many regions, and geo-redundant storage copies data to the pair on its own. Many newer regions have no pair.
No Region has a partner. Nothing leaves a Region unless a replication or copy feature, or your code, sends it, and you choose the recovery Region.
For whole servers, AWS Elastic Disaster Recovery replicates machines from a data centre or another cloud into a low-cost staging area, and launches them within minutes.
Read more Disaster recovery options in the cloud · REL13-BP02 Use defined recovery strategies · What is Elastic Disaster Recovery? · Azure region pairs and nonpaired regions
Losing the data: backups#
Replication copies mistakes too: a deleted ledger row is gone from eu-west-1 within a second. Only a point-in-time backup goes back to before the mistake.
What it does. AWS Backup schedules, copies and keeps backups of EC2, EBS, S3, RDS, Aurora, DynamoDB, EFS, FSx and more, from one place.
How it works. A backup plan sets frequency and retention, and resources join it by tag. Backups land in vaults, and copies go to other Regions and accounts.
When to use it. Every data store, with backup policies applying plans across the organization.
When not to. Alone, for a low RTO or RPO: restores take hours.
Limits. It governs only its own backups, and some resource types copy in full every time.
Cost. Storage per GB-month, restores per GB, and copies between Regions; no service fee.
Security. Copy to a locked vault in another account: not even the root user can delete a backup early.
Gotchas. A compliance-mode lock is permanent once its grace time, at least three days, ends.
Read more AWS Backup documentation · AWS Backup Vault Lock · AWS Backup pricing
Quotas#
Every account starts with default quotas, many of them per Region. Lambda functions in a Region share 1,000 concurrent executions by default as of September 2026. MegaCorp raised that in eu-west-2; in eu-west-1 it is still the default.
What it does. Service Quotas shows each quota, the maximum for a resource or action, and requests increases.
How it works. Services set defaults; adjustable quotas rise by request, which support may approve, deny or partly approve. Automatic Management warns before a quota runs out.
When to use it. Before launch, before big events, and for the recovery Region.
When not to. Fixed quotas: design around them.
Limits. as of September 2026: a request template raises up to 10 quotas in each new account of an organization.
Cost. Free; alarms on quota usage bill as CloudWatch alarms.
Security. Alarm on usage: a sudden climb can mean a runaway job or stolen credentials.
Gotchas. Increases take time to review, and global quotas are raised only from us-east-1.
Read more Service Quotas documentation · What is Service Quotas? · Quota request templates
Raise quotas in the recovery Region too. A pilot light scales up during failover, so the recovery Region's quotas must already allow production capacity.
MegaCorp's design gains its resilience layer: a quarterly game day, and the ledger's live copy and locked backups in payments-dr.
| 1 | Each quarter, FIS fails over the ledger's writer to prove the tasks reconnect |
| 2 | The global database keeps a secondary ledger in eu-west-1, typically under a second behind |
| 3 | AWS Backup copies each night's ledger backup to a locked vault in payments-dr |
As text
- AWS Account payments-prod
- Region eu-west-2
- AWS Fault Injection Service
- Amazon Aurora (global database)
- Region eu-west-2
- AWS Account payments-dr
- Region eu-west-1
- Amazon Aurora (ledger secondary)
- Vault Lock, compliance mode
- AWS Backup (ledger backups)
- Region eu-west-1
- Each quarter, FIS fails over the ledger's writer to prove the tasks reconnect
- The global database keeps a secondary ledger in eu-west-1, typically under a second behind
- AWS Backup copies each night's ledger backup to a locked vault in payments-dr
Read more Service Quotas and CloudWatch alarms
Cost#
AWS bills by use: compute by the second or hour, storage by the gigabyte-month, and data by the gigabyte as it moves.
MegaCorp's first month on AWS brings a surprise line: NAT gateway data processing. The statements job writes every file to S3 through the NAT gateway, paying per gigabyte. Alex adds a gateway endpoint, which has no hourly or processing charge, and that traffic now moves for nothing.
The paths data travels#
| 1 | Zone to zone: billed as it leaves and again as it arrives |
| 2 | To S3 through a gateway endpoint: no charge to move it |
| 3 | The endpoint routes straight to the bucket |
| 4 | Region to Region: billed as it leaves eu-west-2 |
As text
- Region eu-west-2
- Amazon S3 (statements)
- VPC 10.20.0.0/16
- VPC endpoints (S3 gateway)
- Availability Zone a
- AWS Fargate (statements job)
- Availability Zone b
- Amazon Aurora (writer)
- Region eu-west-1
- Amazon S3 (replica)
- Zone to zone: billed as it leaves and again as it arrives
- To S3 through a gateway endpoint: no charge to move it
- The endpoint routes straight to the bucket
- Region to Region: billed as it leaves eu-west-2
Data in from the internet is free, and so is data that stays in one zone. Data that crosses a zone, a Region or the edge of AWS is usually billed.
A NAT gateway charges twice. Every gigabyte through it pays a processing charge on top of the usual data transfer. Send S3 and DynamoDB traffic through gateway endpoints instead.
Read more Overview of data transfer costs for common architectures · Understanding data transfer charges
Pricing models#
On-Demand pays by the second; Spot sells spare capacity cheaply but can be reclaimed at two minutes' notice; a commitment cuts the rate for steady use, as chapter 17 showed.
What it does. Savings Plans lower the price of steady use in exchange for committing to spend an amount per hour for one or three years.
How it works. Compute Savings Plans cover EC2 in any family, size or Region, plus Fargate and Lambda. EC2 Instance plans save more on one family in one Region; Database plans cover Aurora, RDS, DynamoDB and more.
When to use it. Steady use, sized from Cost Explorer's recommendation and bought centrally.
When not to. Use that may shrink: a plan can't be cancelled or changed.
Limits. No capacity reservation, and no discount on Spot.
Cost. All, part or none paid upfront, at a rate fixed for the term.
Security. Only the management account decides which accounts share the discount.
Gotchas. Turning sharing off for an account can raise the bill.
Read more Savings Plans documentation · Savings Plans types · Reserved Instances and Savings Plans discount sharing
Reserved Instances, the older commitment to one instance configuration, still fill many bills; Savings Plans give similar discounts without exchanges.
Seeing and controlling spend#
As text
- Must someone hear when spend passes a line you set? Yes: AWS Budgets. No: the next step.
- Must someone hear about spend nobody planned? Yes: AWS Cost Anomaly Detection. No: the next step.
- Do you need every line item, joined with your own data? Yes: A CUR 2.0 export to S3. No: the next step.
- Explore and forecast in AWS Cost Explorer
What it does. AWS Cost Explorer charts cost and usage by service, account, Region and tag, and forecasts ahead.
How it works. It refreshes at least daily, and recommends Savings Plans and Reserved Instances from your usage.
When to use it. Finding what drives the bill, and sizing a commitment.
When not to. Line-by-line analysis with your own data: export CUR 2.0 to S3.
Limits. as of September 2026: 13 months of history, and forecasts 18 months ahead.
Cost. The console is free; API requests are billed.
Security. Tags appear in cost reports, so keep secrets out of them.
Gotchas. Once enabled, Cost Explorer can't be turned off.
Read more AWS Cost Management documentation · Analyzing your costs with AWS Cost Explorer
What it does. AWS Budgets alerts when cost, usage or commitment coverage crosses a line you set, and can act.
How it works. A budget watches actual or forecast spend; alerts go by email or SNS, and actions apply an IAM policy or SCP, or stop EC2 and RDS instances.
When to use it. Every account, with a forecast alert; sandboxes, with an action.
When not to. A hard cap: it sees charges hours late.
Limits. Data updates up to three times a day.
Cost. Alerts are free, as are two action-enabled budgets a month.
Security. Actions run through a role you grant AWS Budgets.
Gotchas. A budget is visible only in the account that created it.
Read more AWS Cost Management documentation · Configuring budget actions · AWS Budgets pricing
As text
- AWS Budgets to Developers: the month's forecast passes 90% of the budget
- AWS Budgets to Sandbox account: an action applies an SCP that denies new instances
- Sandbox account: running instances need an action in the sandbox's own budget to stop
Cost Anomaly Detection flags unplanned spend with its likely cause; Data Exports sends CUR 2.0, every line of the bill, to S3.
Read more Detecting unusual spend with AWS Cost Anomaly Detection · What is AWS Data Exports?
Tags: who pays for what#
MegaCorp tags every resource with CostCenter and Owner. A tag reaches the bill only after the management account activates it, and earlier costs stay untagged unless backfilled, up to twelve months.
Keys are case sensitive, so CostCenter and costcenter split the bill. A tag policy standardizes them, though it ignores untagged resources; account tags cover everything in an account, even charges that cannot be tagged.
Cost Management can copy subscription and resource group tags onto the usage records of every resource beneath them.
Only account tags flow down, to all usage in the account. Otherwise, tag each resource, and activate each key for cost allocation.
Read more Organizing and tracking costs using cost allocation tags · Backfill cost allocation tags · Tag policies · Using account tags for cost allocation · Group and allocate costs using tag inheritance
Infrastructure as code#
Every MegaCorp resource starts as reviewed code. On AWS that code is Terraform, with its state in S3, or CloudFormation, written as YAML or generated from Java by the AWS CDK.
MegaCorp's platform team already runs Terraform on Azure, so the accounts and networks stay in Terraform. The payments developers write Java all day, so they define their own buckets and queues with the CDK.
As text
- Do your teams already run Terraform, or manage more than AWS? Yes: Terraform, with state in S3. No: the next step.
- Do developers want loops, types and tests in Java? Yes: The AWS CDK, which writes CloudFormation. No: the next step.
- CloudFormation templates, in YAML
Read more What is CloudFormation?
Terraform on AWS#
Terraform works on AWS as it does on Azure; the backend is what changes. State moves from a blob container to an S3 bucket, and locking to a lock file that S3 writes only if none exists.
backend "azurerm" {
use_oidc = true
use_azuread_auth = true
storage_account_name = "megacorptfstate"
container_name = "tfstate"
key = "payments/network.tfstate"
}backend "s3" {
bucket = "megacorp-terraform-state"
key = "payments/network.tfstate"
region = "eu-west-2"
use_lockfile = true
encrypt = true
}Turn on versioning for the state bucket, so a bad write can be undone, and guard the state like a
secret: it can hold values such as initial database passwords. Older configurations lock with
DynamoDB, which is deprecated. The provider's default_tags block puts CostCenter and Owner
on every resource that takes tags, except Auto Scaling groups.
Terraform removes the rule AWS adds. AWS gives a new security group an allow-all outbound rule, and the AWS provider deletes it, so the group sends nothing until you write rules. The payments load balancer needs one to reach its tasks on the listener and health check port.
Read more Terraform S3 backend · Terraform azurerm backend · Sensitive data in Terraform state · AWS provider: default_tags · Terraform: aws_security_group
CloudFormation#
What it does. AWS CloudFormation creates, updates and deletes a set of AWS resources, called a stack, from a template.
How it works. The stack records what CloudFormation created. A change set previews an update, including any replacement, and nothing changes until you execute it; a failed deployment rolls back.
When to use it. AWS-only estates, StackSets that reach every account and Region, and everything the CDK generates.
When not to. Resources outside AWS, or teams already fluent in Terraform.
Limits. as of September 2026: 500 resources per template, and a 1 MB template body from S3.
Cost. No charge for AWS resource types; third-party types and custom hooks are billed per operation.
Security. Review each change set before executing it: it shows every deletion and replacement.
Gotchas. Deleting a stack deletes its resources, and fails on a bucket that still holds objects.
Read more AWS CloudFormation documentation · Update CloudFormation stacks using change sets · CloudFormation quotas
Resources:
Statements:
Type: AWS::S3::Bucket
Properties:
BucketEncryption:
ServerSideEncryptionConfiguration:
- ServerSideEncryptionByDefault:
SSEAlgorithm: aws:kms
PublicAccessBlockConfiguration:
BlockPublicAcls: true
BlockPublicPolicy: true
IgnorePublicAcls: true
RestrictPublicBuckets: trueChanges made by hand show up as drift: CloudFormation reports each changed property but leaves the fix to you. StackSets deploy one template to many accounts and Regions in a single operation.
Read more CloudFormation drift detection · AWS CloudFormation pricing
The AWS CDK, in Java#
What it does. The AWS CDK defines infrastructure in Java, TypeScript, Python, C#, Go or JavaScript, and deploys it through CloudFormation.
How it works. Constructs, from single resources to whole patterns with secure defaults, make up stacks. The CDK synthesizes a CloudFormation template, then deploys it as a change set.
When to use it. Developers who want loops, types, tests and reuse in the language they already write.
When not to. Reviewers who must read every line that deploys: a few lines of Java can become hundreds of lines of template.
Limits. It deploys only through CloudFormation, so CloudFormation's quotas apply.
Cost. Open source; you pay for what it deploys.
Security. By default, a deployment that broadens permissions or security group rules waits for approval.
Gotchas. Each account and Region must be bootstrapped once before stacks with assets can deploy.
Read more AWS CDK documentation · What is the AWS CDK?
public PaymentsStack(final Construct scope, final String id, final StackProps props) {
super(scope, id, props);
Vpc.Builder.create(this, "Payments")
.ipAddresses(IpAddresses.cidr("10.20.0.0/16"))
.maxAzs(2)
.natGateways(2)
.build();
Bucket.Builder.create(this, "Statements")
.encryption(BucketEncryption.KMS_MANAGED)
.blockPublicAccess(BlockPublicAccess.BLOCK_ALL)
.enforceSsl(true)
.build();
}These lines define the payments VPC, across two zones with a NAT gateway in each, and the statements bucket, encrypted and closed to the public. One construct can expand into dozens of resources: AWS's own example of a VPC, cluster and service becomes more than 50.
As text
- Developer to AWS CDK: cdk deploy
- AWS CDK: synthesizes the Java app into a CloudFormation template
- AWS CDK to AWS CloudFormation: creates a change set, then executes it
- AWS CloudFormation: builds the stack, and rolls back if a resource fails
How infrastructure as code grew#
As text
- 2011: AWS CloudFormation
- 2014: Terraform 0.1: AWS and DigitalOcean
- 2019: AWS CDK, generally available
- 2021: Bicep 0.3, supported in production
Terraform's author built it after admiring CloudFormation in 2011, wanting the same idea for every cloud. On Azure, ARM's JSON templates gave way to Bicep, a shorter language that compiles to that JSON, much as the CDK's Java compiles to CloudFormation. Azure keeps Bicep's state, so there is no state file to guard.
resource statements 'Microsoft.Storage/storageAccounts@2023-05-01' = {
name: 'megacorpstatements'
location: resourceGroup().location
sku: {
name: 'Standard_LRS'
}
kind: 'StorageV2'
}Read more What is Bicep? · Bicep frequently asked questions
Delivery pipelines#
A change reaches production in three moves: a pipeline builds and tests it without long-lived keys, a deployment exposes it a little at a time, and feature flags decide who sees it.
The payments service's first pipeline ran in GitHub Actions, trading a signed token for short-lived credentials, as chapter 10 showed. As more teams join, MegaCorp's platform team moves releases into a tooling account running CodePipeline, with GitHub still the source.
Read more Create a role for OpenID Connect federation
Pipelines across accounts#
| 1 | Each commit is built and tested in CodeBuild |
| 2 | The pipeline assumes the deploy role in payments-dev |
| 3 | After an approval, it assumes the deploy role in payments-prod |
As text
- AWS Account tooling
- Region eu-west-2
- AWS CodeBuild (build and test)
- AWS CodePipeline (payments)
- Region eu-west-2
- Workload accounts
- AWS Account payments-dev
- IAM role (deploy)
- AWS Account payments-prod
- IAM role (deploy)
- AWS Account payments-dev
- Each commit is built and tested in CodeBuild
- The pipeline assumes the deploy role in payments-dev
- After an approval, it assumes the deploy role in payments-prod
What it does. AWS CodePipeline runs a release as stages of actions, such as source, build, approve and deploy, on every change.
How it works. A V2 pipeline starts on chosen branches, tags or paths, passes variables between stages, and can roll a stage back. Actions in other accounts run through roles in those accounts.
When to use it. Releases to many accounts from one tooling account, with no long-lived keys.
When not to. Teams already well served by GitHub Actions assuming roles through OIDC.
Limits. An artifact moves between accounts only through the pipeline's own account.
Cost. V2 bills per action execution minute, approvals free; builds and artifacts bill separately.
Security. Let each deploy role trust only the pipeline's service role.
Gotchas. Cross-account pipelines need a customer managed KMS key for artifacts; the default key won't work.
Read more AWS CodePipeline documentation · Create a pipeline that uses resources from another account · AWS CodePipeline pricing
What it does. AWS CodeBuild compiles code, runs tests and produces artifacts, with no build servers to run.
How it works. Each build runs your commands in a prepackaged environment, with tools such as Maven and Gradle, or in your own image, and CodeBuild scales with the builds that arrive.
When to use it. Build and test actions in CodePipeline, and builds nobody wants to host.
When not to. Work that needs a long-running server: builds are time-limited.
Limits. as of September 2026: a build runs for at most 36 hours, and concurrent builds are capped per compute type, often at one by default.
Cost. Per build minute.
Security. Give each project its own IAM role, limited to what its build touches.
Gotchas. Builds beyond the concurrency quota fail; raise it in Service Quotas before a busy release.
Read more AWS CodeBuild documentation · Quotas for AWS CodeBuild · AWS CodeBuild pricing
Deployment strategies#
As text
- Is it a setting or a feature switch, not new code? Yes: An AppConfig rollout, 20% at a time. No: the next step.
- Must the old version keep running until the new one proves itself? Yes: Blue/green, with a bake time. No: the next step.
- Can a few users try it before everyone? Yes: Canary: 10% first, the rest minutes later. No: the next step.
- All at once, rolled back if it fails
Since July 2025, Amazon ECS runs blue/green deployments itself, keeping the old revision ready until the new one proves itself.
As text
- Amazon ECS to Green revision: start its tasks beside the blue ones
- Amazon ECS to Green revision: test it through the lifecycle hooks
- Amazon ECS to Green revision: shift the production traffic to it
- Blue revision: kept through the bake time, ready to take the traffic back
For Lambda functions, and ECS services that use it, CodeDeploy shifts traffic in canary or linear steps and rolls back when a deployment fails or an alarm threshold is met.
As text
- AWS CodeDeploy to Settlement function: send 10% of invocations to the new version
- CloudWatch alarm to AWS CodeDeploy: errors cross the threshold within five minutes
- AWS CodeDeploy to Settlement function: roll back: redeploy the previous version
Blue/green needs room for two. ECS runs both revisions until the bake time ends, which can double the service's tasks, so capacity and quotas must allow both.
Read more Amazon ECS blue/green deployments · Deployment configurations in CodeDeploy · Redeploy and roll back a deployment with CodeDeploy
Configuration and feature flags#
The new instant-refund feature ships hidden behind a flag. The team turns it on for a fifth of its targets every six minutes, AWS's recommended strategy, and an alarm on refund errors would roll it back.
What it does. AWS AppConfig changes how an application behaves, through feature flags and settings, without a redeployment.
How it works. Each change is validated, then rolled out by a deployment strategy; the AppConfig Agent caches it beside the application, and an alarm during the rollout or bake time rolls it back.
When to use it. Feature flags, kill switches, allow and block lists, and tuning values.
When not to. Values fixed at build time, which belong in the code or the environment.
Limits. as of September 2026: AWS's recommended production strategy takes 30 minutes, then watches alarms for 30 more.
Cost. Per retrieval of configuration; the agent's cache keeps retrievals down.
Security. IAM decides who may change or deploy a flag, and CloudTrail records each change.
Gotchas. Rollback on alarms works only once AppConfig has permission to watch them.
Read more AWS AppConfig documentation · Predefined AppConfig deployment strategies
Azure App Configuration stores settings and feature flags in one place; applications read them through client libraries and pick up changes without a restart.
AWS AppConfig treats each change as a release: validated, rolled out to a share of targets at a time, and rolled back on a CloudWatch alarm.
Read more What is AWS AppConfig? · What is Azure App Configuration?
Operations#
Machines still need patches and, now and then, a person. Systems Manager provides both without an open port.
At three in the morning a disk fills on a ledger batch host. Once, an engineer reached it through a bastion on port 22; now Alex opens a session from the console, and no inbound port is open.
As text
- Do you need a shell on one machine, now? Yes: Session Manager. No: the next step.
- Must one command run on many machines? Yes: Run Command. No: the next step.
- Is it a multi-step fix, with approvals, across accounts? Yes: An Automation runbook. No: the next step.
- Patches on a schedule: a Patch Manager policy
What it does. AWS Systems Manager operates fleets of machines, on EC2, on premises and in other clouds, without logging in to each one.
How it works. SSM Agent on each machine connects to the service; tools built on it open sessions, run commands, patch and automate fixes, across accounts and Regions.
When to use it. Any fleet of instances or servers: access, patching and routine fixes.
When not to. Major operating system upgrades, which Patch Manager does not perform.
Limits. as of September 2026: 100 automations run at once per account, with up to 5,000 queued.
Cost. Session Manager, Patch Manager and Run Command cost nothing extra on EC2; Automation bills per step.
Security. IAM decides who may open a session to which machines, and sessions can be logged to S3 or CloudWatch Logs.
Gotchas. A machine counts as managed only when its agent can reach the service, through a route out or VPC endpoints.
Read more AWS Systems Manager documentation · What is AWS Systems Manager? · AWS Systems Manager pricing
Access without open ports#
As text
- Alex to Session Manager: start a session to the batch host
- Session Manager: checks Alex's IAM permissions and the session settings
- Session Manager to Batch host: SSM Agent opens the two-way channel
- Session Manager to S3 bucket: the session's log
Sessions that forward a port or carry SSH are not logged, because Session Manager only tunnels them. Hosts without public addresses reach the service through VPC endpoints.
Read more AWS Systems Manager Session Manager
Patching#
One patch policy covers every account and Region in MegaCorp's organization. It scans daily and installs weekly, against a patch baseline that approves security updates a few days after release.
As text
- Patch policy to Payments hosts: scan every day against the baseline
- Patch policy to Payments hosts: install approved patches in the weekly window
- Patch policy to S3 bucket: compliance report, as CSV
Compliant is not the same as secure. Patch Manager measures each machine against your baseline, not against every known flaw. AWS does not test patches first, and the predefined baselines are examples, not recommendations.
Read more AWS Systems Manager Patch Manager
Runbooks#
When the batch disk fills again, the fix becomes an Automation runbook: steps that snapshot the volume, grow it, and extend the file system, started by an EventBridge rule on the disk alarm.
An Azure Automation runbook is a PowerShell or Python script, run in an Azure sandbox or on a Hybrid Runbook Worker.
A Systems Manager runbook is a YAML or JSON document of steps, which can run scripts, wait for approval, and run across accounts and Regions with rate controls.
Read more AWS Systems Manager Automation · Azure Automation runbook types
- An auditor asks who decrypted last month's settlement files. Where do you look?
Answer
CloudTrail, which records every use of the customer managed payments key, with the caller. - The board wants payments back within an hour, losing at most a minute. Which strategy, and what must you check in the recovery Region?
Answer
Pilot light: data live there, compute deployed on failover. Check that its quotas allow production capacity. - Someone deletes ledger rows by mistake. Why doesn't the global database save you?
Answer
Replication copies the deletion within a second. A point-in-time backup, in a locked vault in another account, brings the rows back. - An engineer needs a shell on a host in a private subnet. How, without opening port 22?
Answer
Session Manager: IAM grants the session, SSM Agent connects out through VPC endpoints, and the session is logged.
The design method#
Every design in this book answers the same thirteen questions. The Well-Architected pillars say why one answer beats another, and a decision record keeps the reasons after the people who chose them move on.
When MegaCorp's payments team started on AWS, each choice was argued in a meeting and forgotten by the next one. Now every design walks the same thirteen questions: six decided once for the company, four for each workload, and three asked of every design.
| Question | What it decides | Covered in |
|---|---|---|
| 1 Accounts | how workloads split into accounts, and how the accounts are organized | chapter 7, chapter 15 |
| 2 Sign-in | how people, pipelines and customers prove who they are | chapter 8, chapter 9, chapter 10 |
| 3 Networks | Regions, address plans, VPCs and how they connect | chapter 11, chapter 12, chapter 13, chapter 14 |
| 4 Guardrails | what no account may do, and how posture is checked | chapter 15, chapter 28 |
| 5 Logs and alerts | what is recorded, and who is woken | chapter 28, chapter 29 |
| 6 Delivery | how code and infrastructure reach production | chapter 32, chapter 33 |
| 7 Where it runs | instances, containers or functions | chapter 16, chapter 17, chapter 18, chapter 19 |
| 8 Data | where each kind of data lives | chapter 20, chapter 21, chapter 22, chapter 25, chapter 26 |
| 9 How parts talk | queues, topics, events and streams | chapter 23 |
| 10 Traffic in | load balancers, APIs, the edge and certificates | chapter 24, chapter 27 |
| 11 Failure | zones, Regions, backups and quotas | chapter 30 |
| 12 Cost | what is billed, and who pays | chapter 31 |
| 13 Attack | keys, secrets, threats and operator access | chapter 27, chapter 28, chapter 34 |
The company answers questions 1 to 6 once, and every workload inherits them. Each workload then answers 7 to 10 with the decision figures from Part III, and tests its answers against 11 to 13. The three designs that follow each show the result the same way: the design, its decisions, how it fails, what drives its cost, and how it resists attack.
Read more AWS Well-Architected Framework
The Well-Architected pillars#
AWS's Well-Architected Framework names six qualities that designs trade against each other: operational excellence, security, reliability, performance efficiency, cost optimization and sustainability. Its questions start a conversation about decisions; AWS says a review is not an audit.
As text
- 2015: The framework, with four pillars
- 2018: AWS Well-Architected Tool, at no charge
- 2021: Sustainability, the sixth pillar
The AWS Well-Architected Tool walks a workload through those questions, records the decisions and recommends improvements. Custom lenses add a company's own questions, such as MegaCorp's rules for card data.
Microsoft's framework has five pillars: reliability, security, cost optimization, operational excellence and performance efficiency, assessed through the Azure Well-Architected Review.
AWS's has six, adding sustainability, and reviews run in the Well-Architected Tool, with custom lenses for a company's own questions.
Read more The pillars of the framework · What is AWS Well-Architected Tool? · Azure Well-Architected Framework pillars
Decision records#
An architecture decision record, or ADR, gives a decision's context, the options, the choice and its consequences, with a status: proposed, accepted, rejected or superseded. MegaCorp keeps its records in the payments repository, beside the code they govern, and reviewers cite them in code reviews.
As text
- Does it shape the structure, such as accounts, Regions or the data store? Yes: Write an ADR. No: the next step.
- Does it trade one pillar against another? Yes: Write an ADR that names the trade-off. No: the next step.
- Would it be hard or costly to undo? Yes: Write an ADR that lists the rejected options. No: the next step.
- Decide in the code review
As text
- Alex to Decision log: ADR-007, proposed: the ledger in eu-west-2, restored from backups
- Payments team to Decision log: accepted after review, with a version and stakeholders
- Alex to Decision log: ADR-019 supersedes ADR-007: a pilot light in eu-west-1
An accepted record is never edited. When the board set a one-hour recovery time, Alex wrote ADR-019 and marked ADR-007 superseded, so the log shows both what changed and why.
Unrecorded decisions come back. Without the reasons written down, the next team reopens the debate, or quietly reverses a choice that had a good reason.
Read more Architectural decision record process · Maintain an architecture decision record
A web platform#
MegaCorp's merchant portal and payments API, designed end to end: the questions answered, the choices recorded, and the design tested against failure, cost and attack.
Merchants sign in to a portal to see payments and refunds, and their systems call the payments API. Questions 1 to 6 were answered once for MegaCorp: the payments-prod account, IAM Identity Center and OIDC, a VPC on the shared hub, SCPs, the organization trail and CodePipeline. This chapter answers the rest for the web platform.
| 1 | Merchants reach the portal and the API through Front Door |
| 2 | Front Door forwards requests to the container app |
| 3 | The app writes the ledger |
| 1 | Merchants reach the portal and the API through CloudFront |
| 2 | CloudFront forwards dynamic requests to the load balancer |
| 3 | The load balancer picks a healthy task |
| 4 | The tasks write the ledger |
As text
On Azure:
- Merchants
- Resource group payments-prod
- Azure Front Door (portal)
- Virtual network 10.20.0.0/16
- Subnet apps
- Azure Container Apps (payments)
- Subnet apps
- Data
- Azure Database for PostgreSQL (ledger)
- Azure Cache for Redis (settings)
- Merchants reach the portal and the API through Front Door
- Front Door forwards requests to the container app
- The app writes the ledger
On AWS:
- Merchants
- Amazon CloudFront (portal)
- Region eu-west-2
- VPC 10.20.0.0/16
- Application Load Balancer (payments)
- AWS Fargate (payments)
- Data
- Amazon Aurora (ledger)
- Amazon ElastiCache (settings)
- VPC 10.20.0.0/16
- Merchants reach the portal and the API through CloudFront
- CloudFront forwards dynamic requests to the load balancer
- The load balancer picks a healthy task
- The tasks write the ledger
The decisions#
| Decision | Chosen | Rejected | Reason |
|---|---|---|---|
| Where it runs | ECS on Fargate, in two zones | EC2 Auto Scaling; Lambda | Java containers the team already builds, steady traffic, and no instances to patch |
| Front door | CloudFront with AWS WAF | the load balancer alone | Caches the portal near merchants and blocks attacks before they reach the Region |
| Ledger | Aurora PostgreSQL, a writer and a reader | DynamoDB; RDS | Transactions across tables, and a reader in the other zone ready to take over |
| Settings | ElastiCache, read before Aurora | Aurora for every read | Settings change rarely and are read on every request |
| Duplicate webhooks | DynamoDB conditional writes | a unique index in Aurora | Absorbs webhook bursts without adding load to the ledger |
| Merchant sign-in | a Cognito user pool | a users table of its own | Hosted sign-in and tokens, with no passwords in MegaCorp's database |
Each row is an ADR in the payments repository. The rejected options stay in the records, so the next team can see what was weighed.
Read more Architectural decision record process
When a zone fails#
| 1 | The load balancer sends requests only to healthy tasks, now all in zone b |
| 2 | The cluster endpoint now names the promoted writer in zone b |
As text
- Region eu-west-2
- VPC 10.20.0.0/16
- Application Load Balancer (payments)
- Availability Zone a (failed)
- Private subnet 10.20.10.0/24
- AWS Fargate (payments)
- Amazon Aurora (writer, lost)
- Private subnet 10.20.10.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- AWS Fargate (payments)
- Amazon Aurora (reader, promoted)
- Private subnet 10.20.11.0/24
- VPC 10.20.0.0/16
- The load balancer sends requests only to healthy tasks, now all in zone b
- The cluster endpoint now names the promoted writer in zone b
Each zone runs enough tasks to carry the whole load, so nothing has to launch during the failure. Aurora's volume already spans three zones, so the promoted reader has every committed write. The settings cache fails over to its replica, and the tasks in zone b keep using their own NAT gateway.
A cache can turn a failure into an overload. If the settings cache fails and every request falls through to Aurora, the database meets a load it never carried. AWS calls this bimodal behaviour. Size Aurora for it, or keep serving the last settings the task read.
Read more REL11-BP05 Use static stability
What drives the cost#
| Driver | Grows with | Lever |
|---|---|---|
| Fargate tasks | vCPU, memory and storage, per second | Size tasks from real use, and cover the steady base with a Compute Savings Plan |
| Aurora | instance hours, storage and I/O | Compare standard and I/O-Optimized clusters, which bill I/O differently |
| CloudFront | data out and requests | Cache the portal's static files; transfer from AWS origins is free |
| Zone to zone | gigabytes between zones, billed both ways | Accept it as the price of two zones, and keep chatty calls within one |
| NAT gateways | hours in each zone, and gigabytes processed | Send S3 and DynamoDB traffic through gateway endpoints |
Read more Understanding data transfer charges
How it resists attack#
| 1 | The web ACL blocks common attacks and floods of requests |
| 2 | CloudFront serves the requests the web ACL allows |
| 3 | Dynamic requests go to the load balancer |
| 4 | The app's group admits only the load balancer's group |
| 5 | The ledger's group admits only the app's group |
| 6 | Tasks read the ledger's credentials when they start |
| 7 | The credentials are encrypted under the payments key |
As text
- Merchants
- AWS WAF (web ACL)
- Amazon CloudFront (portal)
- Region eu-west-2
- Sign-in and secrets
- Amazon Cognito (merchants)
- AWS KMS (payments key)
- AWS Secrets Manager (ledger)
- VPC 10.20.0.0/16
- Security group alb
- Application Load Balancer (payments)
- Security group app
- AWS Fargate (payments)
- Security group ledger
- Amazon Aurora (ledger)
- Security group alb
- Sign-in and secrets
- The web ACL blocks common attacks and floods of requests
- CloudFront serves the requests the web ACL allows
- Dynamic requests go to the load balancer
- The app's group admits only the load balancer's group
- The ledger's group admits only the app's group
- Tasks read the ledger's credentials when they start
- The credentials are encrypted under the payments key
Every hop checks for itself. A request that passes the web ACL still needs a valid token from the Cognito user pool. A task needs the right group to reach the ledger, and the ledger's password lives in Secrets Manager, never in the image. GuardDuty and Security Hub watch the account from the audit account.
Read more Security groups
An event-driven system#
Card-scheme webhooks become ledger entries and events that other teams react to, without any team calling another directly. The same five views test the design.
Card schemes call MegaCorp whenever a payment clears, fails or is disputed, sometimes thousands of times a minute and sometimes twice for the same event. Settlement must post each one to the ledger exactly once, and fraud, statements and merchant alerts each need to hear about it.
| 1 | Card schemes post webhooks to the API |
| 2 | The API puts each one on a queue |
| 3 | A function settles each message |
| 4 | It records the event ID, once |
| 5 | It publishes a payment event |
| 1 | Card schemes post webhooks to the API |
| 2 | API Gateway sends each one straight to SQS |
| 3 | Lambda hands the function batches of messages |
| 4 | A conditional write records the event ID, once |
| 5 | The function puts a payment event on the bus |
As text
On Azure:
- Card schemes
- Resource group payments-prod
- Azure API Management (webhooks)
- Azure Service Bus (webhooks)
- Processing
- Azure Cosmos DB (idempotency keys)
- Azure Functions (settlement)
- Azure Event Grid (payments)
- Card schemes post webhooks to the API
- The API puts each one on a queue
- A function settles each message
- It records the event ID, once
- It publishes a payment event
On AWS:
- Card schemes
- Region eu-west-2
- Amazon API Gateway (webhooks)
- Amazon SQS (webhooks)
- Processing
- Amazon DynamoDB (idempotency keys)
- AWS Lambda (settlement)
- Amazon EventBridge (payments)
- Card schemes post webhooks to the API
- API Gateway sends each one straight to SQS
- Lambda hands the function batches of messages
- A conditional write records the event ID, once
- The function puts a payment event on the bus
The decisions#
| Decision | Chosen | Rejected | Reason |
|---|---|---|---|
| Intake | a REST API sending straight to SQS, behind AWS WAF | an HTTP API; an intake function | HTTP APIs cannot take a web ACL, and an intake function adds code and cost for nothing |
| Buffer | an SQS standard queue, with a dead-letter queue | a FIFO queue | Card schemes need no order, and the consumer already ignores duplicates |
| Settlement | Lambda, with reserved concurrency | an ECS service | Arrivals come in bursts, and the reserved limit protects the ledger |
| Duplicates | a DynamoDB conditional write per event ID, with a TTL | a unique index in Aurora | Keeps the duplicate check off the ledger, and old keys expire by themselves |
| Fan-out | an EventBridge bus, with a rule per consumer | an SNS topic | Many teams, routing on content, and consumers in other accounts |
| Card events | Kinesis Data Streams | EventBridge alone | Fraud models replay an ordered stream of card events |
Read more Choose an API Gateway API integration type
When the ledger is down#
| 1 | Lambda reads batches; a failed message becomes visible again after the timeout |
| 2 | Every write fails while the ledger is down |
| 3 | After the maximum number of receives, a message moves to the dead-letter queue |
As text
- Region eu-west-2
- Amazon SQS (dead letters)
- Amazon SQS (webhooks)
- AWS Lambda (settlement)
- Amazon Aurora (ledger) (failed)
- Lambda reads batches; a failed message becomes visible again after the timeout
- Every write fails while the ledger is down
- After the maximum number of receives, a message moves to the dead-letter queue
Nothing is lost while the ledger is down: the queue keeps messages for four days by default, and up to fourteen. The function reports only the messages that failed, so the rest are not retried, and its reserved concurrency sets the pace when the ledger returns. An alarm on the dead-letter queue's depth tells the team which webhooks need a look.
Every hop can deliver twice. SQS standard queues, EventBridge and Lambda all deliver at least once, so the idempotency keys are not optional: without them, a retry posts a payment twice.
Read more Using dead-letter queues in Amazon SQS · Handling errors for an SQS event source in Lambda
What drives the cost#
| Driver | Grows with | Lever |
|---|---|---|
| SQS | requests, each 64 KB counting as one | Send and receive in batches of up to 10 messages |
| Lambda | requests and GB-seconds | Tune memory to the work; SnapStart starts Java quickly at no extra cost |
| EventBridge | custom events, per million | Events from AWS services on the default bus are free |
| Kinesis | GB written and read on demand, or shard-hours | Switch to provisioned shards once the volume is steady |
| DynamoDB | requests on demand, and storage | Let TTL delete old keys, which uses no write capacity |
Read more Amazon SQS pricing · AWS Lambda pricing
How it resists attack#
| 1 | The web ACL filters webhooks before API Gateway sees them |
| 2 | Allowed requests reach the API |
| 3 | API Gateway's role may send only to this queue |
| 4 | Messages are encrypted at rest under the payments key |
| 5 | The function's role may only receive from this queue |
| 6 | and put events only on this bus |
| 7 | A rule sends card events to the fraud account, whose bus policy admits them |
As text
- Card schemes
- AWS WAF (web ACL)
- Region eu-west-2
- Intake
- Amazon API Gateway (webhooks)
- Amazon SQS (webhooks)
- AWS KMS (payments key)
- Processing
- AWS Lambda (settlement)
- Amazon EventBridge (payments)
- Intake
- AWS Account fraud
- Amazon EventBridge (fraud)
- The web ACL filters webhooks before API Gateway sees them
- Allowed requests reach the API
- API Gateway's role may send only to this queue
- Messages are encrypted at rest under the payments key
- The function's role may only receive from this queue
- and put events only on this bus
- A rule sends card events to the fraud account, whose bus policy admits them
Each service holds only the permissions its step needs, and each resource names who may use it: the queue policy, the key policy and the fraud account's bus policy each admit one caller.
Read more Amazon EventBridge quotas
Hybrid and multi-Region#
The payments platform in two Regions, joined to MegaCorp's data centre: one private link that reaches both, a hub in each Region, and a runbook that moves everything to eu-west-1.
The mainframe in MegaCorp's data centre still receives settlement files, and the board's one-hour recovery objective put a pilot light in eu-west-1. The network must reach both Regions from the data centre, keep working when eu-west-2 is lost, and keep development away from production.
| 1 | The circuit connects to the hub in UK South |
| 2 | and to the hub in UK West |
| 1 | Direct Connect, with a VPN as backup, reaches the gateway |
| 2 | The gateway associates with the hub in eu-west-2 |
| 3 | and with the hub in eu-west-1 |
| 4 | The hubs peer, with static routes |
As text
On Azure:
- Azure ExpressRoute (circuit)
- UK South
- Virtual network gateway (hub gateway)
- Azure Firewall (hub firewall)
- UK West
- Virtual network gateway (hub gateway)
- Azure Firewall (hub firewall)
- The circuit connects to the hub in UK South
- and to the hub in UK West
On AWS:
- Customer gateway (data centre)
- AWS Direct Connect (gateway)
- Region eu-west-2
- AWS Network Firewall (inspection)
- AWS Transit Gateway (hub)
- Region eu-west-1
- AWS Transit Gateway (hub)
- Direct Connect, with a VPN as backup, reaches the gateway
- The gateway associates with the hub in eu-west-2
- and with the hub in eu-west-1
- The hubs peer, with static routes
The decisions#
| Decision | Chosen | Rejected | Reason |
|---|---|---|---|
| Link to the data centre | Direct Connect, with Site-to-Site VPN as backup | a VPN alone | Settlement files are large and steady, and the VPN keeps a path if the link fails |
| Reaching two Regions | one Direct Connect gateway, associated with each Region's hub | a connection per Region | The gateway is global, so one link reaches both Regions |
| Hubs | a transit gateway in each Region, peered | VPC peering between every pair | Route tables keep development from production, and peering joins the Regions |
| Inspection | Network Firewall in each hub | security groups alone | Auditors want data centre traffic inspected and outbound traffic kept to known domains |
| Hybrid DNS | Resolver endpoints and shared forwarding rules | copying data centre names into Route 53 | Each side answers for its own names, and nothing is copied by hand |
| Recovery | a pilot light in eu-west-1, failed over with ARC | a warm standby | Meets the one-hour objective for less, as ADR-019 records |
Read more Direct Connect gateways · Transit gateway peering attachments
When eu-west-2 is lost#
| 1 | The link from the data centre ends on the global gateway, not in a Region |
| 2 | Routes to eu-west-1 keep working |
| 3 | Settlement files reach the recovery VPC, where the promoted ledger takes writes |
As text
- Customer gateway (data centre)
- AWS Direct Connect (gateway)
- Region eu-west-2 (failed)
- AWS Transit Gateway (hub)
- Amazon Aurora (ledger writer)
- Region eu-west-1
- AWS Transit Gateway (hub)
- Amazon Aurora (ledger, promoted)
- The link from the data centre ends on the global gateway, not in a Region
- Routes to eu-west-1 keep working
- Settlement files reach the recovery VPC, where the promoted ledger takes writes
A Direct Connect gateway is global and sits outside the data path, so losing eu-west-2 does not take it down. The on-call engineer runs the runbook from chapter 30: promote the ledger, deploy the tasks, and flip the ARC routing control.
A backup in the failed Region fails with it. A Site-to-Site VPN ends on one Region's transit gateway, so the VPN that backs up Direct Connect in eu-west-2 is gone with it. Give eu-west-1's hub a VPN of its own.
Read more What is AWS Site-to-Site VPN? · AWS Direct Connect Resiliency Toolkit
What drives the cost#
| Driver | Grows with | Lever |
|---|---|---|
| Direct Connect | port hours, and data out of AWS | Buy the port for the steady volume, and keep the VPN for failures |
| Transit gateways | attachments per hour, and gigabytes sent into each | Attach only the VPCs that need the data centre or each other |
| Between Regions | gigabytes replicated, billed as they leave the source Region | Replicate the ledger and backups, not logs and caches |
| Network Firewall | endpoints per zone per hour, and gigabytes inspected | Inspect traffic to and from the data centre, not between trusted VPCs |
| Pilot light | the secondary ledger, running all the time | Keep compute undeployed until a failover |
Read more AWS Transit Gateway pricing · AWS Network Firewall pricing
How it resists attack#
| 1 | MACsec on the Direct Connect link, or a VPN over it, encrypts the traffic |
| 2 | The hub sends data centre traffic through the firewall |
| 3 | Production's route table reaches the data centre |
| 4 | Development's route table has no route to production |
As text
- Customer gateway (data centre)
- AWS Account network
- Region eu-west-2
- AWS Transit Gateway (hub)
- VPC inspection
- AWS Network Firewall (firewall)
- Region eu-west-2
- Workloads
- AWS Account payments-prod
- Region eu-west-2
- VPC 10.20.0.0/16
- Transit gateway attachment (attachment)
- VPC 10.20.0.0/16
- Region eu-west-2
- AWS Account payments-dev
- Region eu-west-2
- VPC 10.30.0.0/16
- Transit gateway attachment (attachment)
- VPC 10.30.0.0/16
- Region eu-west-2
- AWS Account payments-prod
- MACsec on the Direct Connect link, or a VPN over it, encrypts the traffic
- The hub sends data centre traffic through the firewall
- Production's route table reaches the data centre
- Development's route table has no route to production
Direct Connect is not encrypted by default, while traffic between the two Regions' hubs is encrypted by AWS as it leaves a Region. VPCs on hubs that share one Direct Connect gateway could reach each other, so blackhole routes close that path.
Read more Encryption in AWS Direct Connect · What is AWS Network Firewall?
Design review drills#
Four designs, each with a flaw this book has taught you to see. Find it before you read the answer at the end of the chapter.
A review walks a design through the thirteen questions and asks, at every box, what happens when it fails, what it costs, and who could misuse it. Each drill below came from a real kind of mistake, and each fails at least one of those questions.
Read more AWS Well-Architected Framework
Drill 1: the way out#
| 1 | Zone a's tasks send outbound traffic to the NAT gateway |
| 2 | Zone b's tasks use the same NAT gateway |
As text
- Region eu-west-2
- VPC 10.20.0.0/16
- Availability Zone a
- Public subnet 10.20.0.0/24
- NAT gateway
- Private subnet 10.20.10.0/24
- AWS Fargate (payments)
- Public subnet 10.20.0.0/24
- Availability Zone b
- Private subnet 10.20.11.0/24
- AWS Fargate (payments)
- Private subnet 10.20.11.0/24
- Availability Zone a
- VPC 10.20.0.0/16
- Zone a's tasks send outbound traffic to the NAT gateway
- Zone b's tasks use the same NAT gateway
Read more NAT gateway basics · Regional NAT gateways
Drill 2: the ledger's door#
| 1 | Any address on the internet can reach port 5432 |
As text
- Region eu-west-2
- VPC 10.20.0.0/16
- Internet gateway
- Availability Zone a
- Public subnet 10.20.0.0/24
- Amazon Aurora (ledger)
- Public subnet 10.20.0.0/24
- VPC 10.20.0.0/16
- Any address on the internet can reach port 5432
Read more Security groups · Subnets for your VPC
Drill 3: the backups#
| 1 | Each night's ledger backup is copied to eu-west-1 |
As text
- AWS Account payments-prod
- Region eu-west-2
- AWS Backup (backup plan)
- Region eu-west-1
- AWS Backup (vault, unlocked)
- Region eu-west-2
- Each night's ledger backup is copied to eu-west-1
Read more AWS Backup Vault Lock
Drill 4: the consumer#
| 1 | Batches of ten; any error throws, failing the whole batch |
| 2 | Each message posts a payment, with no check for repeats |
As text
- Region eu-west-2
- Amazon SQS (webhooks)
- AWS Lambda (settlement)
- Amazon Aurora (ledger)
- Batches of ten; any error throws, failing the whole batch
- Each message posts a payment, with no check for repeats
A diagram shows only what someone drew. In a review, ask for the route tables, security groups, key policies and quotas behind each box; the flaws in these drills all hide there.
Read more Using dead-letter queues in Amazon SQS · Handling errors for an SQS event source in Lambda
- Drill 1: what fails, and what does it cost?
Answer
This NAT gateway lives in zone a. If zone a fails, zone b loses its way out too, and zone b's traffic crosses zones, billed both ways. Give each zone its own NAT gateway and route table, or use a regional NAT gateway. - Drill 2: what would you change?
Answer
Move the ledger to private subnets, and let its security group admit only the payments tasks' group. Nothing on the internet should reach a database port. - Drill 3: why is this not enough?
Answer
Anyone who can administer payments-prod can delete both copies. Copy to a vault in another account, locked in compliance mode, so not even the root user can delete a backup early. - Drill 4: what goes wrong under load?
Answer
Delivery is at least once and a thrown error returns the whole batch, so payments post twice. Report only the failed messages, check an idempotency key with a DynamoDB conditional write, and add a dead-letter queue.
Service map#
Every AWS service this book names, with its Azure counterparts, in both directions. Start from a service you know and follow it to the figure that shows its AWS match.
The first table runs from AWS to Azure, grouped by category; the second runs from Azure to AWS. "none" means Azure has no single counterpart, and the service's profile or note says how Azure does the job. A service shown in a figure links to the first one, and every row links to its documentation.
From AWS to Azure
| AWS service | On Azure | Category | Where |
|---|---|---|---|
| AWS Glue | Azure Data Factory, Microsoft Fabric | Analytics | shown here · docs |
| AWS Lake Formation | Microsoft Purview | Analytics | shown here · docs |
| Amazon Athena | Microsoft Fabric, Azure Synapse Analytics | Analytics | shown here · docs |
| Amazon Kinesis Data Streams | Azure Event Hubs | Analytics | shown here · docs |
| Amazon MSK | Azure Event Hubs | Analytics | docs |
| Amazon Redshift | Microsoft Fabric, Azure Synapse Analytics | Analytics | shown here · docs |
| AWS Step Functions | Azure Logic Apps, Durable Functions | Application integration | shown here · docs |
| Amazon EventBridge | Azure Event Grid | Application integration | shown here · docs |
| Amazon MQ | Azure Service Bus | Application integration | docs |
| Amazon SNS | Azure Service Bus, Azure Event Grid | Application integration | shown here · docs |
| Amazon SQS | Azure Service Bus, Azure Queue Storage | Application integration | shown here · docs |
| Amazon Bedrock | Microsoft Foundry | Artificial intelligence | shown here · docs |
| Amazon SES | Azure Communication Services | Business applications | docs |
| AWS Budgets | Microsoft Cost Management | Cloud financial management | shown here · docs |
| AWS Cost Explorer | Microsoft Cost Management | Cloud financial management | shown here · docs |
| Savings Plans | Savings plan for compute, Azure Reservations | Cloud financial management | shown here · docs |
| AWS App Runner | Azure App Service, Azure Container Apps | Compute | docs |
| AWS Elastic Beanstalk | Azure App Service | Compute | shown here · docs |
| AWS Lambda | Azure Functions | Compute | shown here · docs |
| Amazon EC2 | Azure Virtual Machines | Compute | shown here · docs |
| Amazon EC2 Auto Scaling | Azure Virtual Machine Scale Sets | Compute | shown here · docs |
| Amazon Lightsail | none | Compute | docs |
| AWS Fargate | Azure Container Apps, Azure Container Instances | Containers | shown here · docs |
| Amazon ECR | Azure Container Registry | Containers | shown here · docs |
| Amazon ECS | Azure Container Apps | Containers | shown here · docs |
| Amazon EKS | Azure Kubernetes Service | Containers | shown here · docs |
| Amazon Aurora | Azure Database for PostgreSQL, Azure SQL Database Hyperscale | Databases | shown here · docs |
| Amazon DocumentDB | Azure Cosmos DB for MongoDB | Databases | docs |
| Amazon DynamoDB | Azure Cosmos DB | Databases | shown here · docs |
| Amazon ElastiCache | Azure Managed Redis, Azure Cache for Redis | Databases | shown here · docs |
| Amazon RDS | Azure SQL Database, Azure Database for PostgreSQL, Azure Database for MySQL | Databases | shown here · docs |
| AWS CDK | none | Developer tools | shown here · docs |
| AWS CloudFormation | Azure Resource Manager templates, Bicep | Developer tools | shown here · docs |
| AWS CodeBuild | Azure Pipelines | Developer tools | shown here · docs |
| AWS CodeDeploy | Azure Pipelines | Developer tools | docs |
| AWS CodePipeline | Azure Pipelines | Developer tools | shown here · docs |
| AWS Amplify | Azure Static Web Apps | Front-end web and mobile | docs |
| AWS AppConfig | Azure App Configuration | Management and governance | shown here · docs |
| AWS CloudTrail | Azure Monitor | Management and governance | shown here · docs |
| AWS Config | Azure Policy, Azure Resource Graph | Management and governance | shown here · docs |
| AWS Control Tower | Azure landing zone | Management and governance | shown here · docs |
| AWS Fault Injection Service | Azure Chaos Studio | Management and governance | shown here · docs |
| AWS Organizations | Azure management groups | Management and governance | shown here · docs |
| AWS Systems Manager | Azure Automation, Azure Update Manager, Azure Arc | Management and governance | shown here · docs |
| AWS Well-Architected Tool | Azure Well-Architected Review | Management and governance | docs |
| AWS X-Ray | Application Insights | Management and governance | shown here · docs |
| Amazon CloudWatch | Azure Monitor | Management and governance | shown here · docs |
| Service Quotas | none | Management and governance | docs |
| AWS Application Migration Service | Azure Migrate | Migration and modernization | docs |
| AWS DataSync | Azure Storage Mover | Migration and modernization | docs |
| AWS Database Migration Service | Azure Database Migration Service | Migration and modernization | docs |
| AWS Transfer Family | SFTP support for Azure Blob Storage | Migration and modernization | docs |
| AWS Cloud WAN | Azure Virtual WAN | Networking and content delivery | docs |
| AWS Direct Connect | Azure ExpressRoute | Networking and content delivery | shown here · docs |
| AWS Global Accelerator | Azure Front Door, Azure Load Balancer | Networking and content delivery | shown here · docs |
| AWS Network Firewall | Azure Firewall | Networking and content delivery | shown here · docs |
| AWS PrivateLink | Azure Private Link | Networking and content delivery | shown here · docs |
| AWS Site-to-Site VPN | Azure VPN Gateway | Networking and content delivery | shown here · docs |
| AWS Transit Gateway | Azure Virtual WAN, Azure virtual network peering | Networking and content delivery | shown here · docs |
| Amazon API Gateway | Azure API Management | Networking and content delivery | shown here · docs |
| Amazon Application Recovery Controller | none | Networking and content delivery | shown here · docs |
| Amazon CloudFront | Azure Front Door | Networking and content delivery | shown here · docs |
| Amazon Route 53 | Azure DNS, Azure Traffic Manager | Networking and content delivery | shown here · docs |
| Amazon VPC | Azure Virtual Network | Networking and content delivery | shown here · docs |
| Elastic Load Balancing | Azure Load Balancer, Azure Application Gateway | Networking and content delivery | shown here · docs |
| AWS Certificate Manager | Azure Key Vault | Security, identity and compliance | shown here · docs |
| AWS IAM | Azure role-based access control, Microsoft Entra ID | Security, identity and compliance | shown here · docs |
| AWS IAM Identity Center | Microsoft Entra ID | Security, identity and compliance | shown here · docs |
| AWS KMS | Azure Key Vault | Security, identity and compliance | shown here · docs |
| AWS Resource Access Manager | none | Security, identity and compliance | shown here · docs |
| AWS Secrets Manager | Azure Key Vault | Security, identity and compliance | shown here · docs |
| AWS Security Hub | Microsoft Defender for Cloud | Security, identity and compliance | shown here · docs |
| AWS Security Token Service | Microsoft Entra ID | Security, identity and compliance | shown here · docs |
| AWS Shield | Azure DDoS Protection | Security, identity and compliance | docs |
| AWS WAF | Azure Web Application Firewall | Security, identity and compliance | shown here · docs |
| Amazon Cognito | Microsoft Entra External ID | Security, identity and compliance | shown here · docs |
| Amazon GuardDuty | Microsoft Defender for Cloud | Security, identity and compliance | shown here · docs |
| Amazon Inspector | Microsoft Defender for Cloud | Security, identity and compliance | docs |
| AWS Backup | Azure Backup | Storage | shown here · docs |
| AWS Elastic Disaster Recovery | Azure Site Recovery | Storage | docs |
| Amazon EBS | Azure managed disks | Storage | shown here · docs |
| Amazon EFS | Azure Files | Storage | shown here · docs |
| Amazon FSx | Azure Files, Azure NetApp Files | Storage | shown here · docs |
| Amazon S3 | Azure Blob Storage | Storage | shown here · docs |
From Azure to AWS
| On Azure | AWS service | Where |
|---|---|---|
| Application Insights | AWS X-Ray | shown here |
| Azure API Management | Amazon API Gateway | shown here |
| Azure App Configuration | AWS AppConfig | shown here |
| Azure App Service | AWS App Runner | |
| Azure App Service | AWS Elastic Beanstalk | shown here |
| Azure Application Gateway | Elastic Load Balancing | shown here |
| Azure Arc | AWS Systems Manager | shown here |
| Azure Automation | AWS Systems Manager | shown here |
| Azure Backup | AWS Backup | shown here |
| Azure Blob Storage | Amazon S3 | shown here |
| Azure Cache for Redis | Amazon ElastiCache | shown here |
| Azure Chaos Studio | AWS Fault Injection Service | shown here |
| Azure Communication Services | Amazon SES | |
| Azure Container Apps | AWS App Runner | |
| Azure Container Apps | AWS Fargate | shown here |
| Azure Container Apps | Amazon ECS | shown here |
| Azure Container Instances | AWS Fargate | shown here |
| Azure Container Registry | Amazon ECR | shown here |
| Azure Cosmos DB | Amazon DynamoDB | shown here |
| Azure Cosmos DB for MongoDB | Amazon DocumentDB | |
| Azure Data Factory | AWS Glue | shown here |
| Azure Database for MySQL | Amazon RDS | shown here |
| Azure Database for PostgreSQL | Amazon Aurora | shown here |
| Azure Database for PostgreSQL | Amazon RDS | shown here |
| Azure Database Migration Service | AWS Database Migration Service | |
| Azure DDoS Protection | AWS Shield | |
| Azure DNS | Amazon Route 53 | shown here |
| Azure Event Grid | Amazon EventBridge | shown here |
| Azure Event Grid | Amazon SNS | shown here |
| Azure Event Hubs | Amazon Kinesis Data Streams | shown here |
| Azure Event Hubs | Amazon MSK | |
| Azure ExpressRoute | AWS Direct Connect | shown here |
| Azure Files | Amazon EFS | shown here |
| Azure Files | Amazon FSx | shown here |
| Azure Firewall | AWS Network Firewall | shown here |
| Azure Front Door | AWS Global Accelerator | shown here |
| Azure Front Door | Amazon CloudFront | shown here |
| Azure Functions | AWS Lambda | shown here |
| Azure Key Vault | AWS Certificate Manager | shown here |
| Azure Key Vault | AWS KMS | shown here |
| Azure Key Vault | AWS Secrets Manager | shown here |
| Azure Kubernetes Service | Amazon EKS | shown here |
| Azure landing zone | AWS Control Tower | shown here |
| Azure Load Balancer | AWS Global Accelerator | shown here |
| Azure Load Balancer | Elastic Load Balancing | shown here |
| Azure Logic Apps | AWS Step Functions | shown here |
| Azure managed disks | Amazon EBS | shown here |
| Azure Managed Redis | Amazon ElastiCache | shown here |
| Azure management groups | AWS Organizations | shown here |
| Azure Migrate | AWS Application Migration Service | |
| Azure Monitor | AWS CloudTrail | shown here |
| Azure Monitor | Amazon CloudWatch | shown here |
| Azure NetApp Files | Amazon FSx | shown here |
| Azure Pipelines | AWS CodeBuild | shown here |
| Azure Pipelines | AWS CodeDeploy | |
| Azure Pipelines | AWS CodePipeline | shown here |
| Azure Policy | AWS Config | shown here |
| Azure Private Link | AWS PrivateLink | shown here |
| Azure Queue Storage | Amazon SQS | shown here |
| Azure Reservations | Savings Plans | shown here |
| Azure Resource Graph | AWS Config | shown here |
| Azure Resource Manager templates | AWS CloudFormation | shown here |
| Azure role-based access control | AWS IAM | shown here |
| Azure Service Bus | Amazon MQ | |
| Azure Service Bus | Amazon SNS | shown here |
| Azure Service Bus | Amazon SQS | shown here |
| Azure Site Recovery | AWS Elastic Disaster Recovery | |
| Azure SQL Database | Amazon RDS | shown here |
| Azure SQL Database Hyperscale | Amazon Aurora | shown here |
| Azure Static Web Apps | AWS Amplify | |
| Azure Storage Mover | AWS DataSync | |
| Azure Synapse Analytics | Amazon Athena | shown here |
| Azure Synapse Analytics | Amazon Redshift | shown here |
| Azure Traffic Manager | Amazon Route 53 | shown here |
| Azure Update Manager | AWS Systems Manager | shown here |
| Azure Virtual Machine Scale Sets | Amazon EC2 Auto Scaling | shown here |
| Azure Virtual Machines | Amazon EC2 | shown here |
| Azure Virtual Network | Amazon VPC | shown here |
| Azure virtual network peering | AWS Transit Gateway | shown here |
| Azure Virtual WAN | AWS Cloud WAN | |
| Azure Virtual WAN | AWS Transit Gateway | shown here |
| Azure VPN Gateway | AWS Site-to-Site VPN | shown here |
| Azure Web Application Firewall | AWS WAF | shown here |
| Azure Well-Architected Review | AWS Well-Architected Tool | |
| Bicep | AWS CloudFormation | shown here |
| Durable Functions | AWS Step Functions | shown here |
| Microsoft Cost Management | AWS Budgets | shown here |
| Microsoft Cost Management | AWS Cost Explorer | shown here |
| Microsoft Defender for Cloud | AWS Security Hub | shown here |
| Microsoft Defender for Cloud | Amazon GuardDuty | shown here |
| Microsoft Defender for Cloud | Amazon Inspector | |
| Microsoft Entra External ID | Amazon Cognito | shown here |
| Microsoft Entra ID | AWS IAM | shown here |
| Microsoft Entra ID | AWS IAM Identity Center | shown here |
| Microsoft Entra ID | AWS Security Token Service | shown here |
| Microsoft Fabric | AWS Glue | shown here |
| Microsoft Fabric | Amazon Athena | shown here |
| Microsoft Fabric | Amazon Redshift | shown here |
| Microsoft Foundry | Amazon Bedrock | shown here |
| Microsoft Purview | AWS Lake Formation | shown here |
| Savings plan for compute | Savings Plans | shown here |
| SFTP support for Azure Blob Storage | AWS Transfer Family |
CLI side by side#
The Azure CLI and the AWS CLI do the same daily jobs with different words. Both query their JSON output with JMESPath, so most of what you know about --query carries over.
Signing in, and choosing where you work#
| Task | Azure CLI | AWS CLI |
|---|---|---|
| Set up sign-in once | <code>az login</code> | <code>aws configure sso</code>, which writes a profile to the <code>config</code> file |
| Sign in | <code>az login</code> | <code>aws sso login --profile my-dev-profile</code> |
| Who am I? | <code>az account show</code> | <code>aws sts get-caller-identity</code> |
| What can I choose from? | <code>az account list</code> | <code>aws configure list-profiles</code> |
| Switch where I work | <code>az account set --subscription "My Demos"</code> | <code>export AWS_PROFILE=my-dev-profile</code>, or <code>--profile</code> on each command |
| Change a setting | per command, such as <code>--subscription</code> | <code>aws configure set region eu-west-2 --profile my-dev-profile</code> |
| Sign out | <code>az logout</code> | <code>aws sso logout</code> |
An AWS profile names one account, one role and a default Region, so switching accounts means switching profiles. While the Identity Center sign-in lasts, the CLI renews the role's credentials by itself.
$ aws sso login --profile my-dev-profile
SSO authorization page has automatically been opened in your default browser.
Follow the instructions in the browser to complete this authorization request.
Successfully logged into Start URL: https://my-sso-portal.awsapps.com/startRead more Configuring IAM Identity Center authentication with the AWS CLI · Configuration and credential file settings in the AWS CLI · Manage Azure subscriptions with the Azure CLI
Querying output#
Both CLIs take --query with a JMESPath expression and --output table for
people. The AWS CLI also passes server-side filters, such as --filters, which each API
defines; they cut what comes back before --query shapes it.
az vm list --resource-group QueryDemo \
--query "[?storageProfile.osDisk.osType=='Linux'].{Name:name, admin:osProfile.adminUsername}" \
--output table
aws ec2 describe-volumes --query 'Volumes[?Size < `20`].VolumeId'[
"vol-2e410a47",
"vol-a1b3c7nd"
]Text output queries each page. With --output text, the AWS CLI splits
the results into pages first and runs --query on each, so a query for the first match
returns one per page. Use JSON output when the query must see everything at once.
Read more Filtering output in the AWS CLI · Query Azure CLI command results
IAM policy cookbook#
Five policies the payments team writes again and again, quoted from the book's Terraform: what each allows, why it is shaped that way, and what to change before you reuse it.
When Alex needs a new permission, the platform team starts from one of these recipes rather than a
blank page. Each is an aws_iam_policy_document data source, which Terraform renders as
the JSON policy that IAM stores. Its statements have the shape shown in chapter 9: an
effect, actions, resources and optional conditions. The account and key IDs are placeholders.
| Recipe | Kind of policy | Attached to | See |
|---|---|---|---|
| Read and write one bucket | identity-based | the settlement job's role | chapter 20 |
| Consume one queue | identity-based | the settlement function's execution role | chapter 23 |
| Let a pipeline deploy | trust policy | the deploy role | chapter 10 |
| Keep every account in two Regions | service control policy | the Workloads OU | chapter 7 |
| Let tags decide | identity-based | every team's role | chapter 9 |
Read and write one bucket#
The settlement job writes each day's statements to megacorp-statements and reads them
back to reconcile. The bucket encrypts objects with the payments team's own KMS key, so the role needs
the key as well as the bucket.
# The settlement job writes statements to one bucket and reads them back.
data "aws_iam_policy_document" "statements_read_write" {
statement {
sid = "ListTheBucket"
actions = ["s3:ListBucket"]
resources = ["arn:aws:s3:::megacorp-statements"]
}
statement {
sid = "ReadAndWriteObjects"
actions = ["s3:GetObject", "s3:PutObject"]
resources = ["arn:aws:s3:::megacorp-statements/*"]
}
statement {
sid = "UseTheKeyOnlyThroughS3"
actions = ["kms:GenerateDataKey", "kms:Decrypt"]
resources = ["arn:aws:kms:eu-west-2:111111111111:key/1234abcd-12ab-34cd-56ef-1234567890ab"]
condition {
test = "StringEquals"
variable = "kms:ViaService"
values = ["s3.eu-west-2.amazonaws.com"]
}
}
}Listing is an action on the bucket, so its resource is the bucket's ARN. Reading and writing are
actions on objects, so theirs is the bucket's ARN followed by /*. IAM's own example
allows s3:*Object, which matches every action whose name ends in Object, including
DeleteObject. Naming the two actions keeps a bug in the job from deleting
statements.
Writing an object encrypted with a KMS key needs kms:GenerateDataKey, and reading one
needs kms:Decrypt. The kms:ViaService condition lets the role use the key
only for requests that come through S3 in London, so the job cannot take the key and call KMS itself.
To reuse the recipe, change the bucket, the key's ARN and the Region in the condition.
The bucket and its objects are different resources. Put
s3:ListBucket on the object ARN, or the object actions on the bucket ARN, and the
statement matches nothing: the job is refused although the policy looks right.
Read more IAM example: read and write objects in one bucket · Using server-side encryption with AWS KMS keys (SSE-KMS) · kms:ViaService condition key
Consume one queue#
Card-scheme webhooks land on the webhooks queue, and the settlement function settles
each one, as in chapter 19. Lambda polls the queue for the function with the function's
execution role, so that role needs the right to receive, delete and inspect messages.
# The settlement function's execution role: Lambda polls one queue on its behalf.
data "aws_iam_policy_document" "settlement_consumer" {
statement {
sid = "ConsumeOneQueue"
actions = [
"sqs:ReceiveMessage",
"sqs:DeleteMessage",
"sqs:GetQueueAttributes",
]
resources = ["arn:aws:sqs:eu-west-2:111111111111:webhooks"]
}
statement {
sid = "ReadEncryptedMessages"
actions = ["kms:Decrypt"]
resources = ["arn:aws:kms:eu-west-2:111111111111:key/1234abcd-12ab-34cd-56ef-1234567890ab"]
}
}These are the three SQS actions in the AWS managed policy
AWSLambdaSQSQueueExecutionRole. That policy also allows three logs: actions
so the function can write its logs; add them too, limited to the function's log group. A queue
encrypted with a customer managed key also needs kms:Decrypt, as here. The function and
the queue must be in the same Region.
The managed policy reaches every queue.
AWSLambdaSQSQueueExecutionRole allows its actions on every resource, so a function given
it can read and delete messages from any queue in the account. Start with it if you must, then
replace it with a policy that names the queue.
Read more AWSLambdaSQSQueueExecutionRole managed policy
Let a pipeline deploy#
The payments pipeline runs in GitHub Actions and deploys with no stored keys, as in chapter 10. This is a trust policy: rather than what the role may do, it says who may assume it. Here, that is a workflow on the main branch of one repository, vouched for by the account's OIDC provider for GitHub.
# The deploy role's trust policy: only the main branch of one repository may assume it.
data "aws_iam_policy_document" "deploy_trust" {
statement {
actions = ["sts:AssumeRoleWithWebIdentity"]
principals {
type = "Federated"
identifiers = ["arn:aws:iam::111111111111:oidc-provider/token.actions.githubusercontent.com"]
}
condition {
test = "StringEquals"
variable = "token.actions.githubusercontent.com:aud"
values = ["sts.amazonaws.com"]
}
condition {
test = "StringEquals"
variable = "token.actions.githubusercontent.com:sub"
values = ["repo:megacorp/payments:ref:refs/heads/main"]
}
}
}The aud condition checks that the token was issued for AWS STS, and sub
names the repository and branch. For GitHub, IAM requires a sub condition that is present
and is more than a wildcard. To reuse the recipe, change the account, the repository and the
branch.
A broad sub hands the role to others. A sub that names
no organization or repository lets workflows in repositories you don't control assume the role. One
that names only the organization lets any of its repositories deploy. Name one repository and the
branch that deploys.
Read more Create a role for OpenID Connect federation
Keep every account in two Regions#
MegaCorp keeps payment data in London, with its recovery copy in Ireland, as in chapter 30. The platform team makes that a guardrail: a service control policy on the Workloads OU refuses any request made to another Region.
# A service control policy: every account works in London and Ireland, except global services.
data "aws_iam_policy_document" "deny_outside_regions" {
statement {
sid = "DenyOutsideLondonAndIreland"
effect = "Deny"
not_actions = [
"cloudfront:*",
"iam:*",
"organizations:*",
"route53:*",
"support:*",
]
resources = ["*"]
condition {
test = "StringNotEquals"
variable = "aws:RequestedRegion"
values = ["eu-west-2", "eu-west-1"]
}
}
}aws:RequestedRegion names the Region a request was made to. A few popular global
services, among them IAM, CloudFront and Route 53, have a single endpoint in us-east-1, so a plain
Region deny would block them; not_actions lists them as exceptions. Like every service
control policy, this one grants nothing, so each role still needs its own permissions.
A deny with not_actions lists exceptions, not permissions. The
statement denies every other action outside the two Regions. When a team starts using another global
service, add it to the list, or its calls are refused everywhere. Attach the policy to a test OU
first: a wrong guardrail stops every account below it at once.
Read more IAM example: deny access based on the requested Region
Let tags decide#
Each team keeps its database passwords in Secrets Manager. Each team's role carries a
team tag, and so does each secret. chapter 9 showed the idea; this is the
policy, one document attached to every team's role.
# One policy for every team's role: a role reads the secrets tagged with its own team.
data "aws_iam_policy_document" "own_team_secrets" {
statement {
sid = "ReadOwnTeamSecrets"
actions = ["secretsmanager:DescribeSecret", "secretsmanager:GetSecretValue"]
resources = ["*"]
condition {
test = "StringEquals"
variable = "aws:ResourceTag/team"
values = ["&{aws:PrincipalTag/team}"]
}
}
statement {
sid = "KeepTheTeamTag"
effect = "Deny"
actions = ["secretsmanager:UntagResource"]
resources = ["*"]
condition {
test = "ForAnyValue:StringEquals"
variable = "aws:TagKeys"
values = ["team"]
}
}
}The first statement lets a role read a secret only when the secret's team tag equals
the role's own. A new secret tagged team=ledger is readable by the ledger role at once,
with no policy change. The second statement stops a team role from removing that tag, because the
tag is now a permission.
Terraform writes IAM's policy variables with &{. IAM spells the
caller's tag ${aws:PrincipalTag/team}, which Terraform would take for its own
interpolation. Write &{aws:PrincipalTag/team} in aws_iam_policy_document,
and IAM receives this condition:
"Condition": {
"StringEquals": {
"aws:ResourceTag/team": "${aws:PrincipalTag/team}"
}
}A tag carries one value, so a person who works for two teams needs two roles. Actions that don't
act on one secret, such as listing secrets, ignore resource tags and need a statement of their own. A
broader policy on the same role, such as AdministratorAccess, is not narrowed by these
conditions.
Read more IAM tutorial: permissions based on tags · aws_iam_policy_document data source
Before a policy ships#
Validate each new policy with IAM Access Analyzer, which runs more than 100 checks on it, and let an external access analyzer watch for anything shared outside the organization. When a request is still refused, the message often names the kind of policy that refused it: look it up in chapter 43.
Read more Using IAM Access Analyzer · Policy evaluation logic
Error index#
Messages you will meet in your first months on AWS, each quoted from the documentation, with what it means and how to fix it. Paste a message into search to land on its entry.
Account IDs, names and ARNs in these messages are the documentation's examples; yours will show your own. The rest of each message is as AWS writes it.
Access denied#
An access denied message names the caller, the action and, often, the kind of policy that refused. When several kinds refuse, it names only one, so check the others too.
User: arn:aws:iam::123456789012:role/HR is not authorized to perform: codecommit:ListRepositories
because no identity-based policy allows the codecommit:ListRepositories actionAllow for the action, and the resource, to a policy attached to the role or
its permission set.User: arn:aws:iam::123456789012:user/John is not authorized to perform: codecommit:ListRepositories
with an explicit deny in a service control policyUser: arn:aws:iam::123456789012:user/John is not authorized to perform: codedeploy:ListDeployments
on resource: arn:aws:codedeploy:us-east-1:123456789012:deploymentgroup:*
because no permissions boundary allows the codedeploy:ListDeployments actionUser: arn:aws:iam::123456789012:user/John is not authorized to perform: sts:AssumeRole
because no role trust policy allows the sts:AssumeRole actionsts:AssumeRole on the role.Credentials and the CLI#
An error occurred (InvalidClientTokenId) when calling the ListBuckets operation: The security token
included in the request is invalid.aws configure list to see which credentials and profile are in use, then sign in
again with aws sso login.An error occurred (SignatureDoesNotMatch) when calling the ListBuckets operation: The request
signature we calculated does not match the signature you provided. Check your key and signing
method.[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failedca_bundle setting,
--ca-bundle or AWS_CA_BUNDLE.Lambda#
Task timed out after 3.00 secondsUser: arn:aws:iam::123456789012:user/developer is not authorized to perform: lambda:InvokeFunction
on resource: my-functionlambda:InvokeFunction on the function. Another service or account needs
permission in the function's resource-based policy.KMSDisabledException: Lambda was unable to decrypt the environment variables because the KMS key
used is disabled. Please check the function's KMS key settings.Containers#
API error (500): Get https://111122223333.dkr.ecr.us-east-1.amazonaws.com/v2/: net/http: request
canceled while waiting for connectionThe task can’t pull the image. Check that the role has the permissions to pull images from the
registry.CannotPullContainerError: pull image manifest has been retried 5 time(s): failed to resolve
ref:latest.EC2#
InsufficientInstanceCapacityInstanceLimitExceededYou are not authorized to perform this operation.ec2:RunInstances or
iam:PassRole for the instance's role.The AWS you'll inherit#
An existing estate carries the AWS of the years it was built in. Each row names something you may find, what it tells you about its age, and what AWS offers today, so you can tell history from a mistake.
Old is not broken: most of what follows still works, and some of it will run for years. Replace it when you are changing that part anyway, or when a date below forces your hand.
| What you find | What it tells you | Today | See |
|---|---|---|---|
| EC2-Classic in old runbooks or scripts | an account from before 4 December 2013; EC2-Classic was retired, and marked deprecated on 31 July 2023 | VPCs | chapter 12 |
| A default VPC in every Region | an account created after 4 December 2013 | VPCs you plan yourself | chapter 12 |
| Classic Load Balancers | built before the Application Load Balancer arrived, on 11 August 2016 | Application or Network Load Balancers | chapter 24 |
| Auto Scaling launch configurations | built before launch templates; accounts created since 1 October 2024 cannot create them | launch templates | chapter 17 |
| gp2 EBS volumes | created before gp3 arrived, on 1 December 2020 | gp3, changed in place with Elastic Volumes | chapter 20 |
| An Origin Access Identity on a CloudFront distribution | set up before Origin Access Control arrived, on 25 August 2022 | Origin Access Control | chapter 24 |
| S3 buckets with ACLs, or open to the public | created before new buckets got Block Public Access and ACLs off, on 28 April 2023 | Block Public Access, with ACLs off | chapter 20 |
| CloudWatch Events rules | written after 14 January 2016; the service became Amazon EventBridge in 2019 | the same rules, in EventBridge | chapter 23 |
| AWS Single Sign-On | set up between 7 December 2017 and its rename on 26 July 2022 | IAM Identity Center, the same service | chapter 10 |
| Reserved Instances | bought before Savings Plans arrived, on 6 November 2019 | Savings Plans | chapter 31 |
| The AWS SDK for Java 1.x | code from before its end of support, on 31 December 2025 | the AWS SDK for Java 2.x | chapter 19 |
| AWS CDK v1 apps | built before v1's support ended, on 1 June 2023 | AWS CDK v2 | chapter 32 |
| Amazon Linux 2 | images from before its end of life, on 30 June 2026 | Amazon Linux 2023 | chapter 17 |
| Terraform state locked with DynamoDB | an older S3 backend; DynamoDB locking is deprecated | the S3 lock file, <code>use_lockfile</code> | chapter 32 |
| ECS blue/green through CodeDeploy | built before ECS added its own, on 17 July 2025 | ECS built-in blue/green | chapter 33 |
| AWS App Runner services | a service no longer open to new customers | Amazon ECS Express Mode | chapter 16 |
| AWS Audit Manager | a service no longer open to new customers | Security Hub's standards and Config's rules | chapter 28 |
| ARC readiness checks | not offered to new customers from 30 April 2026 | ARC Region switch plans | chapter 30 |
| The Systems Manager CloudWatch dashboard | unavailable after 30 April 2026 | CloudWatch dashboards | chapter 29 |
To date something the table doesn't cover, look it up in the service's document history, which lists each change by date, or in AWS What's New. AWS Config shows how a resource's configuration has changed, and CloudTrail's event history shows the last 90 days of calls.
Read more Restrict access to an Amazon S3 origin · Blocking public access to your Amazon S3 storage · Auto Scaling launch configurations · Amazon EBS General Purpose SSD volumes · End of support for the AWS SDK for Java 1.x
Glossary#
The general ideas this book explains in primers, in alphabetical order. Each term links to its primer, where the idea is explained with an example.
| Term | Meaning |
|---|---|
| Availability Zone | An Availability Zone is one or more data centres with their own power, networking and connectivity. The zones in a Region sit far enough apart, up to about 100 km, not to fail together, and close enough for synchronous replication within a few milliseconds. Every Region has three or more, as of September 2026. |
| CIDR block | A CIDR block writes an address range as a base address and a prefix length. 10.20.0.0/16 fixes the first 16 of 32 bits and leaves 16 free: 65,536 addresses. Each extra bit of prefix halves the range, so a /24 holds 256. Ranges that share any address overlap. |
Credits and sources#
Where the icons, the facts and the code come from, and whose names they are.
Icons#
The architecture figures use the official AWS Architecture Icons, package Icon-package_07312026, and the Azure architecture icons, set Azure_Public_Service_Icons_V24, each under its vendor's terms. The icons appear as their packages draw them: never recoloured, rotated or cropped. A failed component is shown by an outline and a label, not by changing its icon.
Read more AWS Architecture Icons · Azure architecture icons
Trademarks#
Amazon Web Services, AWS and the names and icons of AWS services are trademarks of Amazon.com, Inc. or its affiliates. Microsoft, Azure and the names and icons of Azure services are trademarks of the Microsoft group of companies. Other names are trademarks of their owners. This book is not affiliated with, or endorsed by, Amazon or Microsoft.
Primary sources#
Every version, date, status, quota, limit and default in this book comes from the research record, where each fact names its source and the date it was fetched. The sources are AWS documentation, including each service's document history; AWS What's New and the AWS News Blog; Microsoft Learn and Azure Updates; the Terraform registry and HashiCorp documentation; and OpenJDK and Maven Central. Quotas and defaults are stated as of as of September 2026, and costs are described without amounts, because prices change.
No AWS account was used to write this book. The Terraform and Java examples come from a specimen project that is formatted, validated and compiled offline, and command output is quoted only from the documentation, labelled with its source, or marked as illustrative.
Cheat card#
The book on one page: the thirteen design questions with MegaCorp's answers, and the words that mean something else on AWS.
Thirteen questions, and MegaCorp's answers#
Start a design from these answers, and write an ADR wherever yours differ.
| Question | MegaCorp's answer | See |
|---|---|---|
| 1 Accounts | an account per workload and environment, in OUs, with Control Tower's log-archive and audit accounts | chapter 15 |
| 2 Sign-in | IAM Identity Center for people, roles for code, OIDC for pipelines; no long-lived keys | chapter 10 |
| 3 Networks | eu-west-2, recovering to eu-west-1; a VPC per account, attached to a transit gateway hub | chapter 13 |
| 4 Guardrails | Control Tower controls and SCPs on every OU; Security Hub and Config check posture | chapter 15 |
| 5 Logs and alerts | the organization trail into log-archive, GuardDuty, and CloudWatch alarms that wake someone | chapter 28 |
| 6 Delivery | Terraform with state in S3, deployed by CodePipeline with GitHub as the source | chapter 33 |
| 7 Where it runs | ECS on Fargate; Lambda for short work triggered by events; EC2 when the host matters | chapter 16 |
| 8 Data | Aurora PostgreSQL for the ledger, DynamoDB for keys and duplicates, S3 for files | chapter 21 |
| 9 How parts talk | SQS to buffer work, EventBridge to fan out, Kinesis for ordered streams | chapter 23 |
| 10 Traffic in | CloudFront with AWS WAF, then an Application Load Balancer; certificates from ACM | chapter 24 |
| 11 Failure | two zones, each able to carry the load; AWS Backup; a pilot light in eu-west-1 | chapter 30 |
| 12 Cost | tags on every resource, AWS Budgets alerts, and a Savings Plan for the steady base | chapter 31 |
| 13 Attack | customer managed keys, Secrets Manager, least-privilege roles, GuardDuty, and Session Manager for operators | chapter 34 |
Words that mean something else#
| Word | On Azure | On AWS | See |
|---|---|---|---|
| role | permissions assigned to a user, group or identity at a scope | an identity that people and services assume for a limited time | chapter 5 |
| policy | Azure Policy: rules for how resources are configured | a JSON document of permissions; AWS Config rules check configuration | chapter 5 |
| security group | allow and deny rules in priority order, on a subnet or a network interface | allow rules only, on a network interface, never on a subnet | chapter 5 |
| resource group | a container: deleting it deletes what it holds | a saved query over tags or a stack; the account is the container | chapter 5 |
| Application Gateway | a layer 7 load balancer, with an optional WAF | its match is an Application Load Balancer with AWS WAF; API Gateway fronts APIs | chapter 5 |
| service endpoint | a subnet setting that sends service traffic over the backbone | the URL of a service's API; the match is a VPC endpoint | chapter 5 |
| Availability Zone | traffic between zones is not charged | traffic between zones is billed, leaving one and arriving in the other | chapter 11 |
| private subnet | a subnet with default outbound access turned off | a subnet with no route to an internet gateway | chapter 12 |
| region pair | many regions have a pair, which geo-redundant storage copies to | no Region has a partner: you choose where to recover, and replicate | chapter 30 |
| Key Vault | keys, secrets and certificates in one vault | KMS for keys, Secrets Manager for secrets, ACM for certificates | chapter 27 |
| tag inheritance | subscription and resource group tags can flow to usage below | only account tags flow down; tag each resource | chapter 31 |
| App Configuration | settings and flags, read through client libraries | AppConfig: each change is a release, rolled back on an alarm | chapter 33 |
| runbook | a PowerShell or Python script | a Systems Manager document of steps, run across accounts and Regions | chapter 34 |
| Well-Architected | five pillars | six pillars, adding sustainability | chapter 35 |
End of the book