Security reviews love CVE lists and container scans. The landing zone usually fails somewhere quieter: a standing Owner on the management group, a pipeline identity that can edit the hub, a policy exemption with no expiry.
Threat modeling Azure starts with who can change the platform. Not with which npm package is outdated in a spoke.
This is article one of Series 3. Series 2 gave app teams a spoke contract. Series 1 already framed identity as the undersized control plane. This series is whether you can defend that plane when someone asks for proof.

Assets that matter first
Write the list before you invite AppSec to the workshop. If the whiteboard starts at “VM hardening,” you already lost an hour.
- Entra ID tenant settings and privileged directory roles
- Management group hierarchy and Azure Policy assignments
- Platform CI identities (OIDC federated credentials on the service principal or managed identity)
- Hub network, Azure Firewall policy, and the UDRs that force traffic there
- Private DNS zones the whole estate resolves through
- Central Log Analytics workspace and any Microsoft Sentinel workspace that sits on it
- Break-glass accounts and the Conditional Access exceptions that keep them usable
Compromise any of those and workload CVEs become a secondary problem. A stolen Contributor on one spoke hurts that app. A stolen Owner at the platform management group hurts every subscription that inherits from it.
I still walk into estates where the threat model is a spreadsheet of CVE IDs from the last scanner run. The management group still has three standing Owners. Nobody can say which Entra role can reset MFA for a Global Admin. The scanner is not wrong. It is measuring the wrong layer.
Threats worth writing down
Skip the 40-page STRIDE novel nobody maintains. One page per platform asset is enough if it names the threat, the detection, and the response owner.
| Threat | Example | Mitigation you already preached |
|---|---|---|
| Privilege accumulation | Standing Owner, broad Contributor at root MG | Entra PIM, dual pipeline identities, access reviews |
| Pipeline poison | Compromised user merges a PR that edits hub firewall | Code owners, required reviews, OIDC instead of long-lived secrets |
| Policy hollowed out | Permanent exemptions, Audit that never becomes Deny | Expiry on exemptions, canary MG, then Deny |
| Logging blind | Workspace in a sandbox sub, diagnostics off on hub | DINE for diagnostics, locks, central workspace from logging |
| Hub bypass | Spoke-to-spoke peering, UDR removed, public IP on a NIC | Spoke contract, Deny public IP, peer only through platform module |
| Identity federation drift | Federated credential subject widened to * or wrong repo | Inventory federated credentials quarterly, alert on create |
The table is a living document in the platform repo. Security partners who still start at the VM can read it and stop asking you to “scan the landing zone” as if it were a single host.
Entra and PIM are the first chapter
Control-plane threat modeling without Entra is cosplay. Azure RBAC inherits from management groups, but the people who can change Entra roles, Conditional Access, and app registrations sit above that.
Map these questions to a person and a log:
- Who can activate Global Administrator, Privileged Role Administrator, or User Access Administrator?
- Are those activations through PIM with justification, MFA, and a ticket field you actually check?
- How long do eligible assignments sit before access review removes them?
- Who can create app registrations and add client secrets?
- Where do break-glass accounts live, and which Conditional Access policies intentionally skip them?
Identity already said standing Owner at the management group is how Friday subscriptions happen. Threat modeling adds detection: AzureActivity for Microsoft.Authorization/roleAssignments/write at the MG scope, Sign-in logs for privileged users, PIM audit for activations that do not match a change ticket.
If detection is “someone notices in Teams,” you do not have detection. You have hope.
Pipelines are principals
Treat the platform OIDC identity like a highly privileged human who never sleeps. It can create subscriptions, assign policy, and rewrite firewall rules. A compromised developer account that can merge to main on the platform repo is a path to that identity.
Threats that belong on the pipeline page:
- Long-lived client secrets in variable groups
- A single identity that applies both hub and workload state
- Admin bypass on branch protection “so we can hotfix”
- External modules pinned to
ref=main - Local
terraform applyfrom a laptop that still holds Owner
Deployment Model already split platform and workload identities. The threat model records what happens when that split collapses: workload engineer needs a hotfix, platform apply is locked, Owner at the MG comes back for one change. That one change is the incident you will explain later.
Policy and exemptions are attack surface
Policy is how Deny becomes real. An exemption is how Deny becomes optional. Attackers do not need to disable your initiative if a helpful engineer already exempted the resource group “because the vendor needs public storage.”
For every exemption resource, the threat model expects:
- Control ID from MCSB or CIS (or your customer framework)
- Owner and ticket link
- Expiry date that lands on someone’s calendar
- Effect you will restore (Audit vs Deny) when it expires
Permanent exemptions without owners are standing privilege with different branding. Put them in Terraform so the next auditor can read git instead of interviewing three people.
Data plane still matters. Second.
Storage public access, SQL firewall rules allowing 0.0.0.0/0, key-based auth on storage when the standard is Microsoft Entra: those are workload threats. Policy and Defender catch many of them. They sit on top of a control plane that is not already owned by everyone.
Order the workshop that way on purpose. Spend the first hour on Entra, MG RBAC, pipelines, policy assignments, and the hub. Spend the second hour on PaaS defaults the spoke module should already deny. If you reverse the order, you will leave with a container-scanning action item and the same standing Owners you walked in with.
Detection has to name a table
For each asset, write the log source you will query when something goes wrong.
| Asset | Primary signal | Where it lives |
|---|---|---|
| Role assignments at MG/sub | AzureActivity | Platform Log Analytics |
| Privileged sign-ins | SigninLogs / AADNonInteractiveUserSignInLogs | Entra diagnostics to same workspace or Sentinel |
| PIM activations | AuditLogs (PIM) | Entra diagnostic settings |
| Policy assignment changes | AzureActivity on Microsoft.Authorization/policy* | Platform workspace |
| Firewall policy edits | AzureActivity + AZFW* tables | Hub diagnostics you already required |
| Key Vault admin ops | KeyVault logs | Vault diagnostics |
| Pipeline applies | CI logs + git SHA on the change ticket | Repo + Azure DevOps/GitHub |
Logging already fought for a workspace that survives cleanup. Threat modeling is why that fight mattered. Without retained Activity and Sign-in data, your “response” section is fiction.
Output of the exercise
Ship three artifacts from the first workshop. Not a PDF deck that dies in SharePoint.
- Asset register in the platform repo: asset, owner, blast radius, detection query link, response role.
- Top ten misuse cases in plain language (“stolen platform OIDC identity”, “exemption without expiry”, “UDR removed on prod spoke”).
- Quarterly review date on the platform calendar with the on-call rotation, not a one-time tabletop that never repeats.
Link saved Kusto searches from the logging companion. If the query does not run today against real data, delete it from the register. Empty detection rows train people to ignore the document.
Security partners should leave knowing which Microsoft Cloud Security Benchmark controls map to these assets. You do not need to recite every MCSB control ID in the workshop. You need the map that article three will turn into continuous evidence.
Break-glass without mythology
Every estate claims break-glass accounts. Few can answer how they are monitored.
Write the threat model entries for them:
- Two cloud-only accounts, long passwords in a sealed process, not tied to a daily-driver laptop identity
- Excluded from some Conditional Access policies so you can recover when CA is wrong, and included in every monitoring path that still works when CA is wrong
- Sign-in alerts that page immediately
- Use only through the IR procedure in article eight, with a ticket opened before or during the use
- Password rotation after every use, not “when we remember”
Break-glass that nobody has tested is a fictional control. Test sign-in in a controlled window twice a year. Record who watched. If the test fails because MFA devices or CA exclusions drifted, fix that before audit season finds it for you.
Shared responsibility with AppSec
Invite AppSec and cloud security once the asset register exists. Their job is not to rewrite the list into a container-scanning program. Their job is to challenge missing detections and to map MCSB language onto your artifacts.
Useful outcomes from that joint hour:
- Agreement on which MCSB network and identity controls are platform-owned
- A short list of workload threats Policy will constrain (public storage, open management ports)
- Named owners for Defender alert triage vs recommendation backlog (article two)
Useless outcomes: a new 60-page STRIDE document, a demand to “scan all landing zone resources” without defining the control plane, or a meeting that ends with “we should revisit this next quarter” and no git commit.
How you know it worked
Two weeks later, check tickets and access.
- New Owner assignments at the MG require PIM and a ticket, or they do not happen
- Platform PRs that touch firewall or policy still have two reviewers who are not the author
- At least one exemption has an expiry in the next 90 days that someone owns
- On-call can name the workspace and the Activity query for role assignment writes without opening Confluence
- Break-glass test date is on the calendar with a named facilitator
If those fail, the threat model was a meeting. Meetings do not stop standing privilege.
I still walk into workshops where everyone agrees the management group is the crown jewel and then leaves standing Owner in place because “the pipeline is not ready.” The threat model is ready when the privilege changes, not when the slide is pretty.
Next: Defender for Cloud as Platform Signal. Treat recommendations as a backlog owned by platform and workload, not as a Secure Score trophy for the board.
Randy Bordeaux
Azure Trainer | Veteran | Engineer
Helping IT professionals grow through cloud, leadership, and shared knowledge.


