The Junior Climb hardcover mockup
Multi-Team Blast Radius: Split Subscriptions, Not Hope

Multi-Team Blast Radius: Split Subscriptions, Not Hope

Two product teams share one subscription because finance wanted one invoice line. Both have Contributor. One deletes a resource group named temp that held the other’s private endpoints. The outage is total. The retrospective blames communication.

Friday cleanup. Good intentions. Wrong boundary. You already chose this failure in subscription design when you let “shared” mean “everyone.”

Blast radius is a design choice. Communication is a hope.

This is article seven of Series 2. The workload module gave teams a safe toolbox. This article decides how many teams share a toolbox before someone else’s delete becomes your incident.

Subscription is the hard wall

RBAC, policy exemptions, budgets, quotas, and Azure activity log attribution all get cleaner at subscription boundaries. If two teams can hurt each other with Contributor, they should not share a subscription unless one team is clearly tenant of the other. Platform hosting a shared cluster is a deliberate exception. It needs a named owner and a written contract.

Default: one application, or one tightly coupled product, per subscription per environment. Match what subscription design already argued. Series 2 only adds the team lens. When two squads share a backlog and a pager, one subscription can still make sense. When they share only a cost center code, they need walls.

CAF application landing zones assume this wall. Management groups apply policy per environment. Subscriptions apply ownership per product. When you collapse products into one subscription to save on “subscription sprawl,” you trade Azure limits for human failure modes. Sprawl with tags and budgets is cheaper than a combined incident channel and a shared blame thread.

I still walk into tenants proud of “only twelve subscriptions” where each one hosts six products and forty Contributors. That is not FinOps discipline. That is a blast-radius discount that comes due on Friday.

Resource groups are soft walls

Useful for organizing lifecycle and deleting a unit. Not enough if everyone is Contributor at subscription scope. Use resource-group-scoped roles when multiple squads must share a subscription temporarily. Make the temporary part visible with a review date.

Terraform works with a pipeline identity. Humans need Reader plus PIM for break-glass. Identity already said standing Owner is how Friday happens. Standing Contributor for two teams is how Friday becomes someone else’s outage.

Locks on critical resource groups help against accidents. They do not help against a pipeline identity with Contributor that intends to destroy. Design RBAC and pipeline scopes first. Locks second. Hope never.

Name resource groups after the product unit you are willing to delete together. A resource group named shared with three applications inside is a confession that you already know the wall is soft. Soft walls need soft permissions. Soft permissions need review dates.

When one spoke VNet hosts multiple apps

Sometimes network design puts multiple apps in one spoke for address space reasons. That can work when private endpoints and NSGs keep data planes separated, and when subscription walls still separate control planes. Network adjacency is not permission adjacency.

Do not confuse “same VNet” with “same team.” NSGs and ASGs reduce accidental east-west. They do not stop a Contributor from deleting the wrong resource group. If two apps share a VNet and a subscription, you have one failure domain wearing two product names.

Prefer one spoke VNet per subscription in the common case. Shared spokes need platform ownership of the network module and clear subnet allocation. App teams should not carve random subnets into a shared VNet from the portal.

Shared runtimes need a product owner

Shared AKS, shared App Service plans, shared SQL servers, and shared Service Bus namespaces are legitimate platform products when the economics are real. They are disasters when ownership is vague.

Name the owner. Name the blast radius. Name who gets Contributor where. Name the upgrade window. Name the namespace or database isolation model. Vague “we share compute to save money” language is how two products share an outage channel and neither has authority to fix the node pool.

For shared AKS, platform owns cluster lifecycle, node pools, ingress defaults, and cluster identity. Workloads own namespaces, deployments, and app-level secrets. Series 4 will go deeper. For this series, the rule is simple: if the platform team rebuilds node pools for every hotfix and app teams still open DNS tickets, you have an unspoken hybrid. Write the contract or split the clusters.

Shared App Service plans fail the same way. One noisy neighbor starves CPU. One team’s slot swap becomes everyone’s event. If you share, meter and alert per app. If you cannot meter, you cannot operate the share fairly.

Shared SQL logical servers with many databases look efficient until one team’s restore or firewall change takes everyone offline. Prefer separate servers when products have different change windows. Share only when you have a named DBA-style owner and a maintenance calendar app teams already agreed to.

Decision table for co-tenancy

SituationDefaultAllow when
Two products, two backlogsSeparate subscriptionsNever as cost-only consolidation
Two squads, one productOne subscription, RG rolesPipeline identity per squad if needed
Shared AKS for many appsPlatform-owned cluster productContract, namespaces, clear on-call
Shared ACRShared-services subscriptionRBAC per repo, no subscription Contributor
Shared Key Vault for many appsRefusePlatform vault for platform secrets only
Same spoke VNet, many appsPrefer one app per sub + VNetAddress pressure with platform-owned shared spoke
Finance wants one invoiceTags + management group + showbackSeparate subscriptions still, roll up in reporting
“Temporary” shared sub for two teamsRefuse without review dateDated exception with migration plan

Publish the table with the spoke contract. Architecture reviews go faster when co-tenancy answers are already written.

Policy exemptions and shared pain

A policy exemption at subscription scope helps every resource in that subscription. If two products share the subscription, one team’s exemption becomes the other’s posture. Prefer exemptions at resource or resource group scope with expiry. Prefer fixing the module so the exemption dies.

Policy only works when the exception path is slower and louder than the default path. Shared subscriptions make exceptions louder in the wrong direction: one click, many products affected.

Budgets behave the same way. One budget on a shared subscription pages a shared Owner list that nobody treats as personal. Per-product subscriptions make showback and alerting honest. Logging attribution with the four tags still matters inside a subscription, yet tags do not stop delete.

Networking blast radius

Spoke-to-spoke peering couples failure domains in routing. Shared firewall rule collections that allow broad east-west couple them in policy. Networking already argued for intentional east-west. Blast radius adds the team question: if product A’s compromise can route freely to product B’s data plane, your segmentation was a slide.

Private endpoints keep data planes addressable without mesh trust. Prefer them. Hub hairpin with explicit rules when you must. Broad “allow spoke CIDRs to spoke CIDRs” rules recreate a flat network with more billing lines.

Cleanup automation and the delete story

The opener was not hypothetical. Cleanup scripts, orphan hunters, and well-meaning “delete unused resource groups” jobs are high blast-radius tools. Run them only against scopes one product owns. Require dry-run output and a named Owner approval when the target subscription hosts more than one Application tag.

I keep finding automation identities with Contributor on a shared subscription because “orphans were expensive.” They were. So were the private endpoints that automation deleted at 2 a.m. FinOps Series 5 will teach safer orphan hunting. The blast-radius rule starts here: reclaim tools inherit the same walls as humans.

Delete protection for private endpoint resource groups, Key Vaults, and production data stores should be part of the co-tenancy review, not an afterthought after the outage channel fills.

Temporary exceptions that stay temporary

Leadership will ask for a temporary shared subscription. Sometimes the answer is yes for a migration window. Write the exception with four fields: products involved, review date, migration owner, and exit criteria that include separate subscriptions or a named platform-owned shared runtime.

Without those four fields, temporary means permanent. Permanent shared Contributor means you already accepted the delete story from the opener. Put the exception in the same repo as the spoke contract so the next architecture review does not rediscover it as folklore.

Deployment Model keeps pipelines from editing the wrong scope. Blast radius keeps people and products from sharing the wrong scope. A temporary exception that grants both teams Contributor on one subscription fails both rules at once.

Operational signals that your walls are wrong

Incident channels that always include two product teams for one resource group delete. Change freezes that stop unrelated apps because they share a plan or a cluster node pool. Onboarding that requires Contributor on a subscription already full of strangers. Budget alerts ignored because “that is the shared sub.”

Track count of subscriptions with more than one Application tag value in active use. Track Contributor assignments at subscription scope per human. Track shared runtime incidents that impacted multiple apps. Those numbers tell you whether walls are real.

Deployment Model keeps pipelines from editing the wrong scope. Blast radius keeps people and products from sharing the wrong scope. You need both. Perfect pipelines into a shared Owner subscription still produce shared outages.

When leadership asks to consolidate subscriptions for cleanliness, bring the delete story from the opener. Cleanliness that increases Friday risk is not governance. Reporting can roll up. RBAC should not.

Showback can answer the invoice-line question without collapsing RBAC. Cost Management views, management group rollups, and tags already exist so finance can see one number while engineering keeps many walls. If finance tools cannot roll up today, fix the reporting path. Do not flatten the control plane to make a CSV prettier.

The next article closes Series 2 by treating the platform as a product. Catalog items, honest SLAs, and metrics that prove the default path no longer requires a hero on every ticket. Series 3 picks up security and compliance on top of that product.

Randy Bordeaux

Azure Trainer | Veteran | Engineer

Helping IT professionals grow through cloud, leadership, and shared knowledge.


Discover more from Randy Bordeaux

Subscribe to get the latest posts sent to your email.

Discover more from Randy Bordeaux

Subscribe now to keep reading and get access to the full archive.

Continue reading