The Junior Climb hardcover mockup
The Workload Module App Teams May Run

The Workload Module App Teams May Run

Platform publishes a module. App teams wrap it. Somewhere in the middle, someone adds azurerm_firewall_policy_rule_collection_group to the workload repo because the portal wizard made it look easy.

Friday PR. Green plan. Merge. The workload identity suddenly manages hub firewall rules. You already paid for that boundary in Deployment Model. Two pipelines. Two identities. One of them just learned to edit the other team’s job.

That is how you lose the landing zone without a migration project. The diagrams stay pretty. The RBAC story is already compromised.

This is article six of Series 2. Deployment Model split the pipelines. This article names what is allowed inside the workload one.

Workload module: yes

Resources inside the vendored subscription that implement the app:

  • Resource groups for the app with required tags (Environment, CostCenter, Owner, Application)
  • App Service, Functions, Container Apps, and AKS workloads in their namespace when the cluster is already a platform product, or when you explicitly vended per-team clusters
  • Azure SQL, Storage, Cosmos, Service Bus, Event Hubs as the reference pattern needs
  • Private endpoints for those resources in the spoke PE subnet
  • Key Vault for app secrets with purge protection in prod
  • Role assignments in this subscription to the workload identities that need them
  • Diagnostic settings to the workspace resource ID the platform publishes
  • Metric alerts for app SLOs and action groups that page the app Owner
  • Autoscale settings, app settings sourced from Key Vault references, and private DNS records only when the platform pattern allows workload-written records
  • App Configuration for product settings when that is the chosen pattern
  • Budgets at subscription or resource group scope that still page Owner

The module should take the spoke VNet ID, PE subnet ID, DNS zone IDs or rely on already-linked zones, Log Analytics workspace ID, and tags as inputs. It does not create a second VNet. It does not create a second hub peer. It does not invent address space.

CAF calls this the application landing zone content. Platform already placed the subscription, spoke, and guardrails. The workload module fills the subscription. If the module starts creating management groups or hub peers, it stopped being a workload module.

Workload module: never

Anything that changes platform seams or other teams’ blast radius:

  • Hub firewall policy rule collection groups
  • Hub route tables, VPN or ExpressRoute gateways, Bastion in the connectivity subscription
  • Management group policy assignments or exemptions
  • Private DNS zones in the hub created as duplicates in the spoke
  • Peering to other spokes
  • Public IPs in prod when policy Denies them
  • Role assignments above the subscription, including Owner on the management group
  • Central Log Analytics workspace resource itself, as opposed to diagnostic settings that send to it
  • Platform Key Vaults and shared ACR resource groups
  • Subscription vending, management group moves, and budget action groups that page platform on-call for app SKU mistakes

If an app team needs a hub firewall rule, they open a platform catalog request. The platform pipeline applies it after canary. The workload plan must not grow a firewall resource because someone pasted a sample from the internet.

I still walk into workload repos that “temporarily” manage DNS zones, “just until platform catches up.” Temporary ownership of platform DNS is how you get split-brain resolution and a week of onboarding failures. Catch up means a platform PR. Catch up does not mean Contributor on the hub resource group.

Interfaces beat folklore

Publish module inputs with types, defaults, and examples for Patterns A, B, and C. Publish required outputs: resource IDs the app pipeline needs, private endpoint IDs, Key Vault URI, principal IDs for role assignments. Vending should already output spoke network IDs. The workload module should consume them without scraping the portal.

Document breaking changes. Subnet rename is a breaking change. Tag key rename is a breaking change. Quietly changing a default from public network access enabled to disabled is a breaking change that should show up in release notes, even when it is the right security move.

Policy will Deny some module outputs in prod that still pass in sandbox. That is expected. The module should express environment differences as variables so plan predicts the Deny before merge, not after a red pipeline.

A good interface answers three questions without a meeting: what do I pass in, what do I get back, and what will break if I upgrade. Paste sample terraform.tfvars for each reference pattern next to the README.

Versioning the module like a product

Semantic versioning with a changelog. Consumers pin versions in their pipelines. Platform tests modules against canary spokes before tagging a release that every team will pull on Monday.

Floating to main in prod is how a Friday refactor becomes a Saturday restore conversation. Pin. Promote. Same discipline the previous article described for environment promotion, and the same pipeline honesty Deployment Model already required.

Major versions for intentional breaks. Minor for additive resources with safe defaults. Patch for bug fixes. Deprecation windows measured in weeks, not “we deleted the variable yesterday.”

Support policy matters. How many major versions will platform support? How do teams get help? Slack folklore is not support. A short runbook and office hours beat heroics.

I keep finding ten forks of the “official” module with one-line diffs for SKU defaults. That means the module interface is wrong or the docs are invisible. Fix the module. Do not bless forks as self-service.

Pipeline identity and plan review

The workload pipeline identity is Contributor in the target subscription, or a tighter custom role if you have matured that far. It is Reader nowhere that lets it manage hub resources. It cannot edit management group policy. If the plan includes resources outside the subscription ID you vended, fail the pipeline.

Required reviewers for changes that touch networking variables, public network access, and role assignments. Those three areas recreate landing zone failures faster than app settings typos.

Identity already covered PIM for humans. The module story covers machines. Machines need least privilege too. A pipeline with Owner “so apply always works” will eventually apply something you regret.

Ban lists in CI should fail on resource type prefixes that belong to the hub: firewall policy rule groups, VPN gateways, ExpressRoute circuits, Bastion hosts, management group policy assignments. A plan summary that lists azurerm_* types in a comment on the PR is cheap insurance. Reviewers cannot catch what they never see.

Testing before you publish

Platform should maintain a smoke test spoke that applies the workload module for each reference pattern in nonprod. Private endpoint DNS checks. Policy compliance. Basic connectivity through the firewall. If the smoke test is red, do not tag a release.

App teams should have a sandbox path that uses the same module with cheaper SKUs. Sandbox is where they learn. Prod is where they pin.

Logging should be on by default in the module. Optional diagnostics that default off recreate blind spots. Make the secure and observable path the short path.

Smoke tests should assert the never list as well as the yes list. After apply, query Resource Graph for forbidden types. If a firewall policy rule collection group appears, the module failed even if the app URL returned 200.

Starter repos and golden paths

A module without a starter repo still generates tickets. Publish starter repositories for web, data, and integration that already wire module inputs from vending outputs. Include sample pipelines, sample alerts, and the connectivity checklist as a job.

Golden paths beat golden docs. Docs that say “pass the PE subnet ID” lose to a starter that already does it. When teams copy the wrong sample from a random blog, you get firewall resources in workload state. Your sample should be easier than the internet.

Networking and subscription design set the walls. The workload module respects those walls in code. Code is the contract that survives reorgs.

Refresh starters when required inputs change. A stale cookiecutter that still asks for a second VNet ID recreates the anti-pattern you spent a year eliminating. Treat starter drift as a release blocker.

State files and what they prove

Where the workload state lives tells you who owns the blast radius. State for app resources belongs with the workload pipeline, in a storage account the platform hardened: private endpoint, soft delete, versioning, and access limited to the pipeline identity. Platform engineers should not need Contributor on every app state file to “help.” Helping that way recreates the human router.

If an app team’s state file contains hub firewall IDs as managed resources, the module boundary already failed. Remote state inspection during PR review catches that faster than a post-incident blame thread. Require a plan summary that lists resource types. Ban lists belong in CI, not in a wiki.

Backend configuration itself should come from a platform-approved pattern: naming, key prefixes per environment, and encryption. Teams inventing their own state accounts in the portal recreate the snowflake problem the previous article already closed.

Optional features without optional security

Modules grow toggles: enable_apim, enable_cosmos, enable_private_endpoints. Defaults matter. Security-sensitive toggles should default to the posture you will defend in prod. Cheap SKUs can default soft in sandbox. Public network access should not default on in a module you market as landing-zone ready.

Every optional feature needs a tested path in the smoke spoke. Untested toggles become production discoveries. If you cannot staff tests for a feature, do not ship the toggle. A smaller module that works beats a kitchen-sink module that surprises Deny policies at release time.

Document the dependency graph. Enabling Cosmos implies private endpoints, DNS, diagnostics, and identity roles. A boolean that only creates the account leaves teams half-landed and filing tickets for the other half.

Destroy, drift, and partial applies

Document what destroy is allowed to touch. Workload destroy should never peer into the hub. Soft-delete Key Vaults need runbook steps so a bad destroy does not become permanent data loss.

Partial applies happen. Someone cancels a pipeline mid-run or uses -target because “the rest was fine.” Require a full plan before the next merge. Drift detection in nonprod should catch portal clicks before they become the real state file. Policy Deny closes public network access left “temporarily” on. The module should still make the correct apply the path of least resistance.

How you know the module is working

Platform intake shifts from “create my App Service” to “add this PaaS type to the module” and “add this DNS zone to the catalog.” Fork count drops. Time from vend to first green nonprod apply drops. Portal-created resources in workload subscriptions drop.

Track module adoption version distribution. A long tail of ancient majors means upgrades are painful or communication failed. Track failed applies caused by policy Deny. High counts mean the module and policy are out of sync. Fix the defaults.

Track how often workload plans include banned resource types. A rising count means samples on the internet are winning, or your starter repo is stale. Refresh the golden path until copying your repo is easier than inventing firewall Terraform at 4 p.m.

Office hours help when they feed the backlog. If the same question appears three times, it becomes a doc fix, a default change, or a catalog item. Answering the same question forever is how platform engineers become FAQ bots with on-call rotations.

The next article covers multi-team blast radius. Subscription walls, soft resource group walls, and shared runtimes that need a named owner. Contributor shared across products is how cleanup tickets become outages.

Randy Bordeaux

Azure Trainer | Veteran | Engineer

Helping IT professionals grow through cloud, leadership, and shared knowledge.


Discover more from Randy Bordeaux

Subscribe to get the latest posts sent to your email.

Discover more from Randy Bordeaux

Subscribe now to keep reading and get access to the full archive.

Continue reading