Prod is a fork of nonprod that someone copied in the portal last year. Different subnet names. Different tags. A Key Vault SKU that only exists in prod because “security tightened it” by hand. The pipeline promotes code. The infrastructure promotes myths.
Friday before a release: nonprod is green. Prod apply fails on a public network access property that never existed in the nonprod module. Someone “fixed” prod in the portal six months ago. Nobody captured it. You already paid for portal drift in Deployment Model.
Promotion in a landing zone means the same module, different parameters, different management group policy pressure. It does not mean a second architecture discovered during cutover.
This is article five of Series 2. Reference spokes and connectivity gave you shapes. Promotion is how those shapes survive contact with prod.
One module. Three parameter files.
sandbox.tfvars, nonprod.tfvars, prod.tfvars, or YAML equivalents in the workload pipeline. Same source. Same remote module version pin. SKUs change. Replica counts change. Public network access flips from soft to Deny-enforced. Address space comes from the platform pool per environment. You do not reuse prod CIDRs in sandbox and pray.
If prod requires a hand-built resource that is not in the module, you did not promote. You snowflaked. Add the resource to the module or reject the exception. Portal heroes feel fast until the next engineer runs plan and gets a destroy proposal for the “temporary” private endpoint someone loved.
CAF environment separation usually maps to management groups: Sandbox, Non-prod, Prod. Subscription design already put subscriptions under those groups. Promotion inherits that pressure. Prod gets stricter policy. Nonprod gets Audit where you are still learning. Sandbox gets a budget and an expiry story.
I still walk into repos where prod/ is a copy-paste of nonprod/ with twenty unrelated diffs. Half of those diffs are accidents. The other half are tribal security decisions that never made it into variables. Pin one module. Express differences as variables. Review the variable file in the PR. Do not review a second tree of Terraform that drifted because someone was in a hurry.
What is allowed to differ
SKU and scale. Retention and diagnostic verbosity. Whether certain optional components exist, such as APIM in prod only. Policy effects already assigned at the management group, for example prod Deny versus nonprod Audit. Region pairs when you have a real DR story. Secret values and Key Vault references, never secret shapes that only exist in one environment.
What must stay the same: subnet naming schemes, tag keys, private DNS ownership model, hub peering pattern, identity model for the pipeline, and forced tunneling through the hub. Those differences are architecture forks. Architecture forks fail during incidents because runbooks assume one shape.
If a security control only exists in prod, nonprod never proved it. That is how you discover missing private endpoints during the release window. Shift controls left as far as policy and cost allow. Soften SKUs. Keep the control plane model identical.
ALZ expects application landing zones under environment management groups for a reason. The same spoke module under different MG policy pressure is the operable version of “prod is stricter.” Hand-built prod-only private endpoints and portal-only Deny exemptions are the inoperable version.
Gates that earn the word promote
Code review on the module and the tfvars. Plan output reviewed for destroy and replace surprises. Policy compliance check against the target management group. Connectivity checklist from the previous article: DNS resolves private, routes point at the firewall, public access matches intent. Secret scan on the PR. Required reviewers for prod that include someone who understands the spoke contract.
Gates should fail closed for prod. Soft warnings that everyone ignores are documentation with CI green checkmarks. If the gate is too noisy, fix the gate. Do not train teams to click Continue.
Identity matters at promotion time. The prod pipeline identity should be scoped to the prod subscription. It should not be the same standing Owner used for nonprod experiments. PIM for humans. Pipeline identities with least privilege. Break-glass documented and rare.
OIDC federation should differ by environment subject. A federated credential that can deploy to sandbox and prod from any branch is how a feature branch becomes a production change. Scope subjects to the repo, branch or environment protection rules, and the Azure role that matches that environment only.
Data promotion is separate from infrastructure promotion. Restoring a prod database backup into nonprod without scrubbing is a compliance incident wearing a DevOps badge. Infrastructure tfvars promote shapes. Data pipelines promote datasets with their own controls. Confusing those two is how PII shows up in a sandbox that every contractor can Reader.
Data versus infrastructure
Infrastructure promotion answers: will the resources exist with the right network, identity, and policy posture? Data promotion answers: which records move, who approved it, and what was redacted?
Keep them in different pipelines when you can. At minimum, keep them in different stages with different approvers. A green Terraform apply is not permission to copy prod customer tables into nonprod for “realistic testing.” Synthetic data and scrubbed subsets are platform-adjacent products many teams undersize. Undersizing them creates shadow copies on laptops.
Storage account public access, SQL firewall rules, and Cosmos keys are common snowflake sites. Nonprod left open for a migration tool. Prod tightened by hand. The module should express both states as variables so plan shows the flip. Hand edits in prod without a PR are how the next apply becomes a production incident.
Key Vault contents differ by environment. Key Vault shape should not. Same RBAC model, same purge protection expectation in prod, same private endpoint pattern, same diagnostic settings. A prod-only vault with access policies while nonprod uses RBAC is an identity fork. Identity forks fail audits and fail rotations.
Drift is a first-class enemy
Portal changes in prod that never return to git will eventually fight the pipeline. Decide what wins. In a landing zone that takes Deployment Model seriously, the pipeline wins. Portal breaks glass, then someone files the PR before the change ages past the shift.
Detect drift with periodic plan against prod, Azure Policy, and Defender recommendations that show public exposure. Logging helps when diagnostic settings themselves drift. Missing diagnostics in prod only is a promotion failure even when the app still serves traffic.
I keep finding prod resource groups with locks and nonprod without, plus a tribal story about “prod is special.” Locks can be variables too. Special snowflakes that only live in one environment are how onboarding docs lie.
Policy exemptions that exist only in prod are drift wearing a governance badge. If nonprod never needed the exemption, ask whether the control should be in the module as a variable, or whether the exemption should expire when the migration ends. Standing exemptions without expiry become the architecture.
Version pins and promotion of modules
Pin the workload module version in each environment’s pipeline. Promote module versions the same way you promote app versions: nonprod first, then prod. Floating main for prod modules is how a Friday platform commit becomes a Saturday outage for every consumer.
Platform module changes that affect spokes need canary subscriptions. One nonprod spoke first. Then a wave. Then prod. Hub changes need even more caution. Promotion discipline applies to platform products, not only app teams.
Changelog and upgrade notes are part of the module product. If consumers cannot tell whether a minor bump renames a subnet, they will pin forever and you will support fossils.
Azure Verified Modules and internal wrappers follow the same rule. Pin the version. Read the changelog for security default changes: public network access flags, RBAC role assignments, firewall rules. A module default that reintroduces public blob access will create Defender noise across every new spoke the week you bump.
Policy pressure that should tighten toward prod
Map effects deliberately:
| Control | Sandbox | Nonprod | Prod |
|---|---|---|---|
| Public IP on NIC | Audit or Deny | Deny preferred | Deny |
| Storage public network access | Audit | Audit then Deny | Deny |
| Key Vault purge protection | Audit | Deny if cost allows | Deny |
| Required tags | Deny | Deny | Deny |
| Allowed locations | Deny | Deny | Deny |
Audit-then-Deny is still the right path when the module is not ready. Skipping Audit and slamming Deny only in prod is how release Fridays invent architecture. Put the Deny assignment at the management group. Let tfvars express the resource properties that comply. Do not “fix” prod in the portal while nonprod stays noncompliant.
Defender for Cloud recommendations that only appear in prod usually mean nonprod never matched the intended posture. Treat that as a promotion bug. Remediate nonprod first so the next release proves the control.
How you know promotion works
Releases stop inventing infrastructure in the war room. Prod plan looks like a parameter delta, not a rewrite. New engineers can read three tfvars files and explain environment differences in ten minutes. Exception list for snowflake resources shrinks quarter over quarter.
Track mean time from nonprod green to prod green for infrastructure changes. Track count of portal-only prod resources discovered by plan. Track policy exemption count in prod with expiry dates in the past. Those metrics tell you whether promotion is a process or a story you tell in onboarding.
When teams insist prod must differ “because compliance,” ask which control is missing from the module. Compliance controls that only exist as portal folklore fail audits and fail releases. Put the control in code. Let management group policy enforce it. Let tfvars express the allowed differences.
The next article names the workload module app teams may run. Yes lists, never lists, and versioning so Contributor in a spoke cannot quietly edit the hub.
Randy Bordeaux
Azure Trainer | Veteran | Engineer
Helping IT professionals grow through cloud, leadership, and shared knowledge.


