The Junior Climb hardcover mockup
Platform as Product: The Ticket That Should Never Return

Platform as Product: The Ticket That Should Never Return

Platform engineering fails quietly when every capability is still a human workflow. You have modules. You have vending. You still have a queue full of “please create,” “please peer,” “please exempt.”

Friday intake: twelve tickets. Nine of them are App Service plans in subscriptions you already vendored. Two are DNS zones you know you should have cataloged last quarter. One is a real firewall change. Your on-call spends the day on the nine. The real change waits until Monday. You already paid for that queue shape in Deployment Model when the portal stayed easier than the pipeline.

A platform is a product when the catalog is real, the SLAs are honest, and the default path does not require heroics.

This is article eight of Series 2, the close of Spoke / Workload Patterns. Series 1 built the landing zone. This series made spokes operable for app teams. Series 3 picks up security and compliance evidence on top of this product.

Catalog items, not tribal knowledge

Publish a short catalog. Each item has a name, owner, how to request (pipeline or form that opens a PR), SLA, and what it returns.

Minimum viable catalog:

  • Vend subscription and spoke (platform pipeline)
  • Add private DNS zone support for a PaaS type
  • Firewall rule pattern with canary
  • Workload starter repo for web, data, and integration
  • Budget and action group attachment if not included in vend
  • Break-glass path: documented, not whispered
  • Address space expansion request (project timeline, not a two-day SLA)
  • Policy allow-list change for an approved SKU
  • Shared ACR repository onboarding
  • Workload module version upgrade guidance

If it is not in the catalog, the answer is “not yet” or “workload builds it with the module.” The answer is never “sure, I will click it after lunch.” Clicks after lunch become undocumented architecture.

CAF and Azure Landing Zone give you reference implementations. The catalog is how your organization makes those references operable. Management groups, policy, hub-spoke, private endpoints, and identity seams are platform product surfaces. They are not a workshop slide that expired after the pilot.

SLAs that survive contact with Friday

An SLA you cannot meet is a rumor. Publish times you can defend with staffing you actually have.

Examples that work in real estates: two business days for a firewall rule that already matches a pattern, after canary. Five business days for a new private DNS zone type with links to pilot spokes. Same day for subscription vending when the form is complete and address space exists. Weeks for address space expansion that requires redesign.

Say what is excluded. Emergency break-glass is not the standard path. Standard path is the pipeline. Identity break-glass should create a ticket and a follow-up PR, not a quiet portal change that ages into truth.

I still walk into platforms that advertise “24-hour networking SLA” with two engineers and three hundred spokes. That SLA teaches app teams to expect miracles and platform engineers to burn out. Honest slower SLAs beat fictional fast ones.

Publish SLA breach handling the same way you publish the SLA. If a firewall pattern request misses two business days, who gets paged, and what gets deferred to keep the breach from cascading? An SLA without an operating response is a customer promise you already broke in silence.

Metrics that prove the product

Track monthly:

  • Percentage of platform intake that is catalog work versus ad hoc clicks
  • Median time from vend request to usable nonprod subscription
  • Count of portal-created resources cleaned up by policy or the next apply
  • Workload module adoption and version distribution
  • Failed private endpoint onboardings attributable to DNS versus RBAC versus firewall
  • Policy exemption count with expired dates
  • Percentage of app teams that shipped without a platform engineer on the ticket

If more than half of intake is “create my App Service,” the spoke contract and workload module are invisible. If intake is “add privatelink.azurecr.io” and “raise the prod SKU allow-list,” the product is working.

Logging and policy feed these metrics when you bother to query them. Gut feel is how platforms regress into hero culture.

Self-service defaults

The best platform ticket is the one that never opens. Vending outputs feed the workload module. Starter repos already know Pattern A, B, and C. DNS zones for common PaaS exist before the first ask. Deny policies in prod stop public IPs without a conversation. Documentation lives next to the module in git.

Self-service fails when the secure path is longer than the portal path. Make the module shorter. Make the portal harder with policy. Networking already showed what happens when DNS is a human workflow. Productize the seam.

Subscription design productized the boundary. This series productized what lands inside it. Self-service without walls is chaos. Walls without self-service are a ticket factory. You need both.

Intake taxonomy that feeds the roadmap

Tag every platform ticket at close: catalog item, missing product, education (team did not read the contract), or true exception. That taxonomy is how you stop arguing from anecdotes.

Missing product tickets become roadmap candidates. Education tickets become README fixes, starter repo updates, or office hours topics. True exceptions get an owner and an expiry. Catalog items that still require a human for every request need automation investment.

I keep finding platforms that measure only ticket volume. Volume without taxonomy rewards heroes who close the same class of request forever. Product teams measure which features remove future tickets. Platform should too.

Platform roadmap visibility

App teams plan quarters. Platform teams that only react to tickets will always be late. Publish a short roadmap: which PaaS DNS zones come next, which module versions deprecate when, when shared ACR lands, when AKS becomes a platform product if that is the direction.

Roadmap items need owners and exit criteria. “Improve networking” is not a roadmap item. “Add privatelink zones for Azure Container Apps and link pattern for all nonprod spokes” is a roadmap item.

Intake should influence the roadmap. Repeated tickets for the same missing zone are a product signal. Heroically fulfilling each ticket forever is how you avoid building the catalog item.

Share the roadmap in the same place as the catalog. App teams should not need a meeting with the one platform engineer who remembers next quarter’s DNS work. Visibility reduces shadow IT more than another denial in Slack.

On-call and operations as product surfaces

A platform product that cannot be operated is a future outage. Hub firewall, private DNS, vending pipelines, and shared ACR need runbooks, dashboards, and an on-call that knows which pages are real.

Separate app SLO pages from platform path pages. App Service 500s page the app Owner. Hub firewall health pages platform. If the same three people get every page, you do not have products. You have a hero team with phones.

Change management for hub components needs canary spokes and maintenance windows you publish. Silent hub changes on Friday afternoon are how you earn a reputation for breaking the company. Product teams schedule releases. Platform should too.

Track pages that woke a human and resolved as “app misconfiguration” versus “platform defect.” A high app-misconfiguration rate means docs and starter repos are failing. A high platform-defect rate means the product is underbuilt. Both are actionable. Neither is fixed by longer on-call shifts.

Cost and showback belong in the product story

Budgets and the four tags are part of vending. Shared services need their own cost owners. Platform workspace ingestion needs attribution. When FinOps asks why costs jumped, the catalog should already explain which product surfaces drive spend: private endpoints, firewall, logging, shared ACR.

Platform as product includes saying no to free unlimited diagnostics “for compliance” without a cost Owner. Compliance without showback becomes a surprise bill. Surprise bills become shadow IT.

App team journey as acceptance criteria

Walk the journey once a quarter with a real team that is not on your friends list. Start at “we need a nonprod subscription” and end at “first green apply with private endpoints and diagnostics.” Time each step. Note every place a human had to interpret folklore.

If that journey still requires a platform engineer to paste subnet IDs into a Slack thread, the product failed acceptance. Fix the outputs, the starter, or the catalog item. Do not celebrate the CAF workshop completion while the journey still depends on heroes.

The same journey in prod should differ only by SKUs, policy effect, and approval gates, not by inventing a new network design. Environment promotion without snowflakes was article five. Product acceptance is proving that story with a stopwatch.

How Series 2 fits the operating model

Article one wrote the spoke contract. Article two gave three reference spokes. Article three shared services without shared subscriptions. Article four limited connectivity choices. Article five promoted environments without snowflakes. Article six defined the workload module. Article seven set blast-radius walls. This article demands that all of it behave like a product with a catalog, SLAs, and metrics.

Microsoft will keep shipping ALZ reference implementations. Your job is to turn reference into something the next app team can consume on a Tuesday without finding the one engineer who remembers DNS.

I keep finding platforms that completed the CAF workshop and never published a catalog. The workshop was the beginning. The product is the operating system of the cloud estate.

Series 3 moves to Security and Compliance: threat modeling the control plane, Defender for Cloud as platform signal, continuous evidence, secrets at scale, network posture beyond the firewall slide, secure SDLC for IaC, audit season runbooks, and platform incident response. The spoke product you built here is what those controls will measure. If the product is still a ticket queue, compliance will become another queue with sharper language.

Randy Bordeaux

Azure Trainer | Veteran | Engineer

Helping IT professionals grow through cloud, leadership, and shared knowledge.


Discover more from Randy Bordeaux

Subscribe to get the latest posts sent to your email.

Discover more from Randy Bordeaux

Subscribe now to keep reading and get access to the full archive.

Continue reading