OPS 2. How do you define development standards and enforce them automatically?
Last updated on
Standards enforced by review comments decay in a predictable way. They hold while the person who cares is reviewing, they slip when that person is on holiday, and they are abandoned under a deadline. Nobody decides to drop them; the enforcement simply was never reliable.
Automated enforcement changes the failure mode. A rule in the pipeline either passes or blocks, it applies equally to everyone including the person who wrote it, and it does not get tired at the end of a long week.
Best practices
Section titled “Best practices”OPS 2.1Agree the standards that materially affect operability, and write them downOPS 2.2Enforce in the pipeline rather than in review commentsOPS 2.3Make the compliant path the fast pathOPS 2.4Version the standards and change them deliberately
OPS 2.1 Agree the standards that materially affect operability, and write them down
Section titled “OPS 2.1 Agree the standards that materially affect operability, and write them down”Risk if not established: Medium
Not every convention is worth enforcing. The ones that pay for themselves are those that affect whether the system can be operated, understood or recovered by someone who did not write it.
The categories that consistently matter:
- Structured logging and correlation. Without a shared format,
OPS 7cannot correlate anything. - Configuration handling. Where settings come from, how secrets are supplied, what happens on a missing value.
- Dependency policy. What may be added, from where, and how updates are handled.
- Health and readiness semantics. What each endpoint actually asserts, since
REL 10.1depends on the answer. - Error handling and timeouts. The defaults that
REL 5.1requires be explicit.
Formatting and naming conventions are worth automating precisely because they are not worth arguing about. Pick a formatter, apply it, and stop discussing it.
Distinguish the standards that block from those that advise. A rule that blocks a merge needs to be worth blocking for, and a long list of blocking rules of mixed importance trains people to look for ways around all of them.
On STACKIT. Standards are a team decision, not a platform feature.
Where the platform participates is in enforcement, which is OPS 2.2, and in the artefact side:
Container Registry
provides vulnerability scanning and access control for images, which turns a dependency policy
from a document into a check. That is also SEC 10.
Tradeoffs. Operational Excellence. Time to agree, which is mostly the cost of the discussion rather than the writing. Standards imposed without agreement are worked around rather than followed.
Verify. Where are your development standards written, when were they last changed, and which of them are enforced by something other than a person remembering?
OPS 2.2 Enforce in the pipeline rather than in review comments
Section titled “OPS 2.2 Enforce in the pipeline rather than in review comments”Risk if not established: Medium
A check that runs automatically is applied consistently. A check that depends on a reviewer is applied when that reviewer is available, attentive and willing to have the conversation.
Move everything mechanical into the pipeline: formatting, linting, type checking, test execution, dependency scanning, image scanning, infrastructure plan validation. Human review is then free for the thing it is uniquely good at, which is judgement about design.
Two properties decide whether the checks are respected. They must be fast, because a pipeline
that takes forty minutes teaches people to batch changes, which is the opposite of OPS 4.3. And
they must be trustworthy, because a check that fails intermittently gets re-run until it
passes, at which point it is decoration.
Run the same checks locally where possible, so that failure is discovered before the push rather than after. The pipeline should confirm what the developer already knows rather than being the first place a problem appears.
On STACKIT. STACKIT Pipelines is the native CI/CD for STACKIT Git , and it is compatible with GitHub Actions workflows, so existing workflows and community actions can be reused rather than rewritten. That materially lowers the cost of moving enforcement here.
Jobs run on runners in isolated environments, either STACKIT-managed or custom runners you operate. Managed runners are the lower-effort option; custom runners are what you need when a job requires access to a private network or specific hardware.
STACKIT runners are designed to provide reliable and scalable pipeline running.
Scalability
- Runners automatically scale depending on the number of queued jobs.
- Jobs are run as soon as compute capacity is available.
Security
- Each job runs in a fresh, ephemeral running environment.
- The environment is deleted immediately after the job completes.
- No data or credentials persist between runs.
Isolation
- Jobs are isolated from each other to prevent cross-job interference.
Reliability
- Runner infrastructure is distributed across multiple Availability Zones (AZs) to ensure high availability.
What is this?
This section is copied from the STACKIT docs automatically, several times a day. It cannot be changed here. Changes belong in the STACKIT docs.
Nothing obliges you to use it. Where a pipeline already exists elsewhere, the checks matter more
than where they run, and SOV 6 is the question to ask about a CI provider outside the boundary.
Tradeoffs. Cost Optimization. Pipeline minutes and runner capacity are real costs, and comprehensive checks on every commit add up. Scope by branch rather than by removing checks. Performance Efficiency. Slow pipelines cost developer time continuously.
Verify. List the checks that must pass before a change reaches production. How many are automated, how long does the pipeline take, and when did someone last merge past a failing check?
OPS 2.3 Make the compliant path the fast path
Section titled “OPS 2.3 Make the compliant path the fast path”Risk if not established: Medium
Standards that impose friction get circumvented, and the circumvention is invisible until an incident exposes it. This is the same dynamic as the security friction described in the Security tradeoffs, and it has the same answer: invest in the path rather than relying on discipline.
Practically, that means the correct way to do something should also be the easiest. A project template that already has the logging format, the health endpoints and the pipeline configured. A library that supplies the correlation identifier without anyone thinking about it. A generator that produces a compliant service skeleton in a minute.
The measure of success is whether anyone has to know the standard in order to comply with it. If compliance requires reading a document, some people will not.
Watch for standards that are expensive to satisfy and cheap to bypass. That combination reliably produces a codebase where the rule is followed in the parts written when there was time.
On STACKIT. Templates are where this becomes concrete on the platform side. Automation
Service
supports templates for recurring operations, and infrastructure modules under OPS 3 do the same
for resource definitions: a Terraform module that already encodes your standards makes the
compliant deployment the default one.
Tradeoffs. Cost Optimization. Building and maintaining templates and internal libraries is platform work with no direct feature output. It pays back per team that uses it, which means it is worth more in larger organizations and can be over-engineering in small ones.
Verify. How long does it take to create a new service that satisfies all your standards? If the answer is longer than an hour, how many recent services actually satisfy them?
OPS 2.4 Version the standards and change them deliberately
Section titled “OPS 2.4 Version the standards and change them deliberately”Risk if not established: Medium
Standards that change without a decision produce a codebase in several dialects, where the vintage of a component can be read from which conventions it follows. That is not itself a disaster, and it becomes one when nobody can say which convention is current.
Treat the standards as an artefact with the same handling as code: in version control, changed
through review, with the reasoning recorded. OPS 3 argues this for infrastructure and the same
argument applies here.
When a standard changes, decide explicitly what happens to what already exists. Three honest options: migrate everything, migrate on next touch, or leave existing code alone and apply the new rule only to new code. All three are legitimate. Not deciding produces the dialect problem.
Retire standards as well as adding them. A rule whose justification has expired continues to cost compliance effort and to occupy attention that a current rule needs.
On STACKIT. No platform feature applies. The standards document belongs in STACKIT Git alongside the code it governs, which is the same argument as keeping ADRs next to the system they describe.
Tradeoffs. Little. The cost is the discipline of treating a document as a versioned artefact, which is small compared with the cost of nobody knowing which version is current.
Verify. When did your standards last change, what was the reasoning, and what was decided about existing code? Can you find that record without asking someone?
Related
Section titled “Related”OPS 3Everything as code, which applies the same argument to infrastructureOPS 4Deployment automation, which the pipeline also carriesOPS 7Observability, whose correlation depends on a shared logging standardSEC 10Supply chain, which shares the dependency and image scanning checksREL 5Resilient interactions, whose timeout and retry defaults belong in the standards