OPS 7. How do you instrument the workload to answer questions you did not anticipate?
Zuletzt aktualisiert am
Monitoring tells you that a known condition occurred. Observability lets you answer a question you had not thought to ask in advance, which is the only kind of question the expensive incidents produce.
Whether a system is observable is decided in the code, when it is written. Nothing installed afterwards can recover a signal that was never emitted, which is why this is a design question rather than a tooling choice.
Best practices
Section titled “Best practices”OPS 7.1Emit metrics, logs and traces, and make them joinableOPS 7.2Propagate a correlation identifier through every hopOPS 7.3Instrument for the questions you cannot predictOPS 7.4Set retention by the value of each signal rather than uniformly
OPS 7.1 Emit metrics, logs and traces, and make them joinable
Section titled “OPS 7.1 Emit metrics, logs and traces, and make them joinable”Risk if not established: High
The three signal types answer different questions and are weakest alone.
Metrics are cheap, aggregated and good at “is something wrong”. They cannot tell you which request or which customer.
Logs carry detail and context. They are expensive at volume and useless without structure, because a human-readable message cannot be queried across a million lines.
Traces show the path of one request across services, which is the only practical way to find where latency actually goes in a distributed system.
Structure the logs. A message that reads well to a person and cannot be filtered by field is a
message you cannot use during an incident. This is one of the standards OPS 2.1 should enforce,
because a mixed-format log estate cannot be correlated at all.
The joining is the point. Three signal types that cannot be connected produce three partial
accounts of an incident and no narrative, which is what OPS 7.2 exists to prevent.
On STACKIT. Observability is the unified service for metrics, logs and traces with dashboards and alerting, and its architecture page describes the components. Grafana access is available for querying and dashboards.
For log-focused workloads there are two further options. Logs is label-based storage optimized for high-volume querying, with documented read and write limits worth checking against your expected volume before committing to it. LogMe is the centralized ingestion and search alternative.
For Kubernetes specifically, monitoring applications with Observability covers the integration.
The instrumentation itself is yours. The platform collects, stores and queries what your application emits, and no provider can emit a metric that says whether your checkout flow works.
Tradeoffs. Cost Optimization. Telemetry volume grows with the system and its bill is
itemized while its value is invisible until an incident, which is why retention gets cut in cost
reviews. OPS 7.4 is the answer. Performance Efficiency. Instrumentation costs some CPU and
latency, addressed by sampling rather than by removal.
Verify. Take a recent incident. Could you follow one affected request across every service it touched? What did that take, and what was missing?
OPS 7.2 Propagate a correlation identifier through every hop
Section titled “OPS 7.2 Propagate a correlation identifier through every hop”Risk if not established: High
Correlation cannot be added at query time. If a request identifier was not carried from the entry point through every downstream call, no amount of clever querying will reconstruct which log lines belong together.
Generate it at the edge if the caller did not supply one, carry it through every synchronous call, and include it in every log line and every trace span. The part that gets dropped is asynchronous work: messages on a queue, scheduled jobs triggered by a request, retries. Those need the identifier carried in the message rather than inherited from a call stack.
Carry a small amount of additional context with it where the classification permits: tenant, environment, version. Being able to filter an incident to one customer or one release is worth disproportionately more than the effort of adding the field.
Be careful what goes in. Log context is data, and identifiers that are also personal data bring
SEC 3 and SOV 3 with them into a store that usually has broader access than production.
On STACKIT. Propagation is application behaviour, usually supplied by a tracing library rather than written by hand. The platform side is collection and query, which Observability provides.
Making this automatic is the practical route. A shared library or service template that propagates
the identifier without anyone thinking about it is OPS 2.3 applied here, and it is the
difference between correlation that mostly works and correlation that has gaps in exactly the
services written under time pressure.
Tradeoffs. Performance Efficiency. A header on every call and a field on every log line, in practice negligible. Security. Correlation context propagates across trust boundaries, so what it contains needs to be a decision rather than whatever was convenient.
Verify. Pick a request identifier from a recent production log line. Which services can you find it in, and where does the chain break?
OPS 7.3 Instrument for the questions you cannot predict
Section titled “OPS 7.3 Instrument for the questions you cannot predict”Risk if not established: Medium
Dashboards built for anticipated failures are necessary and insufficient, because the incident that costs you is the one nobody anticipated. The property you want is the ability to slice by dimensions you did not think of in advance.
That comes from emitting context rather than pre-aggregating. A metric that counts errors tells you there are errors. The same metric with dimensions for endpoint, version, tenant and region lets you discover that they are confined to one customer on one release, which is usually the whole investigation.
Cardinality is the constraint and it is not optional. High-cardinality dimensions such as user identifiers multiply the stored series and can become the dominant cost. Put unbounded dimensions in logs and traces, which are queried rather than pre-aggregated, and keep metric dimensions bounded.
Instrument the business as well as the infrastructure. Checkouts completed per minute detects a class of failure that CPU and error rate will not, because a flow can be technically healthy and producing nothing.
On STACKIT. Observability stores and queries whatever dimensions you emit, and Logs documents its read and write limits, which is the practical bound on how much high-cardinality detail you can push at it.
Cardinality control is a design decision on your side rather than a platform setting, and it is worth making before the bill or the limit makes it for you.
Tradeoffs. Cost Optimization. Rich dimensions are the main driver of telemetry cost, and the discipline is to choose them rather than to emit everything and hope.
Verify. For your main error metric, which dimensions can you break it down by? Could you determine whether an error spike is confined to one tenant, one version or one zone?
OPS 7.4 Set retention by the value of each signal rather than uniformly
Section titled “OPS 7.4 Set retention by the value of each signal rather than uniformly”Risk if not established: Medium
Uniform retention is wrong in both directions at once. High-volume debugging data is kept far longer than anyone will look at it, while the signals needed for trend analysis or an audit are cut to the same short window during a cost review.
Tier it by what each signal is for:
- Debugging detail, high volume, useful for days. Short retention.
- Operational metrics for trends and capacity planning. Months.
- Audit and compliance records. Whatever
SOV 7requires, which is usually years and is not negotiable in a cost review.
Record the reason next to the number. Retention set to a period because a regulation or an investigation requirement demands it survives scrutiny; retention set to a period because it was the default does not, and it is the one that gets cut.
On STACKIT. Two retention facts are worth planning around.
The audit log records actions by users, service accounts and the platform at organization, folder and project scope, enabled by default, and is retained for 90 days in the STACKIT Portal. Where a longer period is required, records are exported through Telemetry Router to a destination you choose, and there is a worked example of connecting it to Logs. That export is the step that turns a 90-day operational record into compliance evidence, and it has to be set up before the period you will be asked about.
Application telemetry retention is set by the Observability service plan , and the platform already tiers it the way this best practice argues for. Plans also bound ingestion and storage.
STACKIT Observability verwendet offene Standards für die Erfassung von Telemetriedaten.
Alle Metriken, Logs und Traces müssen im OpenTelemetry- oder OpenMetrics-Format übertragen werden.
Metriken
- Standardaufbewahrung: 90 Tage
- Konfigurierbar bis zu 780 Tage (26 Monate)
- Die Aufbewahrung kann im Observability Service Dashboard oder über die Observability API geändert werden
Logs und Traces
- Standardaufbewahrung: 5 Tage
- Konfigurierbar bis zu 30 Tage
- Die Aufbewahrung kann über die Observability API angepasst werden
Was ist das?
Dieser Abschnitt wird mehrmals am Tag automatisch aus der STACKIT-Doku übernommen. Hier lässt er sich nicht ändern. Änderungen gehören in die STACKIT-Doku.
The 30-day ceiling on logs and traces is the number to plan around. It is ample for operations and
short for anything that has to answer a question months later, which is why SEC 11.1 and SOV 7
both route through an export rather than through retention settings.
The export destination is a placement decision under SOV 3, since telemetry inherits the
classification of what it describes.
Tradeoffs. Cost Optimization. This best practice exists to make the cost conversation
possible rather than to avoid it. Sovereignty & Compliance. Longer retention means more data
held for longer, which SUS 5 and data minimization under SEC 3 both push against.
Verify. What is the retention for each class of telemetry, and what is the stated reason for each figure? If you were asked about an administrative action from six months ago, could you answer?
Related
Section titled “Related”OPS 2.1Development standards, which should mandate the logging formatOPS 5.2Health gates, which query these signalsREL 10.1Health model, which turns these signals into a per-flow verdictSEC 11Detection and response, which needs the same records and a longer memorySOV 3Telemetry residency andSOV 7Auditability