- 01Users, outcomes, and service boundaries
- User groups and critical tasks, business processes, products and services, entry points, environments, regions, dependencies, service objectives, demand, latency, correctness, availability, freshness, durability, security, privacy, support expectations, risk tiers, failure consequence, and owners.
- 02System and change context
- Applications, jobs, queues, databases, infrastructure, third parties, identities, networks, data flows, resource names, versions, releases, feature exposure, configuration, migrations, capacity, scaling, deployment records, incidents, known failure modes, recovery paths, and current architecture and responsibility maps.
- 03Current telemetry and tools
- Instrumentation, agents, collectors, exporters, traces, metrics, logs, events, profiles, propagation formats, schemas, dimensions, sampling, aggregation, parsing, redaction, storage, indexes, dashboards, queries, alerts, synthetic checks, status communication, runbooks, access, retention, cost, and vendor constraints.
- 04Response and evidence use
- On-call and support coverage, escalation, paging channels, incident roles, automated actions, tickets, silence and suppression rules, alert history, false and missed detection, diagnosis paths, restore and repair decisions, post-incident reviews, capacity and release decisions, audit needs, training, maintenance, and ownership.