Skip to content

Mainframe Lab 03 · Reference architecture

Observable Infrastructure Baseline

Metrics, logs, alerts, and a runbook as a shared operating layer for small infrastructure stacks.

GrafanaPrometheusLokiRunbooks

System map

Reference architecture flow

Reference architecture flow
  1. 01SystemsServers, services, backups
  2. 02SignalsMetrics, logs, availability
  3. 03EvaluationThresholds, correlation, silences
  4. 04ResponseAlert, owner, runbook
  5. 05ImprovementReview, capacity, maintenance

Starting point

The problem this build examines

Servers and services work until they do not. Without a shared view, users report problems first, alerts become noisy, and ownership remains unclear.

Approach

How the system is structured

The reference architecture collects a small set of relevant signals, maps them to services and owners, and connects alerts to concrete runbook steps. Backups and recovery are tested as separate operating objectives.

Evidence

What this concept build demonstrates

How monitoring becomes an actionable operating routine—with ownership, escalation, and documented recovery.

Acceptance checks

  • Every alert has an owner and a next step
  • Maintenance cannot create an avoidable alert storm
  • Backup status and recovery are tested separately
  • Dashboards represent services, not only machines

Typical handover

  • Signal and service inventory
  • Dashboards and prioritised alert rules
  • Runbooks for common incidents
  • Maintenance and handover plan