Mainframe Lab 03 · Reference architecture
Observable Infrastructure Baseline
Metrics, logs, alerts, and a runbook as a shared operating layer for small infrastructure stacks.
GrafanaPrometheusLokiRunbooks
System map
Reference architecture flow
- 01SystemsServers, services, backups→
- 02SignalsMetrics, logs, availability→
- 03EvaluationThresholds, correlation, silences→
- 04ResponseAlert, owner, runbook→
- 05ImprovementReview, capacity, maintenance
Starting point
The problem this build examines
Servers and services work until they do not. Without a shared view, users report problems first, alerts become noisy, and ownership remains unclear.
Approach
How the system is structured
The reference architecture collects a small set of relevant signals, maps them to services and owners, and connects alerts to concrete runbook steps. Backups and recovery are tested as separate operating objectives.
Evidence
What this concept build demonstrates
How monitoring becomes an actionable operating routine—with ownership, escalation, and documented recovery.
Acceptance checks
- Every alert has an owner and a next step
- Maintenance cannot create an avoidable alert storm
- Backup status and recovery are tested separately
- Dashboards represent services, not only machines
Typical handover
- Signal and service inventory
- Dashboards and prioritised alert rules
- Runbooks for common incidents
- Maintenance and handover plan