Writing, tool comparisons, and explainers on monitoring, observability, and keeping production healthy. Engineer to engineer, no fluff.
Notes on monitoring, alerts, investigation, and on-call for small teams.
Latest: What is a good response time? p95 and p99 latency explained
Honest comparisons of the tools small teams reach for, and what they cost.
Latest: Datadog alternatives and competitors: the small-team guide