Architecture decisions, post-mortems, and pricing essays from the team behind ZnowPulse. New posts every couple of weeks.
A compact monitoring review to find stale targets, misleading health checks, broken alert routes, and missing ownership.
Read postMake your status page useful by deciding who updates it, what evidence they need, and what happens when nobody is immediately available.
Read postA reliable status page needs a different failure path from the systems it reports on. Here is the practical architecture to aim for and the trade-offs to document.
Read postA hosted status page has real operational costs. Here is what it needs to do during an incident, where pricing differs, and how to decide whether managed hosting is worth it.
Read postSeparate customer communication, alert suppression, and health evidence when planning maintenance for a small service.
Read postA calm checklist for investigating disagreements between your browser and an external uptime monitor.
Read postTranslate uptime percentages into minutes, then check the measurement window, sampling method, and exclusions behind the number.
Read postA practical post-deployment checklist covering public routing, read-only smoke tests, asynchronous work, and monitoring evidence.
Read postWhy a healthy container does not prove a reachable application, and how to separate restart checks from customer-facing monitoring.
Read postClear, adaptable incident updates for investigation, mitigation, monitoring, and resolution—without invented ETAs.
Read postReduce alert fatigue by separating symptoms, confirmation, routing, and recovery instead of muting everything.
Read postUnderstand polling intervals, missed short outages, and confirmation delays before choosing faster uptime checks.
Read postA successful HTTP response does not prove your application works. Learn which checks catch maintenance screens, broken dependencies, and empty responses.
Read postStart with customer-facing failures, not a wall of green dots. A practical monitoring plan for a small SaaS team.
Read post