Home / Resources

Resources

What we know, written down.

Practical guides from the engineers who handle these failures for a living. No gated PDFs, no fluff, the same advice we'd give you on a triage call, free. If you're mid-incident, each guide tells you what to do first and when DIY stops being safe.

Statistics

The cost of downtime: statistics, and how to calculate your own number

The widely cited figures on what outages cost and how long recovery takes, presented with honest caveats, plus a worked framework to estimate what an hour of downtime costs your business specifically.

Read the guide →
Comparison

Emergency fix vs reliability retainer: an honest self-selection guide

Erzon's two offers compared honestly, with the real numbers, so you can work out which one actually fits your incident load.

Read the guide →
Tutorial · Software

How to find a memory leak in production

RSS keeps climbing and the OOM killer is restarting your pods. How to prove it's a real leak, capture the evidence, and trace it to the code that's holding on.

Read the guide →
Tutorial · DevOps

How to roll back a bad deploy safely

The rollback decision, the database migration trap that makes rollbacks fail, and the checklist to run before and after you pull the lever.

Read the guide →
Tutorial · Database

How to test your database backups

Nobody wants backups, everybody wants restores. How to run a restore drill, prove the data is intact, and automate the test so it never goes stale.

Read the guide →
Comparison

In-house hire vs on-demand engineering: what coverage actually costs

A full-time SRE, an on-demand incident partner, or a generalist agency: what each costs, how fast each covers you, and when each is the right call.

Read the guide →
Template · DevOps

Incident response runbook template

A copy-paste runbook for production incidents: roles, severity levels, comms, escalation, and the step sequence, with a filled example.

Read the guide →
Tutorial · Database

How to run a zero-downtime database migration

Schema changes on big tables can lock writes for minutes or hours. The expand, migrate, contract pattern and the online DDL tools that make changes safe.

Read the guide →
Comparison · Database

Managed vs self-hosted databases: what you're actually buying

RDS, Aurora, Cloud SQL, and Atlas remove real work, but not the work that causes most database incidents. Here is the honest breakdown.

Read the guide →
Guide

The complete guide to production incident response

End to end incident handling for a business: what to prepare, how to triage and mitigate, what to tell customers, and how each incident makes the next one rarer.

Read the guide →
Template

Production readiness checklist

The checklist to run before trusting a system in production, grouped by database, infrastructure, and application, with pass criteria that are actually testable.

Read the guide →
Template

Root cause analysis report template

A copy-paste blameless RCA template: summary, impact, timeline, root cause, contributing factors, and action items, with a short filled example.

Read the guide →
Guide · DevOps

The 7 most common DevOps failures (and how to fix them)

The failures that account for most infrastructure incidents, expired certs, wedged deploys, red pipelines, runaway bills, with the fix and the prevention for each.

Read the guide →
Guide · Software

Debugging a production bug: a business owner's guide

You don't need to read code to run a good bug hunt. How to manage a production bug like an incident: what to ask, what to preserve, and how to tell real progress from motion.

Read the guide →
Tutorial · Database

How to recover a corrupted production database

The first hour decides how much data you keep. What to do immediately (and what never to do), how to assess the damage, and the recovery paths from least to most invasive.

Read the guide →
FAQ

About these guides.

Do I need to sign up or hand over my email to read these?

No. No gates, no PDFs held hostage, no newsletter ambush. These guides are the same advice we give on triage calls, written down and free, because an informed reader mid-incident makes better first moves, and better first moves save data.

Can I suggest a topic you have not covered?

Please do; send it through the contact form or to [email protected]. The guides come from the failures we see most often, so if you are wrestling with something we have not written up, that is a useful signal. If it is urgent rather than curious, say so and it becomes a triage conversation instead.

Are these written by actual engineers or by a content team?

By the engineers who handle these failures for a living, which is why they contain specific commands, specific pitfalls, and specific moments where DIY stops being safe. If a guide ever reads like it was written by someone who has never watched a restore fail at 3 a.m., tell us and we will fix that too.

Past the point where a guide helps?

That's what we're for. Describe the situation and a senior engineer replies within one business hour.

Book a fix