Home / Resources
ResourcesWhat we know, written down.
Practical guides from the engineers who handle these failures for a living. No gated PDFs, no fluff, the same advice we'd give you on a triage call, free. If you're mid-incident, each guide tells you what to do first and when DIY stops being safe.
The cost of downtime: statistics, and how to calculate your own number
The widely cited figures on what outages cost and how long recovery takes, presented with honest caveats, plus a worked framework to estimate what an hour of downtime costs your business specifically.
Read the guide → ComparisonEmergency fix vs reliability retainer: an honest self-selection guide
Erzon's two offers compared honestly, with the real numbers, so you can work out which one actually fits your incident load.
Read the guide → Tutorial · SoftwareHow to find a memory leak in production
RSS keeps climbing and the OOM killer is restarting your pods. How to prove it's a real leak, capture the evidence, and trace it to the code that's holding on.
Read the guide → Tutorial · DevOpsHow to roll back a bad deploy safely
The rollback decision, the database migration trap that makes rollbacks fail, and the checklist to run before and after you pull the lever.
Read the guide → Tutorial · DatabaseHow to test your database backups
Nobody wants backups, everybody wants restores. How to run a restore drill, prove the data is intact, and automate the test so it never goes stale.
Read the guide → ComparisonIn-house hire vs on-demand engineering: what coverage actually costs
A full-time SRE, an on-demand incident partner, or a generalist agency: what each costs, how fast each covers you, and when each is the right call.
Read the guide → Template · DevOpsIncident response runbook template
A copy-paste runbook for production incidents: roles, severity levels, comms, escalation, and the step sequence, with a filled example.
Read the guide → Tutorial · DatabaseHow to run a zero-downtime database migration
Schema changes on big tables can lock writes for minutes or hours. The expand, migrate, contract pattern and the online DDL tools that make changes safe.
Read the guide → Comparison · DatabaseManaged vs self-hosted databases: what you're actually buying
RDS, Aurora, Cloud SQL, and Atlas remove real work, but not the work that causes most database incidents. Here is the honest breakdown.
Read the guide → GuideThe complete guide to production incident response
End to end incident handling for a business: what to prepare, how to triage and mitigate, what to tell customers, and how each incident makes the next one rarer.
Read the guide → TemplateProduction readiness checklist
The checklist to run before trusting a system in production, grouped by database, infrastructure, and application, with pass criteria that are actually testable.
Read the guide → TemplateRoot cause analysis report template
A copy-paste blameless RCA template: summary, impact, timeline, root cause, contributing factors, and action items, with a short filled example.
Read the guide → Guide · DevOpsThe 7 most common DevOps failures (and how to fix them)
The failures that account for most infrastructure incidents, expired certs, wedged deploys, red pipelines, runaway bills, with the fix and the prevention for each.
Read the guide → Guide · SoftwareDebugging a production bug: a business owner's guide
You don't need to read code to run a good bug hunt. How to manage a production bug like an incident: what to ask, what to preserve, and how to tell real progress from motion.
Read the guide → Tutorial · DatabaseHow to recover a corrupted production database
The first hour decides how much data you keep. What to do immediately (and what never to do), how to assess the damage, and the recovery paths from least to most invasive.
Read the guide →About these guides.
Do I need to sign up or hand over my email to read these?
No. No gates, no PDFs held hostage, no newsletter ambush. These guides are the same advice we give on triage calls, written down and free, because an informed reader mid-incident makes better first moves, and better first moves save data.
Can I suggest a topic you have not covered?
Please do; send it through the contact form or to [email protected]. The guides come from the failures we see most often, so if you are wrestling with something we have not written up, that is a useful signal. If it is urgent rather than curious, say so and it becomes a triage conversation instead.
Are these written by actual engineers or by a content team?
By the engineers who handle these failures for a living, which is why they contain specific commands, specific pitfalls, and specific moments where DIY stops being safe. If a guide ever reads like it was written by someone who has never watched a restore fail at 3 a.m., tell us and we will fix that too.
Past the point where a guide helps?
That's what we're for. Describe the situation and a senior engineer replies within one business hour.
Book a fix