Home / How it works
How it worksFrom "it's broken" to "here's why it won't happen again."
No discovery workshops, no statement-of-work theater. Four steps, each with a concrete deliverable, starting within one business hour of your message.
Triage
What happens. You describe the problem in the form, two sentences is enough. A senior engineer (not a coordinator) replies within one business hour with clarifying questions and an honest first read: what's likely wrong, how urgent it really is, what fixing it involves, and what it costs.
What you get. A diagnosis direction and a flat-scope quote. If it's outside what we do well, we tell you that here, for free, and point you somewhere better.
Fix
What happens. We get access, stabilize first, stop the bleeding, protect the data, then fix the actual fault. Work happens in your systems: your repos, your cloud, your review process. You get progress updates in plain language at agreed intervals.
What you get. A working system. For code fixes: pull requests into your repo. For infrastructure and database fixes: changes logged, explained, and reversible where possible.
Root-cause report
What happens. Every engagement ends with a short written report, typically two to four pages, not forty. What broke, why it broke, what we changed, what we ruled out, and what to watch.
What you get. A document a founder can read and an engineer can verify. It's also your insurance: if anyone touches this system later, the incident's full story exists in writing.
Prevent
What happens. The report ends with a prioritized prevention list, usually two to five specific changes (a missing alert, an untested backup, a query that will fall over at 10x data). Ranked by risk, sized by effort.
What you get. A clear path to "this class of failure is closed." Execute it with your own team, or move to a reliability retainer and we handle it, plus monitoring and a guaranteed SLA.
How access works.
- Least privilege, always. We start with read access, logs, metrics, dashboards, a read replica. Write or admin access is requested explicitly, scoped to the fix, and revoked when we're done.
- Your accounts, not ours. Access via your IAM roles, temporary credentials, or a screen share with your engineer driving. We don't ask you to hand over root and hope.
- Everything is logged. Commands run and changes made are recorded and included in the report.
- NDA before anything sensitive. We'll sign yours, or provide ours, before you share a single credential. Standard practice, not a special request.
- No team? No problem. If there's no engineer on your side, we work from whatever access you can grant, hosting panel, cloud console invite, and document as we go.
What you walk away with.
- The fix, live in production, verified.
- The root-cause report (what broke, why, what changed, what to watch).
- The prevention list, prioritized by risk.
- Logged, reviewable changes, in your repos and infrastructure, owned by you.
- A team that now knows your stack, if you ever need us again.
How the work actually happens.
We have no technical staff. Can this process still work?
Yes, it is a situation we handle regularly. You grant whatever access you can (a hosting panel login, a cloud console invite) and we take it from there, documenting everything as we go. The root-cause report is written so a founder can read it, so you will understand what happened without needing an engineer to translate.
How do you get into our systems without creating a new risk?
Least privilege, explicitly granted, fully logged. We start with read access such as logs, metrics, or a read replica, and request write access only when needed, scoped to the fix and revoked after. Everything flows through your IAM roles or temporary credentials, never shared root passwords, and we sign an NDA before you share anything sensitive.
How long does a typical fix take?
It depends on the failure, which is why triage gives you a specific timeline along with the quote. As a pattern: most emergencies are stabilized the same day we engage, with the full fix and report landing within days, not weeks. Stabilize first, then fix properly, is the order we always work in.
What happens if the problem turns out to be bigger than the quote?
The quote is the ceiling for the scope we agreed, full stop. If diagnosis reveals a genuinely different problem than triage indicated, we stop, show you what we found, and re-quote before doing anything more. You never discover a scope change on the invoice, and you keep every finding regardless.
The first step takes two sentences.
Tell us what's happening. Triage is how every engagement starts, and where we tell you honestly if you even need us.
Book a fix