Measure what users feel
A Service Level Indicator is a number that tracks the user’s experience — request success rate, latency of successful responses, freshness of data. Not CPU. Not “server up”. What the customer would notice.
ICAN Consultancy · Executive Field Guide
SRE as a Service — Platform Engineering & Reliability
The SRE playbook for platforms that can't afford to fail: SLOs and error budgets, observability, incident response, and the engineering discipline behind zero downtime — from the team that has kept Cardano infrastructure up since the network's foundation.
Cardano infrastructure operated by ICAN’s team since the network’s foundation, in prior roles and engagements.
“As reliable as possible” is not a target — it’s an unlimited budget with no owner. SRE replaces it with three connected instruments.
A Service Level Indicator is a number that tracks the user’s experience — request success rate, latency of successful responses, freshness of data. Not CPU. Not “server up”. What the customer would notice.
A Service Level Objective is the promise: e.g. 99.9% of requests succeed each month. It should be as reliable as the business needs — and no more. Every extra nine multiplies cost; users can’t tell the difference past a point.
100% minus the SLO is the error budget — unreliability you’re allowed. Budget healthy? Ship fast, take risk. Budget burned? Feature work pauses, reliability work takes over. The trade-off becomes a rule, not an argument.
A target you never miss is not ambition — it’s an SLO set too low. A target you always miss is fiction. Both waste money.
Senior SREs are scarce, expensive, and hard to keep challenged on a single estate. Buying reliability as a service can beat hiring — if you hold the provider to the same bar this guide sets.
| Demand | Why it separates real SRE from rebadged support |
|---|---|
| SLO-backed accountability | Explicit reliability targets in the engagement — uptime, latency, correctness — with performance reported against them. Effort-based contracts are support tickets wearing an SRE badge. |
| Design-build-operate, one team | Whoever operates the platform should be able to build it — and vice versa. Reliability bolted onto someone else’s architecture is always weaker than reliability engineered in from day one. |
| Engineers, not just operators | People who can write the script, the exporter, the migration tool — on the fly when needed. Ask directly: “when the runbook runs out, what do your people do?” |
| Everything as code, clean handover | The whole estate code-defined and reproducible, yours to take in-house whenever you choose. A provider who resists this is building lock-in, not infrastructure. |
| 24/7 that’s real | Follow-the-sun or genuine on-call rotations, symptom-based paging, incident command discipline, blameless postmortems you get to read. |
| Workload-native expertise | If you run blockchain infrastructure, demand node and validator operations experience — consensus, staking mechanics, protocol upgrades. Generic cloud ops learns those lessons on your stake. |
The outcome you’re buying is a number on a dashboard the board can read: the platform stayed up, and nobody on your team had to think about it.
SLOs and error budgets, the four golden signals, toil and automation, SREs who code, incident discipline and change management — the full 10-page guide, free to read and share.
You are welcome to share this guide in full, with attribution.
We architect for reliability from day one, build the estate as code — cloud, nodes, validators, RPC endpoints — and run it to defined targets with observability and 24/7 on-call.
Our engineers don’t stop at the runbook. They script fixes, tooling and migrations on the fly — then fold every improvisation back into the automated estate.
Node and validator operations, consensus behaviour, staking mechanics and protocol upgrades — expertise earned in production, on workloads where downtime is slashed stake.
Tell us about the platform you’re running — or the one you need built. We’ll come back with a pragmatic plan to design, build and operate it to clear reliability targets.