True

The Incident Was Due at 4 PM

We Have Seen It Before. At 1 PM.

The Incident Was Due at 4 PM. We Should Have Seen It at 1 PM.

A critical healthcare or pharmaceutical platform doesn't usually fail at the moment the business notices the failure. The failure tends to start much earlier.

A data pipeline stops processing correctly at 13:00. A scheduled workload starts piling up retries. An integration keeps returning HTTP 200 but quietly stops delivering complete records. A storage queue begins to grow. A certificate creeps toward expiration. A downstream service gets slower and slower.

Then at 16:00, the operational process that depends on that platform misses its expected output, and that's when the business reports the incident.

But 16:00 wasn't when the problem began. It was just the moment the problem became visible to the business.

So the engineering question should really be: what signal existed at 13:00 that could have told us the 16:00 failure was already on its way? That one question changes how observability should be designed.


Availability is not the same as operational health

Plenty of enterprise environments monitor whether systems are technically available. CPU is under threshold, memory is stable, containers are running, the database responds, the API returns 200, the Kubernetes cluster looks healthy. Every dashboard is green.

And the business process is already failing.

That happens because infrastructure monitoring answers "is the technology running?", while critical platforms need a second question answered too: "is the technology still producing the outcome the business expects?" Those two are not the same thing.

What the dashboard says      What might actually be happening
Service is available      It's processing increasingly stale data
Pipeline is running      Throughput has dropped below what the next deadline needs
Integration succeeds      Mandatory fields are being dropped
Platform is online      Latency is slowly eating the margin before a critical process fails

This is where observability has to go beyond infrastructure metrics.


Start with the failure the business cannot accept

The first mistake is usually starting with tools. Prometheus, Grafana, Datadog, CloudWatch, OpenTelemetry, Elastic, Splunk. They all matter, but none of them is the starting point.

Start with the business event instead. For example: at 16:00, a validated downstream process must receive the complete dataset produced during the previous processing window.

Then work backwards. What has to be true at 15:30? At 15:00? At 14:00? And what signal at 13:00 would tell us that meeting the 16:00 requirement is becoming unlikely?

That gives you a completely different monitoring model. Instead of only watching component health, you start watching the distance between what the platform is doing now and a future business failure.


Step 1: Define the critical service outcome

For every critical workflow, define what "healthy" means from the consumer's point of view. Here's how the shift looks in practice:

Instead of      Define it as
Pipeline is running      99.9% of expected records are processed and available downstream within 15 minutes of arrival
API is available      99.95% of requests required by the critical workflow complete successfully within the latency budget
Database is online      Required data is complete, current and queryable within the operational window

These become the basis for meaningful Service Level Indicators.


Step 2: Measure freshness, completeness and throughput

Healthcare, pharmaceutical and MedTech platforms lean heavily on data movement, which makes three signals especially important.

Freshness. How old is the latest successfully processed data? If normal lag is five minutes and it climbs to twenty, the system may still be "up", but operational risk is going up with it.

Completeness. Did we receive and process everything we expected? A workflow that successfully processes 98,000 records can still have a serious problem if 100,000 were expected.

Throughput. Is processing capacity enough to clear the workload before the deadline? Say the queue holds 60,000 records, the system handles 5,000 an hour, and the downstream deadline is six hours away. The infrastructure might look perfectly healthy, but mathematically the process can't finish in time. That's an incident already in progress.


Step 3: Watch the error budget before it disappears

Traditional alerting usually reacts once a fixed threshold gets crossed: CPU above 90%, disk above 85%, latency above 500 ms. That's useful, but it isn't enough.

For critical services, teams should understand how fast reliability is being consumed. If a platform has an availability or processing objective, every failure eats into the tolerance available for that period. That's the logic behind SLOs and error budgets.

So the important question becomes: are we burning through our reliability margin faster than expected? A rapid burn rate is often far more useful than waiting for the final SLA to be breached. By the time the SLA is missed, the alert is just describing history.


Step 4: Build predictive operational signals

The best alerts don't just say "something is broken." They say "if this continues, something important will break."

Here are some examples of what that looks like:

Signal What the alert says
Queue growth Processing rate is lower than arrival rate. If sustained, queue capacity runs out in 2.3 hours
Deadline risk Remaining workload divided by current throughput points to completion at 16:42 against a 16:00 deadline
Data freshness Expected update interval is five minutes, current freshness is 18 minutes and rising
Retry acceleration Retry volume is up 400% over baseline in the last 20 minutes
Dependency degradation A downstream service is still available, but p95 latency has doubled across three consecutive windows

These are operational signals. They tell engineers where the system is heading, not just where it is right now.


Step 5: Connect technical signals to business impact

An alert like "Kafka consumer lag > 50,000" means something to the platform engineer. It might mean almost nothing to the incident commander, the application owner, or a business stakeholder.

A better alert adds context: consumer lag indicates roughly 75 minutes of processing delay, and if current throughput continues, the 16:00 downstream processing window will be missed.

Now the engineering signal has operational meaning, and that makes prioritisation much easier. Not every infrastructure anomaly deserves the same response. A 90% CPU alert on an autoscaling stateless service may need no action at all. A slowly growing processing delay on a regulatory or operationally critical workflow may need immediate attention even when infrastructure utilisation is low.


Step 6: Define ownership before the incident

Observability without ownership just becomes another dashboard nobody acts on. Every critical signal should have clear answers to a handful of questions:

  • Who owns it?
  • Who receives the alert?
  • Who can diagnose it?
  • Who can decide to intervene?
  • Which service owns the failing dependency?
  • When does it escalate?
  • What happens outside business hours?
  • What information does the responder get automatically?

If those questions get answered during the incident, the response is already too slow.


Step 7: Automate the first safe response

Not every incident needs a human to take the first corrective action. Some failure modes are predictable enough to automate safely, for example:

  • restarting failed workers
  • increasing processing capacity
  • pausing a problematic deployment
  • rerouting traffic
  • activating fallback infrastructure
  • clearing known transient states
  • scaling queue consumers
  • rolling back a release after defined health conditions fail

The goal isn't "automate everything." It's removing the repetitive operational actions that slow recovery down, while keeping sensible controls around critical systems.


Step 8: Test the detection path

A monitoring setup that has never been tested is really just a hypothesis. Teams should deliberately check that:

  • the right metric actually changes
  • the alert actually fires
  • the right person actually receives it
  • the alert carries enough diagnostic context
  • the runbook is usable
  • escalation works
  • recovery restores the expected service outcome

This matters most after changes to architecture, infrastructure or deployment. Operational readiness belongs to delivery, not to something bolted on after production.


The metric that matters is detection margin

Most organisations track MTTD, Mean Time to Detect. That's useful. But for time-critical workflows there's another concept worth adding:

                Detection Margin = Expected Business Failure Time − Engineering Detection Time
              

If business impact hits at 16:00 and engineering spots the trajectory at 15:55, the margin is five minutes, which is probably useless. If engineering spots it at 13:00, the margin is three hours, and now there's real time to step in.

That changes what observability is for. The goal isn't only to detect failures quickly. It's to detect the conditions that will cause failures while there's still enough time to prevent them.


A practical review for critical platforms

Pick one critical workflow. Don't start with the monitoring stack, start with the business deadline or outcome, and ask:

  • What exactly must happen, and by when?
  • Which technical conditions must stay true for that to happen?
  • Which of those can degrade silently?
  • Which signal changes first?
  • How much intervention time would that signal give us?
  • Who owns the response?
  • Can the first corrective action be automated safely?

This exercise often exposes more operational risk than reviewing yet another infrastructure dashboard.


Where m-tech1 fits

This kind of problem sits across several disciplines at once. Architecture determines where failure can propagate. Platform Engineering shapes how teams consume infrastructure and operational capabilities. DevOps decides how changes reach production. SRE brings the reliability model, the SLOs and the operational response. Cloud engineering provides telemetry, automation and resilient infrastructure. Observability ties all of those layers together.

m-tech1 can work with enterprise engineering teams to identify critical workflows, map technical dependencies, define meaningful service indicators, improve observability, and build the automation needed to catch operational risk earlier.

The objective is simple: the business should never be the monitoring system.


Detect Problems Before They Become Incidents

To install this Web App in your iPhone/iPad press and then Add to Home Screen.