Fortnightly · For the executive committee
Your organisation is putting probabilistic systems into paths where failure has a physical, clinical, financial or political consequence. This is a briefing on what that changes — written for the people who will answer for it.
Why this exists
Every issue takes one real failure mechanism, explains it precisely enough that an engineer would recognise it, and translates it into the language an executive committee actually decides in — exposure, obligation, and the cost of continuing as you are.
Not “observability matters.” A mechanism — why a threshold alert cannot detect a feed that silently stopped, why a retry policy becomes an amplifier under load, why an AI investigator never reports uncertainty.
Quantified where it can be, and honestly labelled where it is modelled. Regulatory penalty, downtime, engineer-hours, missed windows, and the second-order costs that never reach the incident report.
Autonomy compresses diagnosis and removes toil. It also introduces confidently wrong conclusions, log-borne prompt injection, unbounded spend and explainability obligations. Both halves, in every issue.
The pattern first, the product only where it is the honest answer. Where the right move is organisational or architectural rather than a purchase, the issue says so.
The cost of not tackling it
Each figure is drawn from an issue in the archive, where the reasoning and assumptions behind it are published in full. We do not print a number without its method.
Latest issue
Commercial SaaS observability cannot follow classified data across the boundary. The result is that the estates with the highest consequence of failure frequently have the least visibility into it.
Issue 12 · Government, Defense & Aerospace · 10 min
The organisations with the most at stake are frequently the least instrumented, and the reason is architectural rather than cultural.
Coverage
Each has a characteristic way of failing, a different regulator, and a different reason that generic advice does not apply. One issue per sector, with the application landscape, the exposure, and the mitigation.
Downtime is clinical. Diagnosis must be defensible before it is fast.
Supervised resilience, and a trace that stops at the mainframe boundary.
A year of engineering graded across five days that cannot move.
Selling reliability to everyone else, with an SLA credit attached.
No buffer, a federal enforcer, and silence that looks like health.
The clearest per-minute cost in the economy, and no degraded mode.
Failures cascade through a scheduled network for days.
Statutory availability on one side, an unrepeatable moment on the other.
Small estates, and transactions larger than the entire IT budget.
Classification creates an observability gap that procurement cannot close.
What every issue contains
So a returning reader can go straight to the part they came for — and so the argument cannot skip the inconvenient section.
The event, filing or change that makes this the subject this fortnight.
How the failure works at the level of the system. The section that earns the reader’s trust.
Where the exposure actually lands — regulatory, operational, commercial, reputational.
What autonomy changes, both the upside and the new failure modes it introduces.
What to change, ordered by cost, cheapest first.
Hosted or customized — how the tooling is actually deployed in that sector.
The engagements that move an organisation from assessment to operating practice.
The strongest case against that issue’s advice, made honestly. When it does not apply, and what it costs.
Tools and a partner
Thalamus Advisory builds enterprise resilience practice, supported by two products that can be spun up as managed instances or deployed inside your own boundary.
OTLP-native ingest for traces, metrics and logs. Golden signals and Apdex per service, critical-path attribution, baseline-relative log regression, and correlated problems that name one root-cause entity instead of paging nine teams.
An agent fleet over that telemetry. Seven-stage investigation with an adversarial verifier on a separate model, guardrails enforced outside the prompt, and postmortems generated as evidence rather than reconstructed afterwards.
Both products as managed instances, pooled or dedicated, provisioned per region — or self-hosted inside your VPC, data centre or accredited enclave where the data is not permitted to leave.
Subscribe
Written for chief information, technology, risk and operating officers in sectors where failure is expensive. No gated content, no logos as argument, and no more than one product mention per issue.
Fortnightly · Tuesday 07:00 ET · Unsubscribe in one click