Fortnightly · For the executive committee

Enterprise resilience in the age of AI

Your organisation is putting probabilistic systems into paths where failure has a physical, clinical, financial or political consequence. This is a briefing on what that changes — written for the people who will answer for it.

Every two weeks12 issues published10 sectors covered8–12 minute read

Why this exists

Most resilience reporting tells you what happened. This tells you what it costs, and what to do.

Every issue takes one real failure mechanism, explains it precisely enough that an engineer would recognise it, and translates it into the language an executive committee actually decides in — exposure, obligation, and the cost of continuing as you are.

01

The challenge, named specifically

Not “observability matters.” A mechanism — why a threshold alert cannot detect a feed that silently stopped, why a retry policy becomes an amplifier under load, why an AI investigator never reports uncertainty.

02

The cost of leaving it alone

Quantified where it can be, and honestly labelled where it is modelled. Regulatory penalty, downtime, engineer-hours, missed windows, and the second-order costs that never reach the incident report.

03

The AI upside, and the pitfalls

Autonomy compresses diagnosis and removes toil. It also introduces confidently wrong conclusions, log-borne prompt injection, unbounded spend and explainability obligations. Both halves, in every issue.

04

What mitigates it

The pattern first, the product only where it is the honest answer. Where the right move is organisational or architectural rather than a purchase, the issue says so.

The cost of not tackling it

4
Teams typically paged for a single upstream fault. Three spend the first hour proving it is not them.
Zero
Alerts generated by a data feed that stops entirely, under threshold-based monitoring.
4h 20m
Typical gap between incident onset and investigation — against a 60-minute diagnostic window.
Days
Duration over which the recovery cost of one missed logistics window accrues.

Each figure is drawn from an issue in the archive, where the reasoning and assumptions behind it are published in full. We do not print a number without its method.

Latest issue

Inside the Boundary

Commercial SaaS observability cannot follow classified data across the boundary. The result is that the estates with the highest consequence of failure frequently have the least visibility into it.

Issue 12 · Government, Defense & Aerospace · 10 min

The organisations with the most at stake are frequently the least instrumented, and the reason is architectural rather than cultural.

Coverage

The ten largest sectors of the US economy

Each has a characteristic way of failing, a different regulator, and a different reason that generic advice does not apply. One issue per sector, with the application landscape, the exposure, and the mitigation.

See the challenge and exposure for each sector →

What every issue contains

Seven blocks, in the same order, every time

So a returning reader can go straight to the part they came for — and so the argument cannot skip the inconvenient section.

01

The Signal

The event, filing or change that makes this the subject this fortnight.

02

The Mechanism

How the failure works at the level of the system. The section that earns the reader’s trust.

03

The Cost

Where the exposure actually lands — regulatory, operational, commercial, reputational.

04

The AI Stakes

What autonomy changes, both the upside and the new failure modes it introduces.

05

The Remedy

What to change, ordered by cost, cheapest first.

06

In Practice

Hosted or customized — how the tooling is actually deployed in that sector.

07

Getting There

The engagements that move an organisation from assessment to operating practice.

Always

The Counter-Argument

The strongest case against that issue’s advice, made honestly. When it does not apply, and what it costs.

Tools and a partner

Reading about a failure mechanism is not the same as closing it

Thalamus Advisory builds enterprise resilience practice, supported by two products that can be spun up as managed instances or deployed inside your own boundary.

Observability

ThalamusTrace

OTLP-native ingest for traces, metrics and logs. Golden signals and Apdex per service, critical-path attribution, baseline-relative log regression, and correlated problems that name one root-cause entity instead of paging nine teams.

Autonomy, bounded

Thalamus SRE

An agent fleet over that telemetry. Seven-stage investigation with an adversarial verifier on a separate model, guardrails enforced outside the prompt, and postmortems generated as evidence rather than reconstructed afterwards.

Deployment

Thalamus AI Cloud

Both products as managed instances, pooled or dedicated, provisioned per region — or self-hosted inside your VPC, data centre or accredited enclave where the data is not permitted to leave.

How the practice works, and how to engage it →

Subscribe

One mechanism, its cost, and what to do about it. Every two weeks.

Written for chief information, technology, risk and operating officers in sectors where failure is expensive. No gated content, no logos as argument, and no more than one product mention per issue.

Fortnightly · Tuesday 07:00 ET · Unsubscribe in one click