Thalamus Advisory

Tools are necessary. A practice is what makes them work.

Buying an observability platform does not create resilience, any more than buying a treadmill creates fitness. We build the operating practice — measurement, escalation policy, evidence, and the organisational habits that survive contact with a bad night — supported by two products designed around it.

The position

Autonomy is only safe where it can be wrong safely

Every organisation in the sectors we cover is putting probabilistic systems into paths with physical, clinical, financial or political consequence. The question that decides whether that goes well is not how good the model is. It is whether the system around the model can detect and survive a wrong answer.

01

Calibrated, not confident

An investigation agent that always produces an answer is one an attacker can choose the answer for. Conclusions must arrive with evidence and a confidence score, and low confidence must change what happens next.

02

Independent verification

Self-verification produces correlated errors — the same training distribution, the same blind spots. Challenge must come from a different model, with authority to force human escalation.

03

Policy outside the prompt

Safety implemented as a sentence in a prompt is advisory. Guardrails belong at the execution boundary, enforced by infrastructure the model cannot talk its way past.

04

Evidence, not summaries

Retain the probes, their results and the agent-to-agent chain. The artefact a regulator asks for is the reasoning; the conclusion is only its summary.

Products

Two products, one loop

ThalamusTrace raises a correlated problem. Thalamus SRE investigates it, challenges its own conclusion, and routes a decision. Either works alone; together they close the loop from signal to remediated change without a human in the middle of the boring part.

Observability & AIOps

ThalamusTrace

OTLP-native ingest for traces, metrics and logs — the OpenTelemetry standard rather than a vendor agent, which is also what keeps your data portable back out again.

  • Golden signals and Apdex per service, with fleet health scoring
  • Span timeseries, failure breakdowns and critical-path attribution
  • Normalized DB statements and slow-query analysis
  • Log error-group regression measured against each group’s own baseline
  • SLOs on error rate, p95 latency and throughput — including floors
  • Correlated problems with a root-cause entity and affected-entity set
  • Signed webhooks to downstream receivers; multi-tenant throughout
Autonomy, bounded

Thalamus SRE

An agent fleet over that telemetry, with the boundary between automated and escalated set by you rather than by us.

  • SLO and error-budget monitoring with a problem queue and auto-dispatch
  • Seven-stage investigation: detect, triage, specialist fan-out, correlate with code, synthesize, challenge, route
  • Adversarial verifier running on a different model from the synthesis
  • Guardrails enforced at the execution boundary — confidence gates, blast-radius limits, human-in-the-loop approval, prompt-injection filtering
  • Application Resiliency: Auditor, Gap Tickets, Tuner, Vulnerability Detective
  • Blameless postmortems with action items, generated as evidence
  • ServiceNow, Jira, Azure DevOps, GitHub, Slack and Confluence integration
Deployment

Thalamus AI Cloud

Both products provision as managed instances — pooled or dedicated tenancy, per region, from one console — or deploy self-hosted inside your VPC, data centre or accredited enclave where the data is not permitted to leave. In regulated and classified estates, self-hosted is not a preference but an eligibility requirement, and the architecture was designed around it rather than accommodating it afterwards.

How we engage

Measurement first. Procurement is a consequence, not a starting point.

Most organisations do not need a tool first — they need to find out how bad the problem currently is, which is a measurement exercise. Every engagement below produces a number you did not previously have.

2–3 weeks

Calibration Audit

Reconstruct 30–50 recent incidents and measure the overturn rate of initial hypotheses, time lost to misdirected remediation, and how often the eventual cause sat outside the observable estate.

3–4 weeks

Telemetry Readiness

Whether your estate can support automated investigation at all: instrumentation coverage, retention against investigation latency, trace continuity across seams, and the diagnostic query paths that will be exercised at volume.

2 weeks

Injection Exposure Review

Where user-controlled data enters logs, which agents read those logs, and what authority those agents hold — delivered as a ranked remediation list.

2–3 weeks

Retention Economics

Your time-from-onset-to-investigation distribution plotted against current retention, with a tiered model sized to your real error ratio. Usually reduces cost while increasing diagnosable coverage.

3–4 weeks

Incident Archaeology

Re-investigate incidents previously closed as unreproducible, using correctly anchored windows. Typically recovers several genuine root causes — and writes the business case for everything else.

8–14 weeks

Guarded Rollout

Deployment against a bounded service domain, guardrails configured to your risk posture, escalation treated as a success state, and the overturn rate tracked from day one.

Sector-specific engagements — clinical path mapping, regulatory evidence readiness, pre-peak resilience audits, OT/IT boundary observability, enclave deployment design — are described in each sector’s issue.

What we will not do

The engagements we decline

A partner who will not say no is a supplier, and you already have those.

Start a conversation

The first question is always the same: what does an incident actually cost you?

Most executive teams have three of the four inputs and have never assembled the fourth. If you would like help putting the number together — or simply want the newsletter — both start here.

hello@thalamusadvisory.com · thalamusadvisory.com