SoW:
DREW incident responder — automatic response to #ax-incidents
Tracking: A-4750.
DREW's dependency-remediation trunk is deployed. It runs supervised
on the mac mini (composer plus the supervisor and worker images from
build-and-release-drew.yml). This SoW replaces the old
docs/plans/drew.* plan, which we deleted in the same PR.
The plans format is deprecated (#3357). The old plan listed incident
triage as deferred work. That work starts now.
Objective. Build a bot that watches #ax-incidents
and reacts to events. When an incident appears, the bot investigates it:
it reads logs and databases through bard, reads incident.io, and checks
GitHub for recent changes. It posts what it finds in the incident
thread. If someone needs to act, it tags a human. Version 1 never
changes anything.
Decisions (agreed 2026-08-26)
- Use Slack Socket Mode. Socket Mode only makes
outgoing connections. The mini does not need to accept traffic from the
internet, and the connection survives IP changes. The bot is a small
Bolt (Python) app.
- Give the bot its own package. The responder is not
dependency remediation. It lives in its own
py/ package. It
reuses DREW's container, compose, and egress setup, but its code stays
separate from py/drew.
- Run it as an always-on service under composer. The
listener is a compose service with no schedule and a restart policy —
like the bare
drew service, but always up. Composer's label
watching, error alerts, and the GHCR image pipeline work for it without
changes. Keep the listener small: each investigation runs in its own
worker container, so a stuck investigation cannot take down the
listener.
- Read-only. No exceptions. The worker can: query
logs (ClickHouse), query the database (bard, read-only), read
incident.io, read GitHub (what merged or deployed recently), and post to
the thread. It cannot restart, change, or write anything. This rule also
protects us from prompt injection: alert text and log lines are
untrusted input, and a worker that cannot change anything cannot be
tricked into changing anything. Automatic fixes are a future SoW with
its own approval process.
- Escalate by tagging a human. When the bot concludes
that someone must act, it tags the person on the L2 support schedule and
stops. That schedule does not exist yet. Creating it is a separate
ticket, not part of this SoW. Until it exists, the tag target is a fixed
config value.
- One investigation at a time. One session per
incident. The session key is the incident.io page's Slack message, which
is the thread head. Alerts that arrive during an active investigation
join the open incident or wait in a queue. They never start a second
investigation. Alert storms collapse into one thread.
- Find open incidents in incident.io, not in Slack
history. The listener never pages back through #ax-incidents to
learn what is open — that read has no bound. On start and on reconnect,
it queries incident.io for the relevant open alerts, then searches Slack
pointedly for each one's thread head (the incident.io page message).
incident.io is the source of truth for what is open; Slack reads stay
small and bounded.
- Stay quiet by default. If a transient error fixed
itself, the bot only adds an emoji reaction. The bot posts findings only
when it is confident. "I investigated and found nothing useful" is a
valid silent result. This behavior decides whether people accept the bot
in the channel.
- Keep the conversation in the thread. The daemon
maps each
thread_ts to an agent session id. When a human
replies in the incident thread, the daemon passes the reply into the
same session. The bot answers with its full investigation context, in
the channel where responders already are.
- Offer a "migrate session" button. With its
findings, the bot posts an action that gives the user a command to copy
and paste. The command resumes the investigation session on the user's
own machine. How the transfer works is an open question — the session
transcript lives on the mini.
- Get runbooks by shallow clone, not from the image.
The mini has no repo checkout on purpose. But triage knowledge (the
incident skills, runbooks) changes faster than DREW's code, and an old
runbook during an incident is worse than none. So each worker does a
shallow read-only clone when it starts.
- Put limits on every investigation. The daemon
enforces a wall-clock timeout per investigation (the
worker_timeout_s lesson: a limit that is defined but not
enforced is no limit). Each investigation also gets a token budget, so
an unlimited log store cannot absorb unlimited spend before the answer
turns out to be "brief network blip."
PR breakdown
PR A — responder package
skeleton
The Socket Mode listener, joined to #ax-incidents. It detects
incident thread heads (the incident.io page message) from live events;
on start and on reconnect it recovers open incidents from incident.io
and finds their thread heads with a pointed Slack search, never by
scanning channel history. It keeps the thread_ts →
session-id map, runs the one-at-a-time queue with
fold-into-open-incident behavior, and has the emoji-ack path for
transients. No real investigation yet — a stub worker proves the spawn,
timeout, and exit-code plumbing.
PR B — investigation worker
The worker image: a headless agent harness plus a shallow runbook
clone at start. MCP wiring for bard (Postgres and ClickHouse,
read-only), logs, incident.io, and read-only gh. Posting
findings and tagging for escalation. Timeout (exit code 124) and
token-budget enforcement. The confidence gate on posting. The responder
gets its own egress allowlist (Slack, incident.io,
bard, ClickHouse Cloud, GitHub, Anthropic). The jail forks per worker
type; the dependabot workers' profile does not change.
PR C — mini integration
The compose service definition (always on, restart policy) under the
existing composer stack. The responder images added to
build-and-release-drew.yml. New .env entries:
Slack app and bot tokens, Anthropic key, bard credential, GitHub token.
The known mini launchd, TCC, and PATH problems apply to everything new
here.
PR D — conversation and
handoff
Relay thread replies into the live session (resume by session id).
Build the migrate-session action with whatever transfer method the open
question settles on.
Ops items (no PRs)
- A bard service credential for the responder (see
Open questions).
- The L2 support schedule — file as its own ticket;
this SoW only uses it.
- A replay corpus: pick about 10 past incidents with
known root causes from #ax-incidents history, for the gate-1 test
harness.
Open questions
- Which identity gets log and database access?
Workers are fresh containers with no built-in credentials. Proposal:
create a dedicated read-only service user on bard (and ClickHouse) with
the smallest possible access, and put its credential in the mini's
.env. Then the bot's queries are attributable to the bot,
and no human credential is borrowed. Needs provisioning on the bard
side.
- How does migrate-session work? The session JSONL
lives on the mini. Option 1: the posted command copies the transcript
from the mini with scp, then runs
claude --resume locally.
Option 2: the bot exports the transcript to a place the user can fetch
it, and the command pulls it and resumes. Decide in PR D; ssh plus scp
needs the least machinery.
- Where is the confidence bar? What separates an
emoji ack from a posted finding, and what counts as "transient, fixed
itself"? Start strict — post only with a concrete, correlated cause —
and tune against the replay corpus.
Gates
- The replay test passes. Against the replay corpus
of past incidents, the worker finds the known root cause, or correctly
stays quiet, before we turn the bot on in the live channel.
- Read-only is verified. No tool that changes
anything is reachable from the worker, and the egress allowlist holds.
Same verification shape as the dependabot jail.
- The storm test passes. A flapping alert produces
one thread, one investigation, and one summary — never parallel
investigations.
- The handoff round-trip works. A human reply in the
thread gets an answer with full investigation context, and
migrate-session produces a working local resume.
Deliberately deferred
- Automatic fixes. A separate SoW. The bot gets no
tool that changes anything until an approval process exists.
- Incident memory (a store of learned root causes).
Explicitly "not yet."
- Parallel investigations. One at a time until the
quiet and confidence behavior is proven.
- The L2 schedule itself. Separate ticket.