1 Operator Guide
Ope Olatunji edited this page 2026-05-14 21:50:13 -04:00
This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

Operator Guide

Running AgenticMail in production. Aimed at someone who already has it installed and wants to keep it healthy.

Process model

AgenticMail in production runs three long-lived processes:

  1. The API server — listens on :3829, started by agenticmail (the CLI). One process per host.
  2. The dispatcher daemon — agenticmail-claudecode-dispatcher (or whichever host package is active). PM2-managed.
  3. Stalwart — bundled IMAP/SMTP server, managed by AgenticMail itself.

All three need to be running. The web UI on :3829 and the agent-coordination layer both depend on the API + Stalwart; the dispatcher is what makes agents wake on incoming mail.

PM2 layout

The dispatcher is the only AgenticMail process designed for PM2. Recommended ecosystem entry:

// ~/.agenticmail/ecosystem.config.cjs
module.exports = {
  apps: [
    {
      name: 'agenticmail-claudecode-dispatcher',
      script: '/opt/homebrew/lib/node_modules/@agenticmail/cli/node_modules/@agenticmail/claudecode/dist/dispatcher-bin.js',
      autorestart: true,
      max_restarts: 10,
      min_uptime: '30s',
      kill_timeout: 5000,
      env: {
        // Optional overrides
        // AGENTICMAIL_API_URL: 'http://127.0.0.1:3829',
        // AGENTICMAIL_DISPATCHER_MAX: '50',
        // AGENTICMAIL_DISPATCHER_SYNC: '30000',
      },
    },
  ],
};

Bring it up: pm2 start ~/.agenticmail/ecosystem.config.cjs && pm2 save.

Healthy-state checks

You want to see:

pm2 list
│ agenticmail-claudecode-dispatcher │ default │ 0.2.8 │ fork │ online │ 0 restarts (recent) │
curl http://127.0.0.1:3829/api/agenticmail/health
{ "ok": true, "version": "0.9.11" }
pm2 logs agenticmail-claudecode-dispatcher --lines 5 --nostream
[dispatcher] starting (maxConcurrent=50, syncEvery=30000ms)
[dispatcher] opening SSE for "vesper" (vesper@localhost) (restored 4 seen UIDs, lastSeenUid=57)
[dispatcher] opening SSE for "orion"  (orion@localhost)  (restored 7 seen UIDs, lastSeenUid=57)
...

If you don't see the restored N seen UIDs line on startup, you're on a pre-0.9.8 build — upgrade so restart recovery works.

Upgrade procedure

Standard:

npm install -g @agenticmail/cli@latest
pm2 restart agenticmail-claudecode-dispatcher
pm2 logs agenticmail-claudecode-dispatcher --lines 30

The npm install -g updates the CLI which pulls in the latest API + claudecode via semver-compatible deps. The PM2 restart picks up the new dispatcher binary because PM2's script path resolves to the freshly-updated dispatcher-bin.js.

You also need to restart the API process (it doesn't auto-update — there's no PM2 entry for it by default). Stop whatever's running on :3829 and re-launch your usual agenticmail command.

Known operational gotchas

Wake-budget circuit breaker

Each (agent, thread) pair has a cap of 10 wakes per 24 hours. When tripped, you'll see:

[dispatcher] wake-budget exhausted for "vesper" on thread "..." (count=10, cap=10); muted for ~1300min

This is by design — it stops a runaway reply loop. The budget is in-memory; today (as of 0.9.14) a dispatcher restart resets it. This will move into dispatcher-state.json in the next milestone so restarts don't bypass the breaker.

Dispatcher silent on broadcast

If you send to: alice, cc: bob,carol,dan without an explicit wake: [...], only Alice wakes (since 0.9.0). Bob/Carol/Dan get the mail in their inboxes but no Claude turn fires for them. To wake everyone, pass wake: 'all'. To wake nobody, pass wake: [].

Auto-compact wipes Claude's context

Claude Code's auto-compact keeps the same session_id but wipes the model's working context. Pre-0.9.12 the AgenticMail capabilities preamble was injected once per session, so the model would lose awareness of the toolbelt after compact. Since 0.9.12 we hook SessionStart (which Claude Code fires on startup / resume / compact) so the preamble re-lands after every compact event. You'll see this in your ~/.claude/settings.json:

{
  "hooks": {
    "SessionStart": [{ "matcher": "", "hooks": [{ "type": "command", "command": "node \"...\"" }] }],
    "UserPromptSubmit": [...],
    "Stop": [...]
  }
}

If any of those three are missing after an upgrade, re-run agenticmail-claudecode install.

Receiver cache and the "one IMAP connection per agent" problem

The API caches one ImapFlow connection per agent (10-minute TTL). If you have 50 agents all active in the web UI at once and you click message-detail rapidly across them, requests serialize on the per-agent mailbox lock. Not a correctness issue, just a latency surprise. Mid-term roadmap: add a small per-agent connection pool.

Web UI's 6-connection limit (FIXED in 0.9.9)

Browsers cap at 6 connections per origin. Pre-0.9.9, the UI opened one SSE per agent, saturating the cap on accounts with ≥5 agents. Refresh / message-fetch / attachment-download would all hang waiting for an SSE slot to free. Fixed in 0.9.9 by multiplexing all per-agent events through one shared /system/events stream. If you still see hangs after 0.9.9, hard-refresh the browser to clear cached JS.

Disk usage

AgenticMail writes to ~/.agenticmail/:

Path Grows with Bounded?
agenticmail.db accounts + tasks + tag rows Yes (rows have natural ceilings)
thread-cache/ distinct threads × K envelopes Yes (K is small, ~5)
agent-memory/ agents × threads × markdown Loosely (one markdown per (agent, thread))
dispatcher-state.json accounts × 256 UIDs each Yes (hard cap, ~50 KB total)
worker-logs/ worker turns Not bounded — clean periodically
logs/ API + dispatcher stdout Rotate manually

worker-logs/ is the only thing that needs explicit operator attention. Run a weekly find ~/.agenticmail/worker-logs -mtime +7 -delete cron.

Diagnostics

The MCP tool check_activity (from inside any Claude Code session) is the operator's primary view:

mcp__agenticmail__check_activity()

Returns the dispatcher's process state (uptime, channels, queue size, recent worker history). The same data is on GET /api/agenticmail/dispatcher/activity for scripted checks.

The web UI's activity badges (between the search bar and notification bell) show real-time worker state on every active agent. If you don't see badges when you know an agent is running, hard-refresh — pre-0.9.11 the badges only painted on the next heartbeat (~30s); 0.9.11 added a one-shot backfill.

When the dispatcher won't dispatch

The four most common causes:

  1. Wake budget tripped — pm2 logs | grep wake-budget. Reset by waiting for the window or restarting (caveat below — 0.9.14 still in-memory).
  2. Wake allowlist empty/wrong — pm2 logs | grep wake allowlist excludes. Means the sender's wake list didn't include this agent. Check what was sent.
  3. SSE channel dropped — pm2 logs | grep SSE for "...".ended. Auto-reconnects with backoff; if it stays disconnected, the API is down.
  4. Dispatcher itself crashed — pm2 list shows errored or stopped. Pre-0.9.6, this could be a bundling error; pre-0.9.7, the lone-leading-edge crash. Both fixed. If you still see crashes on a current build, the uncaughtException guard logs the error and continues — pm2 logs | grep uncaughtException to find it.