Ask — search your own work

Dashboard → Ask. A conversation over everything you have synced — memory entries, prompts and session records — and HarnessLink's own guides.

shell
what did we decide about autobank backfill?
  → and when was that agreed?
how do I restore HarnessLink on a new machine?

What happens when you ask.

  1. A small, fast model pulls the keywords out of your question (and resolves

    a follow-up such as "and when was that?" against the thread). They are shown under your question as Searched for. Your own words are still searched, at lower weight, so a missed keyword cannot hide a result.

  2. Memory entries, prompts, session titles, the guides (this file,

    INSTALL.md, MEMORY_BANKING.md, MULTI_HARNESS.md) and the semantic index are searched at once.

  3. Sources are ranked so that what answers beats what asked: a bank

    entry's result or a guide section outranks the prompt that requested the work, and a prompt whose banked turn is already a source is dropped. Temporary-folder entries, generated indexes and daily summaries are never sources.

  4. The sources appear straight away, each with a link to the exact place —

    the memory entry, the prompt inside its session, or the guide section — and the answer streams in under them.

Answers are grounded. The answering model uses only the sources, cites them as [n] (click one to jump to it), never expands an acronym the sources do not expand, and treats a prompt as a request, not an outcome. When the sources do not answer the question it says "I couldn't find this in your synced work or the HarnessLink guides." and shows that as a notice rather than an answer. Citations pointing outside the source list are discarded.

Filters narrow before the search runs: kind (memory / prompt / session / guides), project, and a date range. A project or date filter leaves the guides out.

Details of sources opens the full evidence: one card per source, cited entries marked. Check the answer against them — that is what they are for.

Keeping the index current

The semantic index embeds new memory entries (bank entries first), sessions and prompts every hour by itself, on plans with semantic search. **Rebuild index** runs a pass now. Until an entry is embedded it is still found by its words.

Typical timings: sources in about 2–3 seconds, the first words of the answer about a second later, the whole answer in under 7 seconds.

History — every prompt, from every harness

Dashboard → History. Every prompt you typed, newest first. Each row has a badge naming the harness it was typed into (omp, claude-code, codex, pi, antigravity, and superpi for prompts from HarnessLink's old built-in agent); click the badge, or a name in the Harnesses panel, to show only that harness.

The HarnessLink service reads new prompts from each harness with capture on about every 30 seconds and uploads queued rows every 5–10 seconds, so a prompt usually appears within a minute. superpi cloud sync does both now. Cursor prompts are not captured.

Policies — rules for your organization

Dashboard → Policies. Write the rules your organization wants its harnesses to follow. Tool and command rules are enforced before every tool call in Claude Code, Codex and omp (see Org policy — enforced before every tool call); pi, Cursor and Antigravity have no HarnessLink tool hook.

Each rule is effect + target + optional conditions.

EffectMeaning
Explicit denyBlocked, always
Explicit promptAlways ask first
Explicit allowOutranks a lower-precedence deny or prompt; never auto-approves
Conditional deny / prompt / allowThe same, only when conditions match

Explicit beats conditional; within each, deny > prompt > allow.

Targets are a tool (bash, write, edit, computer), a shell pattern (rm -rf *), or a model (claude-opus-*, anthropic/*). Conditions are cwd_prefix and model for tools and commands, and harness (claude-code, codex, omp, pi, antigravity, other) for models. A model rule only allows or denies: once any allow applies to you, models no allow matches are refused.

Usage limits (Policies → Usage limits) cap spend in USD or tokens (input, output, cache reads and cache writes) per UTC day or UTC month, either per machine or per person across all their machines, optionally for matching models only. USD is HarnessLink's estimate from list prices, not your provider's bill. A per-person limit counts a person's other machines once they sync, usually within about a minute.

Model rules and limits are enforced by HarnessLink 17.5.50 and later, in the local gateway, before a model call leaves the machine: every call from Claude Code and Codex, and calls through the Anthropic and OpenAI Codex providers in omp and pi, once superpi gateway connect has routed them. Cursor, Antigravity and other providers go direct and are not covered.

A refused call gets a 403 from the gateway naming the rule or limit, your organisation and the Policies page; for a limit it also says how much is used, of how much, and when the window resets (UTC midnight, or the first of the month). Limits are checked before each call: a response that starts under a limit finishes even if it goes over, and the next call is refused. If the policy file on a machine cannot be read, model calls are refused until the HarnessLink service rewrites it; on a machine that has never fetched a policy, nothing is enforced. superpi and superpi cloud status show each limit with its current usage.

shell
Explicit deny        command   rm -rf *
Explicit deny        command   terraform apply *
Conditional prompt   tool      bash        when cwd_prefix=/Users/you/Workspace
Conditional allow    tool      read        when cwd_prefix=/Users/you/Workspace

Memory

Dashboard → Memory. What your harnesses' sessions left behind, with full-text search and dropdown filters for project (by folder name), harness, type, machine and time; the filters live in the page address, so links and Back work. Each row shows the entry's title. Temporary folders (/tmp, the OS temp dir) and generated files (each project's MAIN.md index and daily summaries) are hidden unless you ask for them. Click an entry to read it; its file and folder paths are under Details. Turns run in a temporary folder are not banked at all.

What syncs, once setup has run (all through the memories sync category):

  • HarnessLink's memory tree (~/.superpi/agent/memories): bank entries,

    MAIN.md, daily summaries and state/ snapshots — see MEMORY_BANKING.md.

  • Each captured harness's own memory files, in place, about every 5

    minutes and on superpi cloud sync: Claude Code CLAUDE.md (global and per project) and auto memory; Codex AGENTS.md, AGENTS.override.md and ~/.codex/memories/; omp and pi AGENTS.md and memories/; Antigravity GEMINI.md, rules, knowledge and brain artifacts; Cursor ~/.cursorrules. A file whose project is known is filed under that project (harness/<id>/…); the rest go to a project named after the harness. Only .md files up to 1 MiB.

Every file is redacted client-side before upload, and an unredactable file is skipped rather than sent.

Importing older memory

bash
superpi cloud import-memories          # preview
superpi cloud import-memories --apply  # write (setup runs this once for you)

Reconstructs documents from a MemPalace store and ~/.hermes (MEMORY.md, SOUL.md, USER.md) into the normal memory tree, with a provenance header recording where each came from. They then sync like any other memory. Each machine files its imports under its own machine-<id>/ folder, so two machines with the same home path (two servers as ubuntu) keep separate copies instead of replacing each other's. Imports from earlier releases are moved into that folder when their content matches; anything else is left where it is. Idempotent — re-running rewrites nothing unchanged. Live harness memory such as CLAUDE.md is not imported: it already syncs in place.

Recall — ask before you re-derive

bash
superpi cloud recall "rate limit"              # this folder's project + linked projects
superpi cloud recall "rate limit" --all        # every project you own
superpi cloud recall "rate limit" --project=<slug>

Full-text search over your banked memory. Every finished turn in each harness with autobank on (omp, pi, Claude Code, Codex; Antigravity per conversation) becomes a bank entry automatically (memory.autoBank, default on; subagent turns are not banked), so the outcome of past investigations is one query away — far cheaper than an agent re-deriving it. Entries banked by HarnessLink's old built-in agent stay searchable. superpi cloud bank --title=… [--tags=…] records something by hand; superpi cloud link-memory --project=<slug> makes another project's memory part of this folder's default recall scope.

Context brief — seed a session instead of replaying one

bash
superpi cloud brief                            # print to stdout
superpi cloud brief --budget=6000 --out=brief.md
superpi cloud brief --recall="deploy freeze"   # add a recall section (local index, no network)

A deterministic, budgeted summary of the current folder built from memory, not transcripts: the human-written part of MAIN.md, the newest ten bank entries, the harness's own memory files (omp's raw_memories.md and newest rollout summary — folders worked in upstream omp carry context there and may have no bank at all), and an optional recall section. Sections are kept whole or dropped, with a note naming what was dropped, so the brief never exceeds --budget characters (default 8 000).

Paste it into a fresh session in any harness. Re-establishing context is the largest token cost in agent work; a brief costs a few hundred tokens where replaying the transcript it stands in for can cost millions.

The local index — memory search without the network

The HarnessLink service keeps a SQLite full-text index of this machine's memory tree at ~/.superpi/agent/memory-index.db. memory_search, memory_get, context_brief's recall and the context hooks read it first, so they answer offline, in milliseconds, at no cost.

  • What is indexed: every .md file under ~/.superpi/agent/memories.

    A bank entry is split the way the dashboard's Ask reads it — title, tags, ## Prompt, the ## Actions step log (one line per tool call) and ## Result — and ranked so a match in the Result counts most, then the step log, then the prompt. Words match across inflections ("permissions" finds "permission"). Each project's generated MAIN.md index and daily summaries are kept out of results, as on the dashboard.

  • How it stays current: one pass when the service starts (files whose

    size and modification time are unchanged are not read again), then the service's file watcher applies each add, edit, rename and delete as it happens, and a full check every 5 minutes catches anything it missed. Without the service, superpi mcp, the hooks and superpi cloud brief bring the index up to date for the projects they are about to search before searching.

  • Cost: on 2,533 memory files (15 MB) the first build took about 0.5 s

    and the index is 12 MB; a later service start checks every file in about 30 ms. A search takes 1–6 ms.

  • Other machines: memory_search also asks the cloud, for entries from

    your other machines and linked projects, and waits for it at most 1.5 s. When the cloud doesn't answer in time, you're offline or not signed in, you get this machine's results and one line saying so.

The file is a cache. Deleting it is safe: it is rebuilt from the memory files, and an unreadable one is rebuilt automatically.

Resuming sessions across harnesses

bash
superpi resume

One picker over every session on the machine that a harness can resume natively (Claude Code, Codex, pi and omp; Antigravity and Cursor transcripts are not offered, since no harness can open them), plus sessions left by HarnessLink's old built-in agent — newest first, each row ending with its replay cost. Resuming does not re-send the file: the harness rebuilds the model context, which compaction keeps bounded, so the cost shown is the context the provider actually saw on the session's last request — read from the transcript's own last usage record (~69k tok). A transcript with no usage record shows a bytes-based ceiling instead (≤425k tok). Measured live, the byte heuristic alone overstated a 17 MB session by 63×. Pick a session, then pick where to open it: omp is preselected for omp-format sessions (including the old HarnessLink agent's), claude/codex for their own.

Handoff copies the session file into the target harness's store — never moves or overwrites it — then launches the harness's native resume. Cross- family conversion is refused rather than approximated: a claude-code transcript is not an omp session file.

At 50 000 tokens and above the harness picker gains a fifth option: *brief — seed a fresh session from project memory instead of replaying ~N tokens*. It runs superpi cloud brief in the session's folder and works for any family, because it is built from the folder's memory rather than the transcript. The native resume stays the default; the cheaper path is simply visible at the moment the cost is being paid.

Profile sync — moving machines

bash
superpi cloud push-profile             # snapshot this machine now (the HarnessLink service also does it every 5 minutes when something changed)
superpi cloud push-profile --dry-run   # print exactly what would be uploaded; uploads nothing
superpi cloud restore                  # show what this machine would get
superpi cloud restore --apply          # write it (config.yml is backed up first)
superpi cloud restore --approve-mcp    # approve restored MCP servers left pending

The device profile holds HarnessLink's own config.yml, the MCP servers its hub runs and the skills and agents the old built-in agent left in its own folder — and the setup of every harness installed here (including skills superpi add put there), Markdown files only:

HarnessWhat travels
Claude Code (~/.claude)agents/, commands/, rules/, output-styles/, skills/, CLAUDE.md
Codex (~/.codex)prompts/, skills/, AGENTS.md
omp (~/.omp/agent), pi (~/.pi/agent)agents/, commands/, rules/, prompts/, skills/, AGENTS.md, SYSTEM.md

Skill scripts, extensions and anything else that could run are never captured. Each file is redacted and capped at 32 KiB, and the whole profile stays under 1 MiB (instruction files first, skills last; anything left out is counted). Harness config files (settings.json, config.toml, config.yml, models.yml …) are listed by key name only and never restored — superpi cloud setup redoes the wiring instead.

restore reads the newest profile pushed by another device unless you name one with --device <id>. Harness files are restored additively: a missing file is created, an existing one is never overwritten, and a harness not installed on this machine is skipped with a note to install it and re-run.

Secret values are never uploaded. MCP env/headers entries are reduced to their key names, so restore tells you which secrets to re-supply and cannot supply them itself. Syncing credentials would put a fleet-wide secret store behind one web session.

Restored MCP servers stay off until you approve them. Nothing from the cloud runs on this machine without your consent: every MCP server a restore adds is written with "enabled": false and recorded in ~/.superpi/agent/mcp-pending.json. restore --apply then shows each one with the command or URL it would run and asks Approve? [y/N]; without a terminal nothing is approved. --approve-mcp approves them all without asking — with --apply for this restore, alone for what an earlier restore left pending. An organization can pre-approve servers in its policy bundle with "mcp": { "allow": ["uvx", "https://mcp.example.com/sse"] }: entries match the exact command or URL, never the server name (the name is chosen by whoever pushed the profile). superpi cloud setup restores the harness files (additively) but never config.yml or MCP servers — run superpi cloud restore --apply for those.

Machine-bound servers stay disabled even when approved. A restored stdio server whose command is an absolute path that does not exist on this machine (typically /Users/<other-user>/…/some-mcp) keeps "enabled": false, and the apply report lists it with the re-enable instruction. A config that fails on every session start is worse than one that says why it is off. Servers that need a per-machine credential (an http MCP answering 401) still surface at connect time; the report names the secrets they need.

Templates — the registry

bash
superpi registry search intent
superpi add intent-discovery              # shows where it will write, then asks; --yes skips
superpi add intent-discovery --harness claude-code   # one harness only
superpi add --list                       # what HarnessLink installed, and which harness reads it
superpi remove intent-discovery           # removes only what HarnessLink wrote

Dashboard → Templates lists reusable skills and agents. superpi add <name> copies each skill into the user-level skill folder of every harness enabled in cloud.harnesses or detected on this machine (its folder exists or its CLI is on PATH), so the harness loads it in its next session:

HarnessSkill folderSource
Codex~/.agents/skills/<name>/Codex: Build skills
Claude Code~/.claude/skills/<name>/Claude Code: Skills
omp~/.omp/agent/skills/<name>/omp config-usage
pi~/.pi/agent/skills/<name>/, also reads ~/.agents/skills/pi: Skills
Antigravity CLI~/.gemini/antigravity-cli/skills/<name>/Antigravity: Agent skills
Cursor~/.cursor/skills/<name>/, also reads ~/.agents/skills/, ~/.claude/skills/, ~/.codex/skills/Cursor: Agent Skills

A harness that also reads a folder another harness's copy goes to (pi and Cursor read ~/.agents/skills/; Cursor reads ~/.claude/skills/) uses that copy instead of getting a duplicate it would load twice. Agent definitions (agents/*.md) go to omp only (~/.omp/agent/agents/); when omp is not enabled or detected they are listed as skipped, never dropped silently.

HarnessLink records every path it writes in ~/.superpi/agent/addons.json. A skill folder of the same name that HarnessLink did not write is yours: add skips that harness and says so (exit code 1), even with --force. Running add again is a no-op when nothing changed and updates HarnessLink's own copy when the registry has a newer one; a copy you edited is left alone unless you pass --force. superpi remove <name> deletes only the files HarnessLink wrote, unedited (or with --force), keeps anything you added to the folder, and keeps a folder another installed item still uses. The AI-native SDLC pack brings the intent → spec → plan artifact chain to every project:

SkillStageProduces
intent-discoveryPlanintent/<date>-<slug>.intent.md — interview the originator to completion; intent/ is the backlog
spec-from-intentDesignspec/<slug>.spec.md — requirements traced to the intent, org instructions honoured by construction
plan-from-specBuildplan/<slug>.plan.md — files, ordered/parallel steps, blast-radius interrogation, proof per step; implementable with no other context
maintenance-intentMaintaina diagnosed intent from an alert, log or ticket — evidence quoted, suggestions never actions

Slugs match across the three folders so a chain is greppable end to end, and the committed/done stages bank to memory — the chain plus recall is the project's institutional record.

superpi (superpi add superpi) is the cross-harness playbook. It teaches the agent in any harness — Claude Code, codex, omp, or a plain shell — to reach cloud recall and briefs, find and resume where a session stopped (superpi sessions list --json), and sync on demand (superpi cloud sync --json).

Extensions — code that runs inside omp

bash
superpi add context-compress        # asks for consent; --yes skips

Registry items of kind extension install into omp's own extension directory (~/.omp/agent/extensions/), because they are code the harness executes, not prose; no other harness runs omp extensions, so without omp they are reported as skipped. Only the official registry and your org's shared artifacts can supply them; third-party @alias sources are refused.

context-compress keeps large tool outputs out of the model's context. On omp's tool_result hook, any bash/eval/grep/glob/MCP output over 8 KB is replaced by its first 40 lines, last 20 lines, every line that looks like an error or warning, and a marker:

shell
…3943 lines omitted (18.3 KB) — call retrieve_output({"id":"<toolCallId>"}) for the full output

The original is written to ~/.superpi/agent/outputs/<toolCallId>.txt (0600, pruned after 14 days) and a retrieve_output tool returns it, or a line range of it, on demand. read, edit and write results are never touched — hashline anchors and diffs are what the agent edits against. The compression is a pure function of the output, applied once when the result arrives, so the provider's prompt cache stays valid. Measured on real sessions, tool results are about half of billable context and the largest 6% of them carry half of those bytes; expect a 10–20% reduction in input-side spend. Live: a 4 000-line seq entered context as 335 characters and the model retrieved lines 2000–2002 exactly.

qmd — local search over your memories

bash
superpi add qmd            # shows the plan, then asks once; --yes skips (required without a terminal)
superpi add --list         # qmd's binary, collection, models and MCP entries
superpi remove qmd         # removes only the MCP entries and the collection HarnessLink set up

qmd is a local search engine for Markdown: BM25 keywords, vector search and LLM reranking, all on this machine with GGUF models. superpi add qmd:

  1. installs qmd with bun install -g @tobilu/qmd when qmd is not on

    PATH (qmd needs Node.js 22 or newer);

  2. creates the qmd collection harnesslink-memories over

    ~/.superpi/agent/memories (**/*.md) with a context that tells search what the entries are; re-running updates it, never adds a second one;

  3. registers qmd's stdio MCP server (qmd mcp, server name qmd) in every

    enabled or detected harness HarnessLink writes MCP config for: omp (~/.omp/agent/mcp.json), Cursor (~/.cursor/mcp.json), Codex ([mcp_servers.qmd] in ~/.codex/config.toml) and Claude Code (claude mcp add --scope user qmd -- qmd mcp; skipped when the claude CLI is not on PATH);

  4. downloads qmd's three models with qmd pull and runs the first

    qmd embed: about 2.3 GB (embedding 334 MB, reranker 639 MB, query expansion 1.3 GB) into ~/.cache/qmd/models. The confirmation states the size before anything is downloaded.

A qmd MCP entry you added yourself is never overwritten or removed (the harness is listed as skipped, exit code 1); the same rule as skills, recorded in ~/.superpi/agent/addons.json. superpi remove qmd leaves qmd itself, its models and your own collections in place (bun remove -g @tobilu/qmd uninstalls qmd).

Freshness. While the add-on is installed, the HarnessLink service runs qmd update 30 seconds after memories stop changing (and once at service start), so new entries are found by keyword search within a minute. It does not embed in the background: run qmd embed now and then to add vectors for new entries. qmd update re-indexes every collection and runs each collection's update command, so when one of your collections has one (qmd collection update-cmd, e.g. git pull) the service does not run it and says so in its log; run qmd update yourself.

Delegation — one harness hands a task to another

bash
superpi run --harness codex "why does the build fail?"                 # headless, in this directory
superpi run --harness claude-code --cwd ~/src/api --timeout 5m "…"     # another directory, stopped after 5 min
cat task.md | superpi run --harness omp --json                         # task on stdin, result as JSON

superpi run runs the harness headless, streams what it does, and ends with its status, exit code and a summary; the command exits with the harness's exit code. Inside any harness, the delegate MCP tool does the same: {harness, task, cwd?, timeout_s?, model?} → {harness, status, exit_code, summary, transcript_ref, run_id}. A call waits about 25 s, under the 30 s MCP timeout of omp and others; a longer run comes back as status: "running" with its run_id, and delegate_status {run_id} waits again for the result (cancel: true stops it). The tools run in the calling harness's own superpi mcp relay, never in the HarnessLink service, so the child inherits the caller's environment and goes away with the caller's session.

How each harness is run. No permission-bypass flag is ever passed (--dangerously-skip-permissions, --yolo, --full-auto, --auto-approve, --dangerously-bypass-hook-trust, …), and terminal wrapper shims on PATH (cmux's) are skipped because they add such flags. Each harness keeps its own default non-interactive permissions:

HarnessInvocationWhat a delegated run may doHarnessLink hooks in the headless run
Claude Codeclaude -p --output-format stream-json --verbose [--model=<id>], task on stdinStarts in Claude Code's built-in mode for -p (default, or auto where feature flags aren't fetched, as behind the gateway) unless your settings set defaultMode. Anything that would prompt is denied, and the summary lists the refused tools.Yes: -p loads ~/.claude/settings.json hooks (no --bare); org policy "prompt" rules become asks nobody answers, so they're denied.
Codexcodex exec --json [--model=<id>] -, task on stdinRead-only sandbox, no approvals: it reads and answers, edits are refused. Refuses to start outside a git repository (no --skip-git-repo-check).Only if you trusted HarnessLink's hooks in Codex's /hooks; untrusted hooks are skipped and no bypass is passed.
ompomp -p --mode json [--model=<id>], task on stdinYour tools.approvalMode (omp's default lets tools run).Yes: the superpi-context.ts extension loads in print mode; "prompt" rules block (no UI).
pipi -p --mode json [--model=<id>] <task>pi asks before nothing; untrusted project files (extensions, settings) are skipped headless.None: HarnessLink installs no pi hook.
Antigravityagy --output-format stream-json [--model=<id>] -p=<task>request-review: anything that needs permission (shell commands, for example) is auto-denied; allow it with a permissions.allow rule in Antigravity's settings.None: Antigravity has no documented hook.

The gateway is wired in each harness's own config (superpi gateway connect), so a delegated Claude Code, Codex, omp or pi run goes through it and its model rules and usage limits like any session. A harness that isn't installed is refused with <harness> is not installed (no <binary> on PATH); skipped.

Limits.

  • Depth: the child gets SUPERPI_DELEGATION_DEPTH = caller's + 1; a session

    at depth 2 may not delegate. A harness that strips the env from its tools (Codex gives MCP servers a minimal one) doesn't escape it: a caller whose ancestor process is a running delegated harness counts from that run.

  • 3 delegations at a time per machine, across every session.
  • Every run has a timeout (15 min by default, at most 4 h). On timeout or

    cancel (Ctrl-C, cancel: true, the calling session ending), the harness's whole process tree gets SIGTERM, then SIGKILL after 5 s.

  • cwd must exist. superpi run stays in the current directory unless you

    pass --cwd, which may name any existing directory. The delegate tool only accepts a cwd inside the caller's project root (its git top level).

Runs and memory. The child runs with SUPERPI_RUN_ID, SUPERPI_PHASE = delegate-<harness> and SUPERPI_ROSTER, so it shows on the Runs page: in the caller's run when the caller has a SUPERPI_RUN_ID, else in the calling Claude Code or Codex session's own run (<harness>-<session id>, the id subagent runs use, with the caller linked as phase main), else in a run named after the delegation. A finished run is banked at once as one memory entry tagged delegated, with the delegated prompt and its result, through the autobank's parsers and claim store (the service never banks it a second time); transcript_ref is its {project, rel_path} for memory_get. It follows memory.autoBank and superpi cloud autobank like any session. A failed, timed-out or cancelled run banks nothing; so does a run that exits 0 without an answer (Antigravity does when its permissions refused every step), which is reported as failed. Each run's record and raw output stay in ~/.superpi/agent/delegations/ for 7 days.

Runs — watch a software factory from outside

bash
export SUPERPI_RUN_ID=feature-42 SUPERPI_PHASE=build SUPERPI_ROSTER=cheap
claude                                          # or codex, omp — this session joins run feature-42
superpi cloud runs                              # runs, phases, tokens, cost

Dashboard → Runs. Sessions that carry a run id are grouped into runs: one row per run with its roster, harnesses, session count, tokens and cost, and a phase strip in start order (plan → build → test → review). Expand a run to see each phase's session. Every phase, tool call and response is already in the synced transcript, so the run view is the trace the factory pattern asks you to read instead of reaching into the box.

A session joins a run in one of two ways. Neither needs a command.

1. A run in the environment. Set SUPERPI_RUN_ID (letters, digits, ., _, -; up to 64), and optionally SUPERPI_PHASE and SUPERPI_ROSTER (lowercase letters, digits, -; up to 32), before you start the harness. Invalid values are ignored.

  • Claude Code, Codex, omp: HarnessLink's context hook sees the variables when

    the session starts or you send a prompt, and records them for that session in ~/.superpi/agent/run-links.db. The HarnessLink service stamps the run on the session's next upload (within about 5 minutes) and on its subagents'. A session uploaded before the link existed is sent again so it joins the run. This needs the context hooks (superpi harness context; Codex runs a new hook only after you trust it once in /hooks), transcript sync for that harness, and the sessions sync category. A session started before the hooks were installed is not linked.

  • pi, Antigravity and Cursor have no context hook, so environment runs

    do not apply to them.

2. Subagents. A session that started subagents forms a run with them automatically: run id <harness>-<parent session id>, phase main for the parent and the subagent's name for each child (omp and pi: the subagent file name; Codex: the agent role; Claude Code: subagent). This works for omp, pi, Claude Code and Codex. Subagent transcripts upload as sessions of their own but stay out of superpi resume and session lists. An environment run wins over subagent grouping and also covers the subagents, even when the session already reached the cloud inside its subagent run.

Tokens and cost per session are computed server-side from the transcript's own usage records, for omp and Claude Code transcripts alike.

Run tokens — the credential boundary

bash
export SUPERPI_CLOUD_TOKEN=$(superpi cloud run-token --run=feature-42 --ttl=2h)
superpi cloud run-token --run=feature-42 --revoke   # at teardown

A run token is a short-lived child of your device token: sync:write only by default, never broader than its parent, at most 24 hours, invalid the moment its parent is revoked, and unable to mint children of its own. Hand it to a sandbox and the box can stream traces up but cannot read your memory, sessions or profile.

The token is also the run's identity: every session uploaded with it is stamped with that run id server-side, whatever the client says. A sandboxed Claude Code or Codex needs no environment plumbing to be attributed — the credential does it. One level of nesting, enforced by credentials rather than by trust.

Cost

Dashboard → Overview shows month-to-date spend by model. Usage is the analytical view: time series, per-model breakdown, gateway savings. Both read your harnesses' model calls through the HarnessLink gateway, plus any usage the old built-in agent recorded.

$/Mtok is blended across input, output and cache tokens. A headline input price is misleading when cache writes dominate, which they often do.

The Cache savings card shows how much of your context is served from the provider's prompt cache: hit rate (cache reads ÷ (cache reads + fresh input)), the cached-token count, and a tokens-weighted *≈N× cheaper than uncached* estimate — cache reads bill at roughly a tenth of fresh input, so N = (input + cache) ÷ (input + 0.1 × cache). It is labelled an estimate and uses no per-model price table; it is hidden until cache data exists.

Costs reported from your machine (the gateway's catalog estimates, and the old built-in agent's own figures) are cross-checked against a server-side rate card. Where they disagree, both are shown — the divergence is the signal. A model with no list price on file is marked unverified rather than being given a fabricated rate.

A weekly email flags specific, actionable findings — cache writes that never get read, spend concentrated on one model, prompt caching left unexploited, run-rate spikes. Each cites the numbers that triggered it. Findings are throttled per issue, so a standing one stays quiet for 28 days.

Harnesses — installing and updating

HarnessLink accompanies any of the supported coding-agent harnesses. Use superpi harness to see what is installed on the current machine and to install additional ones.

bash
superpi harness list           # table: id, binary state, real path, capture, install hint
superpi harness install <id>   # install a harness via bun install -g <pkg>
superpi harness install <id> --yes  # skip the confirmation prompt

Valid harness ids and their packages:

idbinarynpm package
ompomp@oh-my-pi/pi-coding-agent
claude-codeclaude@anthropic-ai/claude-code
codexcodex@openai/codex
geminigemini@google/gemini-cli
pipi@mariozechner/pi-coding-agent
dshdsh@deepseek-ai/dsh

Binary state column

  • installed vX.Y.Z — binary found, realpath is a genuine install.
  • shim (cmux) — which <binary> resolves to a PATH shim managed by

    cmux-cli (usually under the OS temp directory). The harness is not actually installed here; select install to install it properly.

  • missing — binary not found in PATH.

The update picker (superpi update) also shows shim and missing harnesses and lets you install them in one step.

Capture state is separate

The capture column shows whether HarnessLink's cloud sync is reading this harness's sessions, memory and prompt history. That is managed independently:

bash
superpi cloud harnesses              # list capture state
superpi cloud harnesses enable <id>  # turn on cloud sync for a harness
superpi cloud harnesses disable <id> # turn it off

Autobank (the memory bank of finished turns) is a separate switch per harness, because it works locally and does not depend on cloud sync:

bash
superpi cloud autobank               # status: per harness on/off, turns banked, today, last banked
superpi cloud autobank enable <id>   # turn it on; banks that harness's last 30 days right away
superpi cloud autobank disable <id>  # stop banking (existing entries stay)
superpi cloud autobank backfill [--days=N] [id…]   # re-scan the last N days (default 30); never banks a turn twice

Installing a harness with superpi harness install turns on neither. Run superpi cloud setup again, or the commands above. The HarnessLink service picks up a harness turned on after it started, without a restart.

Installing at setup time

Pass --with=<id,id,...> to the installer to install harnesses alongside omp:

bash
curl -fsSL https://harnesslink.sh/install.sh | bash -s -- --with=claude-code,codex

The installer verifies the registry's bin field before running bun install -g <pkg>, and --yes is implied by --with (consent is given by including the flag).

Command reference

bash
# Home and housekeeping
superpi                        # home screen: service, sign-in, sync, each harness's wiring, next step (first run: cloud setup; plain text when not a terminal)
superpi cleanup                # remove what the old built-in agent left (git worktrees in ~/.superpi/wt, background-job logs in ~/.superpi/run/daemons), with consent; never silently drops uncommitted work

# Setup and the HarnessLink service
superpi cloud setup            # the onboarding: plan → sign-in (required) → sync, service, every harness, restore, verify
superpi gateway start          # install the HarnessLink service (launchd / systemd --user / logon task) and start it; restarts an older build
superpi gateway service install|uninstall|restart|status   # manage that supervised service directly
superpi gateway status         # gateway health, harness wiring, today's requests and savings, cloud sync health
superpi gateway connect [id]   # route a harness's model calls through the gateway (all detected when omitted; starts the service)
superpi gateway disconnect [id]  # put back the harness's own provider settings
superpi gateway run [--port=<n>]  # the gateway alone, in this terminal (debugging)
superpi daemon install|uninstall  # same as gateway service install|uninstall
superpi daemon status         # is the service process running, and which harnesses it captures
superpi daemon start [-f]     # unsupervised background copy (no auto-restart, no OS limits); -f runs it in this terminal
superpi daemon stop           # stop the running service process

# Harnesses
superpi harness list           # binary state, path, capture, install hint
superpi harness install <id> [--yes]
superpi harness select [id] [--enable|--disable] [--hook] [--all]   # capture on/off and MCP registration, interactive without flags
superpi harness hook [id]      # register HarnessLink's MCP memory tools (every installed harness when omitted)
superpi harness context [id] [--remove]   # context hooks for Claude Code, Codex and omp
superpi cloud harnesses [enable|disable <id>]   # transcript and prompt-history capture per harness

# Cloud account and sync
superpi cloud login [--url=<api>] [--token]   # link this machine (turns cloud sync on); --token reads a dashboard-minted spk_ token from stdin
superpi cloud status           # sync on/off, every sync category on/off, token, API base, per-store pending / failed / last success
superpi cloud sync [--json] [--retry-dead]   # ingest harness prompt history, then upload everything pending now; --retry-dead requeues failed rows first
superpi cloud sync --disable <category>      # stop one category (history|sessions|memories|usage|gatewayUsage|profile); its queued rows are discarded
superpi cloud sync --enable <category>       # turn it back on (from now; nothing typed meanwhile is backfilled)
superpi cloud logout           # unlink and stop capturing
superpi cloud backfill [--force]

# Moving machines
superpi cloud push-profile [--dry-run]   # snapshot machine setup (HarnessLink's and every harness's); --dry-run prints it
superpi cloud restore [--device=<id>] [--apply] [--approve-mcp] [--install-tools] [--yes]
superpi cloud pull-sessions [--all]
superpi cloud pull-history
superpi cloud pull-memories [--all]
superpi cloud adopt [--from=<slug>]   # map another machine's folder to this one

# Memory
superpi cloud autobank [status|enable <id…>|disable <id…>|backfill [--days=N] [id…]] [--json]   # memory bank of every harness's finished turns
superpi cloud import-memories [--apply] [--gateway=<path>]   # MemPalace and ~/.hermes (setup runs it once)
superpi cloud recall <query> [--project=<slug>|--all]
superpi cloud brief [--budget=<chars>] [--recall=<query>] [--out=<file>]
superpi cloud bank --title=<t> [--tags=a,b] [--body=<text>]
superpi cloud link-memory --project=<slug>

# Sharing, runs, sessions
superpi cloud share | shared | pull-shared [--yes]   # org-shared artifacts (Team and Enterprise plans)
superpi cloud runs [--limit=<n>]                 # runs: phases, tokens, cost
superpi cloud run-token --run=<id> [--ttl=2h] [--scopes=sync:write] [--revoke]
superpi sessions list [--project=<dir>] [--limit=<n>] [--json]   # sessions from every harness on this machine (read-only)
superpi sessions search <query> [--project=<dir>] [--limit=<n>] [--json]
superpi resume                 # cross-harness session picker
superpi run --harness <claude-code|codex|omp|pi|antigravity> [--cwd <dir>] [--timeout <dur>] [--model <id>] [--json] "<task>"   # hand a task to another harness, headless (see Delegation)

# Updates, templates, MCP
superpi update [--check] [--force]   # harness picker: omp, HarnessLink itself, claude, codex, …; offers to clean up ~/.superpi/wt
superpi add <name> [--harness <id>] [--yes] [--force]   # install a registry skill into every enabled or detected harness's skill folder (agents, extensions: omp only)
superpi add --list             # what HarnessLink installed, per harness
superpi add qmd [--harness <id>] [--yes] [--force]   # qmd: local BM25 + vector search over your memories, as an MCP server in each harness (downloads ~2.3 GB of models)
superpi remove <name> [--yes] [--force]   # remove only what `superpi add` wrote
superpi registry search|list|sources|add-source   # browse registries, add a third-party source (`add` is an alias of `registry`)
superpi mcp [--cwd=<dir>] [--harness=<id>]   # what a harness runs: relay into the per-machine MCP hub
superpi mcp register [--harness=omp,claude-code,codex] [--yes]   # wire into harnesses
superpi mcp register --status # show registration status
superpi mcp print [--harness=<id>] [--format=json|toml]          # print the MCP entry for any client
superpi context-hook           # what a harness's context hook runs (reads hook JSON on stdin); installed by `superpi harness context`

# Other
superpi kb search|ingest|status …   # alias `knowledge`: a separate semantic knowledge base in your own Cloudflare account (Workers AI + Vectorize)
superpi telegram link|status|unlink # link a Telegram chat to your cloud account to ask about your memory and sessions from it
superpi completions bash|zsh|fish   # print a shell completion script

# Root flags
superpi --help | --version
superpi --smoke-test                # load every command and exercise the gateway and the MCP relay end to end
superpi --profile <name> <command>  # run against a named profile instead of the default

Commands that belonged to HarnessLink's old built-in agent are gone: launching a coding session, any first argument starting with - other than the root flags above (-p, --model …), and acp, agents, auth-broker, auth-gateway, bench, browser-relay, cleanse, commit, compress, config, dry-balance, eval, factory, gallery, gc, grep, grievances, install, integrate, join, marketplace, models, plugin, provider, q, read, say, search, setup, share, shell, ssh, stats, ste, tiny-models, token, ttsr, usage, worktree and wt. Running one prints a single line saying it was part of the old built-in agent and to use omp or your harness instead, and exits 2.

Memory inside every harness — the MCP server

Every harness registers the same command, superpi mcp. Wire it into a harness once and your memory is a tool the model can call directly — no manual superpi cloud recall round trips.

The hub. superpi mcp is a small relay. It forwards the harness's JSON-RPC 2.0 stdio traffic to the MCP hub, one per machine, inside the HarnessLink service. The hub:

  • serves HarnessLink's own tools (below) plus the MCP servers in HarnessLink's own

    ~/.superpi/agent/mcp.json, starting each of those servers once and sharing it across every session;

  • hides from a session the servers its harness already runs itself

    (--harness=<id> tells the hub which harness is asking; register writes it for you);

  • serves at most 64 sessions at a time; the 65th gets HarnessLink's own tools

    in-process, without the shared servers;

  • starts nothing while the containment watchdog is tripped (see

    The gateway).

When the service is not running, superpi mcp serves HarnessLink's own tools in-process, so memory tools keep working.

Tools

ToolDescription
memory_searchFull-text search over memory: this machine's from the local index (always, offline included), plus your other machines' from the cloud when it answers within 1.5 s; an entry held by both appears once. Default scope is this project plus its linked projects. Use scope: "all" only when the user asks to search globally or names something that may belong to another project.
memory_getPull one entry's full content by project + rel_path (values come from a memory_search hit): from this machine when the file is here (an entry banked moments ago included), otherwise from the cloud.
memory_bankRecord a durable memory for this project — a decision, an outcome, a fact. Syncs to the cloud and becomes searchable everywhere.
context_briefA budgeted brief of this project from memory (human notes, recent work, harness memory, optional recall from the local index). Use at the start of a session instead of replaying transcripts. Never waits on the network.
session_listSessions from every harness on this machine, newest first, with whether each can be resumed here.
session_getOne session by id: owning harness, transcript path, and the exact resume command.
cloud_syncUpload everything pending now (queued rows, finished sessions, memory files) and report what was pushed.
kb_search, kb_ingestA separate semantic knowledge base for documents and code, in your own Cloudflare account (CF_API_TOKEN / CLOUDFLARE_API_TOKEN and CF_ACCOUNT_ID). Without credentials it falls back to local embeddings. It is not your HarnessLink Cloud memory.
mcp_list_serversThe downstream MCP servers the hub proxies, their status and tools.
delegateHand a task to another harness on this machine (claude-code, codex, omp, pi, antigravity) and wait for its answer: {harness, task, cwd?, timeout_s?, model?} → {harness, status, exit_code, summary, transcript_ref, run_id}. Runs with that harness's default permissions, in a cwd under your project root. See Delegation.
delegate_statusA delegate run that outlived the call (status: "running"): waits up to ~25 s and returns its result; cancel: true stops it.
recall_memory, record_learning, search_knowledgeAliases of memory_search, memory_bank and kb_search.

The hub also lists the tools of the MCP servers it proxies.

Scope semantics for memory_search (the same for the local index and the cloud):

  • "linked" (default) — current project + projects listed in

    <memoriesRoot>/<project>/linked.txt. Stay here unless the user asks otherwise.

  • "project" — current project only.
  • "all" — no project filter; searches everything in the account. Use when

    the user says "all my memory", "across projects", or names something that clearly belongs elsewhere.

Signed out, memory_search still answers from this machine and says that other machines' memories need superpi cloud login.

Register in your harnesses

bash
superpi mcp register            # default: omp + any other detected harness binaries
superpi mcp register --yes      # non-interactive (skip confirmation prompts)
superpi mcp register --harness=omp,claude-code,codex
superpi mcp register --status   # show what is already registered

Config written per harness:

HarnessFileEntry
omp~/.omp/agent/mcp.jsonmcpServers.superpi = { type: "stdio", command: "superpi", args: ["mcp", "--harness", "omp"] }
claude-codeuser scopeclaude mcp add --scope user superpi -- superpi mcp --harness claude-code
codex~/.codex/config.toml[mcp_servers.superpi] block running superpi mcp --harness codex
cursor~/.cursor/mcp.jsonsame shape as omp, with --harness cursor

A preset is only shipped for a client whose config path and format we verified against that client's own documentation — today that is Cursor. Every other client still gets the canonical block: print it and paste it into that client's config under the key it documents (VS Code: servers, Zed: context_servers, OpenCode: mcp with command as one array).

bash
superpi mcp print                 # canonical `mcpServers` block (JSON)
superpi mcp print --format toml   # TOML form
superpi mcp print --harness=omp   # exactly what register writes for a known harness

The block is the only thing on stdout; the note naming where it goes is written to stderr, so superpi mcp print | jq and superpi mcp print > mcp.json stay clean. An unknown --harness id exits non-zero with the valid ids listed rather than falling back to the generic block.

In the onboarding

superpi cloud setup registers the memory tools for every detected omp, Claude Code and Codex as part of its automatic checklist ("Memory tools (MCP)" per harness; turn it off per harness under c customize). A failure shows the exact superpi mcp register --harness=<id> --yes command to retry it.

Context hooks — memory before the model thinks

MCP tools only run when the model decides to call them. Context hooks run first, on every session and prompt, so the model starts with your memory instead of having to go looking for it.

bash
superpi harness context              # install into every detected harness
superpi harness context codex        # one harness
superpi harness context --remove     # undo (only HarnessLink's entries are touched)
HarnessWhereMechanism
Claude Code~/.claude/settings.jsonSessionStart + UserPromptSubmit + PreToolUse + Stop + PreCompact command hooks
Codex~/.codex/hooks.jsonthe same minus Stop; Codex runs a new hook only after you trust it once in /hooks
omp~/.omp/agent/extensions/superpi-context.tsa managed extension on before_agent_start, tool_call, agent_end and session_before_compact

Each hook calls superpi context-hook, which injects:

  • On session start: your organization's policy instructions (the

    global text plus every scoped instruction whose conditions match), then the project's context brief (the same one superpi cloud brief prints), opened by the handoff of an unfinished task when there is one (see Session handoff). The whole injection stays under 9,500 characters for Claude Code and omp (Claude Code's cap is 10,000) and 7,000 for Codex (under its 2,500-token preview limit).

  • On each prompt: up to three bank entries whose title or tags share at

    least two distinctive words with the prompt, ranked by how well their full text matches it, each with the project and rel_path that memory_get takes. Words that appear in most of the project's entries (usually the project's own name) don't count. Nothing is injected when nothing matches.

It reads the local memory index and memory files only (about 75 ms per prompt including process start). It never calls the cloud, gives up after 1.5 seconds (10 seconds at a turn's end or before a compaction, where it reads the transcript), and exits 0 on any error, so a broken or missing HarnessLink can't block the harness. Existing hooks and settings are kept, running it again changes nothing, and a settings file that isn't valid JSON is left alone and reported instead of overwritten. superpi cloud setup installs them for every detected harness as its "Context hooks" step.

Org policy — enforced before every tool call

The same hooks enforce the rules your org sets on the dashboard's Policies page. Before each tool call the harness asks superpi context-hook, which reads the policy the HarnessLink service keeps in ~/.superpi/agent/cloud-policy.json (refreshed every minute while you're signed in, whatever your sync settings). It never calls the cloud on this path.

RuleClaude CodeCodexomp
deny (tool or command)blockedblockedblocked
promptyou're askedblocked (Codex hooks cannot ask)you're asked; blocked with no UI
allowoutranks a lower-precedence restriction; never auto-approvessamesame

Tool names on the dashboard use omp's vocabulary (bash, write, edit, read, web_search, task, …); Claude Code's Bash, Write, Edit, WebSearch, Task and Codex's Bash and apply_patch (counts as edit and write) map onto them, and MCP tools match by their full name. Command rules use * wildcards and match any part of a compound command (cd x && rm -rf y meets rm -rf *). A rule conditioned on a model applies when the harness's model is unknown, unless it is an allow.

If the policy can't be fetched, the last one keeps applying, however old it is. If the cached file is unreadable, every tool call is blocked until the service rewrites it. On a machine that has never fetched a policy, nothing is enforced. pi, Cursor and Antigravity have no HarnessLink tool hook, so rules are not enforced there, and Codex's hosted web search bypasses hooks. Plain superpi and superpi cloud status show the active rules, when they were last fetched, and where they are not enforced. A refusal names the rule and links the Policies page.

You don't need to reinstall after an update. When the HarnessLink service starts, it adds the hooks added since (policy, session handoff) wherever HarnessLink's context hooks are already installed, and refreshes its omp extension. It never adds hooks to a harness where you haven't installed them. Codex runs a new hook only after you trust it: until you open /hooks in Codex and trust HarnessLink's entries, Codex skips them, and plain superpi shows "run /hooks in Codex" as the next step.

Session handoff — a new session instead of compaction

A long session gets slower and dearer on every request, and compaction then squeezes it into a lossy summary. HarnessLink instead moves the task to a new session that starts from memory.

Checkpoint. For every session the autobank keeps a task checkpoint in state/<harness>--<session>.md under the project's memory, updated as the transcript grows: the goal (the session's first prompt, plus the latest request), the next step (the todo in progress, else the last assistant message), the latest tool calls, decisions (assistant and, where the transcript keeps it, reasoning sentences that state a choice), the files edited, and the todo list or plan. It is capped at 4,000 characters. Its status is open, done (every todo finished) or picked_up (a new session resumed it).

When to hand off. The hooks read the context size from the transcript's last usage record (Claude Code, omp, pi) or Codex's token_count, never from a character estimate; omp reports its own. The handoff point is context.handoffAt in ~/.superpi/agent/config.yml:

yaml
context:
  handoffAt: auto     # default: half the context window or 100k tokens, whichever is smaller
  # handoffAt: 80000  # a token count
  # handoffAt: "40%"  # a share of the window
  # handoffAt: off    # never steer, never block compaction

The window is the one the harness states (Codex, omp), else the model's family default (Claude 200k, or 1M once a session is past 200k).

MomentClaude CodeCodexomp
A turn ends past the handoff point, task openthe model is asked to save a handoff with memory_bank (tag handoff) and tell you to start a new session (/clear); asked again only after the context grows by another half of the handoff point, never while a Stop hook is already continuing the turn—a handoff is saved and a notice tells you to start a new session (/new)
Automatic compaction, context under 90% of the windowa handoff is saved and the compaction is skipped; the reason is shownthe same (continue: false)the same (cancelled); an overflow recovery is never cancelled
Automatic compaction at 90% or more, or /compacta handoff is saved, then it compactsthe samethe same

Each safe point also banks the checkpoint as a bank entry tagged handoff (one per session, rewritten in place), so the handoff syncs to your other machines. omp's extension API cannot open a new session from a hook, so HarnessLink only tells you to.

Resuming. When a session starts fresh (startup), after /clear, or after a compaction, the newest open checkpoint of the project (this session's own first, at most three days old) and the newest handoff note go first in the brief, under "Resume: unfinished task from an earlier session", capped at 5,000 characters and never dropped for budget. Both are then marked picked up, so the next session starts clean; a checkpoint reopens if its session does more work. Resuming a transcript (--resume) injects no handoff.

HarnessLink runs a local gateway on 127.0.0.1:4747. Once a harness is connected, its model calls go through HarnessLink first, then to the provider. Your harness keeps its own login: HarnessLink passes the credential through and never stores provider keys. A HarnessLink Cloud account is required — without a sign-in the gateway answers every call with "HarnessLink Cloud sign-in required".

bash
superpi gateway connect           # start the HarnessLink service if needed, then point every detected harness at it
superpi gateway connect omp       # or just one harness
superpi gateway start             # start the service on its own (now and at every login)
superpi gateway status            # service, sign-in, per-harness wiring, today's savings
superpi gateway disconnect        # restore direct provider access
HarnessWhat connect changes
Claude Codeenv.ANTHROPIC_BASE_URL in ~/.claude/settings.json (your claude.ai login keeps working)
Codextop-level openai_base_url in ~/.codex/config.toml (ChatGPT login or API key, detected from auth.json)
ompproviders.anthropic.baseUrl and providers.openai-codex.baseUrl in ~/.omp/agent/models.yml
pithe same two providers in ~/.pi/agent/models.json
Antigravitynot supported yet — signed-in Antigravity sends its traffic straight to Google

Restart a harness after connecting it. disconnect puts back exactly what was there before, including a base URL you had set yourself.

The service. superpi gateway connect and superpi gateway start register superpi daemon run with the platform's supervisor: a LaunchAgent on macOS, a systemd --user unit on Linux, a logon task on Windows. Where there is no supervisor (a container, WSL without systemd) it starts the daemon in the background and says it will not survive a reboot. connect needs a HarnessLink Cloud sign-in and waits until the gateway answers before it changes any harness; if the service cannot start, nothing is connected. The service (superpi daemon run) hosts the gateway, the MCP hub, the containment watchdog, gateway-usage sync, prompt-history ingest from every harness with capture on, cloud sync and the harness transcript watcher. When it is down, connected harnesses get "connection refused" — there is no silent bypass around HarnessLink. superpi daemon start runs the same loop without a supervisor: no restart after a crash or reboot and no OS resource limits — use it only where gateway start cannot install a service.

Limits. The daemon cannot run away, however many sessions use it. A watchdog inside it checks the daemon's child-process tree every 5 seconds; above 200 processes or 2048 MB of combined resident memory it kills that tree (never the daemon), refuses to start MCP servers for 5 minutes, and records the trip in ~/.superpi/agent/run/watchdog.json. superpi gateway status shows a red line for a trip in the last 24 hours, and /_superpi/health reports the watchdog. The OS adds a backstop: the systemd unit sets TasksMax=512, MemoryHigh=2G and MemoryMax=3G (TasksMax counts threads — the daemon alone runs about 40 and each node-based MCP server about 12, twice that under npx, so 256 would cap you at about nine MCP servers). The LaunchAgent sets NumberOfProcesses to 4096: macOS applies that limit to every process the user owns, not just the daemon's tree, and a desktop session already runs about 500, so a tree-sized value would make every process start in the daemon fail. Windows has no OS limit (Bun cannot place the daemon in a job object); the watchdog is the bound there.

Disk. The service writes its output to ~/.superpi/logs/daemon.log. Past 5 MB (checked at start and hourly) it is copied to daemon.log.1, replacing the previous copy, and emptied in place, so the log never holds more than about 10 MB. gateway.db gives back the space of cache entries it drops: once more than 8 MB is free it shrinks the file. A gateway.db created before this was in place is rewritten once by the service when it holds more than 32 MB of free space; the service log records the size before and after, and gateway requests wait for it (about 0.7 s for a 550 MB file).

Token savings. Two mechanisms, both reported per harness on the dashboard's Usage page under Gateway savings:

  • Exact-match cache. An identical model request replays the stored

    response without calling the provider (response header x-superpi-gateway: hit). Per-request identifiers that harnesses stamp on every call (Codex item ids and turn metadata, the Claude-Code billing hash, metadata, prompt_cache_key) are ignored when matching. Streams over WebSocket (Codex's default transport) and requests chained with previous_response_id are never cached, because they depend on server-side state. Entries live 24 hours in gateway.db, capped at 128 MB (gateway.cache.ttlHours, gateway.cache.maxMB, gateway.cache.enabled). Send x-superpi-cache: bypass to skip it for one call.

  • Provider prompt caching. Most tokens in an agent loop are the same

    conversation prefix re-sent every turn; providers bill those at a large discount when they are cache reads. The gateway records every cache read, and adds Anthropic cache breakpoints to requests that set none.

Costs are estimates from the bundled model catalog; a model the catalog does not price shows tokens saved without a dollar figure.

Flag notifications