Hindsight + OpenClaw: A Working Setup Guide

How we run Hindsight as the durable memory layer for OpenClaw — transport, tagging, provenance, and the failure modes that cost us real time.

Verified against a live deployment on 2026-10-07. Plugin @vectorize-io/hindsight-openclaw 0.13.0 · Hindsight server API 0.10.2


1. The shape of it

Two pieces, and keeping them separate is the whole trick:

The plugin talks to the server over HTTP. Point hindsightApiUrl at the root of the server — not /v1/....

Install

# 1. Install the plugin (verify the install actually succeeded before touching config)
openclaw plugins install @vectorize-io/hindsight-openclaw

# 2. Run the interactive setup wizard
npx --package @vectorize-io/hindsight-openclaw hindsight-openclaw-setup

# 3. Validate and start
openclaw config validate
openclaw gateway

The wizard offers three modes:

The wizard stores credentials inline in openclaw.json for convenience. For anything production-shaped, use a SecretRef instead:

openclaw config set plugins.entries.hindsight-openclaw.config.hindsightApiToken \
    --ref-source env --ref-id HINDSIGHT_API_TOKEN

2. What we are actually running

Thing Value
Plugin @vectorize-io/hindsight-openclaw 0.13.0
Server API 0.10.2
Server features on observations, mcp, worker, bank_config_api, file_upload_api
Bank layout one shared default bank, dynamicBankId: false
Fact units in bank ~51,500
Recall / retain retain on every turn; auto-recall injection off, knowledge tools on
Embeddings omniroute provider, openrouter/qwen/qwen3-embedding-4b (2560-dim)

3. The plugin config

Lives at plugins.entries.hindsight-openclaw.config in ~/.openclaw/openclaw.json.

{
  "enabled": true,
  "hooks": { "allowConversationAccess": true },
  "config": {
    "hindsightApiUrl": "http://YOUR-HINDSIGHT-HOST:8889",
    "bankId": "default",
    "dynamicBankId": false,
    "enableKnowledgeTools": true,

    "autoRetain": true,
    "retainRoles": ["user", "assistant"],
    "retainFormat": "json",
    "retainToolCalls": true,
    "retainEveryNTurns": 1,
    "retainOverlapTurns": 0,
    "excludeProviders": ["heartbeat"],
    "skipStatelessSessions": true,

    "retainTags": [
      "origin_scope:session",
      "source_system:openclaw",
      "host:primary",
      "source:chat"
    ],
    "retainSource": "openclaw",
    "retainContext": "This content is an AI-assistant conversation transcript from OpenClaw. Retain request metadata may include routing identifiers such as 'sender_id' (an opaque user ID, not a human name), 'channel_id' (a chat identifier), and 'provider' (the messaging platform name). These are operational routing metadata, not semantic actors or people. Messages with role 'assistant' are from the AI assistant; first-person statements in assistant messages refer to the AI, not the human user. Messages with role 'user' are from the human user. Bank IDs, session keys, agent IDs, thread IDs, source systems, and tags in metadata are also operational routing identifiers, not human names, project names, or organizations. Starting directory, host and source tags are provenance only. The bank classifies semantic project and scope per fact. Preserve original message dates.",

    "autoRecall": false,
    "recallBudget": "mid",
    "recallMaxTokens": 1024,
    "recallTypes": ["observation"],
    "preferObservations": false,
    "recallRoles": ["user", "assistant"],
    "recallContextTurns": 1,
    "recallMaxQueryChars": 800,
    "recallTimeoutMs": 10000,
    "recallInjectionPosition": "user",
    "recallPromptPreamble": "Relevant memories from past conversations (prioritize recent when conflicting). Only use memories that are directly useful to continue this conversation; ignore the rest:",

    "dynamicBankGranularity": ["agent", "channel", "user"],
    "retainQueueMaxAgeMs": -1,
    "retainQueueFlushIntervalMs": 60000,
    "logSummaryIntervalMs": 300000,
    "logLevel": "info",
    "debug": false
  }
}

4. The five decisions that actually matter

4.1 autoRecall: false + knowledge tools on

This is our biggest deviation from defaults. By default the plugin silently injects recalled snippets into the prompt every turn. We turned automatic injection off and use the explicit tools instead:

Why: silent injection burns tokens on irrelevant hits and muddies the transcript. Explicit tools let the agent decide when to reach for memory.

If you want the "it just magically remembers" feel, leave autoRecall on — but tune recallBudget and recallMaxTokens down or it gets chatty.

4.2 Recall observations, not raw messages

recallTypes: ["observation"] means the server's consolidation layer hands back synthesized observations instead of raw transcript chunks. That is the difference between "we talked about X once" and "here is what we concluded about X."

Also set a finite reflect_source_facts_max_tokens at the bank level (4096–8192). Never leave reflection source facts unbounded.

4.3 One shared bank, tagged by scope

dynamicBankId: false keeps everything in default; project and scope get classified semantically per fact.

The alternative — a bank per agent/channel/user — gives hard isolation, but every new bank cold-starts with server defaults and you end up stamping the same contract N times.

Our call: tags route; banks isolate. Flip to per-bank only when you genuinely need data isolation (different people, different tenants), not just organization.

4.4 Keep noise out at the door

Cheap, and it makes everything downstream better.

4.5 The provenance paragraph in retainContext earns its keep

Without it, extraction reads host, channel_id, session_key, bank_id as if they were real entities, and reads assistant first-person text as if the user said it.

That paragraph tells the extractor which tokens are operational metadata versus real facts, and instructs the bank to classify project/scope semantically. If your entity graph is polluted with fake entities named primary, a model name, or a UUID — this paragraph is the fix.


5. How tagging actually works

Tags reach a document through three separate mechanisms. The plugin only owns one of them. Mixing these up causes most confusion.

5.1 Plugin-stamped tags on auto-retain

Every document the plugin retains gets retainTags stamped automatically:

origin_scope:session   source_system:openclaw   host:primary   source:chat

Flat key:value strings. host: is per-host — yours will differ.

5.2 Tagging manual ingests — you must PATCH after ingest

agent_knowledge_ingest accepts only title and content — there is no tags parameter. (Verified against the live tool schema.)

Tags embedded as YAML frontmatter in the content are NOT auto-extracted. Probed live 2026-10-07 on Hindsight API 0.10.2: a document ingested with a frontmatter tags: block came back tags: [] from the API, while every plugin-retained conversation still carried its retainTags. Frontmatter prose is useful for a human reader, but the server does not lift it into document tags.

The working path for a hand-written page is ingest, then PATCH:

# 1. Ingest (title becomes the document ID)
#    agent_knowledge_ingest { title, content }

# 2. Apply tags via the API
curl -X PATCH \
  "$HINDSIGHT/v1/default/banks/default/documents/$DOC_ID" \
  -H 'Content-Type: application/json' \
  -d '{"tags":["scope:global","source:live-runtime","domain:openclaw",
               "knowledge:procedure","host:primary","project:openclaw"]}'

# 3. Verify — the PATCH body is NOT authoritative (see 5.3)
curl "$HINDSIGHT/v1/default/banks/default/documents/$DOC_ID"

The title becomes the document ID, so re-ingesting with the same title replaces the document. That is your update path.

5.3 Direct API for backfill and repair

PATCH /v1/default/banks/{bank_id}/documents/{document_id}

This is how we retro-tagged an existing corpus.

Gotcha: PATCH normalizes asynchronously. The PATCH response body is not authoritative — the server rewrites the tag set after writeback and can drop tags. Always confirm with a follow-up GET.


6. The tag taxonomy we run

Namespace-first, key:value, lowercase, colon separators. Counts are from a real 1,285-document corpus.

Namespace Values Meaning
source: chat · session-reflection · sleep_consolidation · mcp · upload · external · correction · live-runtime · preference how the fact entered the bank
source_system: openclaw · codex · repository-inspection which tool produced it
scope: global · session · repository · unclassified the primary routing key
domain: hindsight · security · storage · infrastructure · operator subject area
knowledge: decision · failure · incident · component · recovery · feature-work fact kind
host: primary · dev provenance (machine)
harness: codex runtime that produced it
project: openclaw · omniroute · repo slugs repository binding
project_scope: none explicitly not repo-specific
lifecycle: time-bounded · superseded · current staleness control
freshness: verify-at-source must re-check upstream before trusting
authority: nist · openzfs · truenas · cis · cisa external issuing body
platform: openzfs · truenas-scale · docker tech platform
document_type: guidance · reference · benchmark · framework · standard · tag-registry document shape
correction: operator-confirmed a human explicitly corrected this

The two that carry the most weight


7. Rules that came from real screwups

Classifier/LLM-output failures are transport errors, not classifications

Malformed or failed model output → skip the document, record the error, retry later. Never convert a failure into a fallback tag.

The unclassified tag is only for documents the classifier genuinely evaluated and could not classify. The moment a failure silently becomes "unclassified," you have poisoned your routing key with garbage that looks legitimate.

A script must never flip its own dry-run/apply gate

Our tag-repair script stamped wrong tags because it self-set applyApproved: true on a clean dry-run, and a later run inherited it.

The apply gate lives in deployment source (APPLY_ENABLED constant), not runtime state. Same principle for any mutating automation.

Batch + cursor + recorded skips — not a big-bang rewrite

The pattern that made retro-tagging tractable: a batch LLM sweep that reviews, compresses, and commits in one pass — 10 documents at a time, with a cursor file, resumable, and a permanent-skip list carrying recorded reasons.

Our 1,288-document retro-tag ended at exactly 1 untagged (CONTRIBUTING.md — ambiguous provenance, deliberately skipped with the reason logged). That is the shape to copy.


8. The other customization surface


9. How to actually use the wisdom

The loop we teach the agent:

  1. Recall before writing — agent_knowledge_recall
  2. Ingest raw — agent_knowledge_ingest (full content, never pre-summarized)
  3. Verify — agent_knowledge_get_page or recall again

agent_knowledge_reflect is for a synthesized answer rather than a document dump. It defaults to budget: "low", max_tokens: 1024, fact types world/experience/observation.


10. Gotchas checklist


11. API quick reference

Method Path
GET /health, /version, /openapi.json
POST /v1/default/banks/{bank}/memories — retain
POST /v1/default/banks/{bank}/memories/recall
GET /v1/default/banks/{bank}/memories/list
POST /v1/default/banks/{bank}/memories/dry-run-extract
GET/PATCH/DELETE /v1/default/banks/{bank}/documents/{id}
POST /v1/default/banks/{bank}/documents/{id}/reprocess
GET/PATCH /v1/default/banks/{bank}/config
GET /v1/default/banks

Package assembled from a live Hindsight + OpenClaw deployment. All versions, paths, and config values were read from the running system, not from memory.