---
name: pp-exa
description: "Every Exa search API feature, plus local history, spend tracking, and diffing no other Exa tool has. Trigger phrases: `search the web for`, `get the latest news on`, `answer this question with sources`, `find similar pages`, `monitor this topic`, `exa search`, `run exa`."
author: "Som Samantray"
license: "Apache-2.0"
argument-hint: "<command> [args] | install cli|mcp"
allowed-tools: "Read Bash"
metadata:
  openclaw:
    requires:
      bins:
        - exa-pp-cli
    install:
      - kind: go
        bins: [exa-pp-cli]
        module: github.com/mvanhorn/printing-press-library/library/ai/exa/cmd/exa-pp-cli
---

# Exa — Printing Press CLI

## Prerequisites: Install the CLI

This skill drives the `exa-pp-cli` binary. **You must verify the CLI is installed before invoking any command from this skill.** If it is missing, install it first:

1. Install via the Printing Press installer. It defaults binaries to `$HOME/.local/bin` on macOS/Linux and `%LOCALAPPDATA%\Programs\PrintingPress\bin` on Windows:
   ```bash
   npx -y @mvanhorn/printing-press-library install exa --cli-only
   ```
2. Verify: `exa-pp-cli --version`
3. Ensure the reported install directory is on `$PATH` for the agent/runtime that will invoke this skill.

If the `npx` install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.6 or newer). This installs into `$GOPATH/bin` (default `$HOME/go/bin`), so add that directory to `$PATH` instead:

```bash
go install github.com/mvanhorn/printing-press-library/library/ai/exa/cmd/exa-pp-cli@latest
```

If `--version` reports "command not found" after install, the runtime cannot see the binary directory on `$PATH`. Do not proceed with skill commands until verification succeeds.

exa-pp-cli wraps the full Exa API — search, contents, answer, find-similar, monitors, agent runs, websets, webhooks, imports — and adds a local SQLite store so every result, run, and cost lands on disk. Re-run past queries, diff monitor runs, track entity first-seens, and watch spend in real time.

## When to Use This CLI

Use exa-pp-cli whenever an agent needs live web research with grounded citations, clean page content, scheduled monitoring, or entity tracking — and wants every search persisted locally for later querying.

## Anti-triggers

Do not use this CLI for:
- Do not use exa-pp-cli for general web crawling at scale; use a dedicated crawler.
- Do not use exa-pp-cli to bypass paywalls or scrape restricted content.
- Do not use exa-pp-cli for real-time monitoring dashboards at sub-second cadence; use the API directly.

## Unique Capabilities

These capabilities aren't available in any other tool for this API.

### Local state that compounds
- **`spend`** — See cumulative API spend across every Exa call, broken down by day and resource.

  _Agents should reach for this when they need to know how much Exa usage costs before running another batch._

  ```bash
  exa-pp-cli spend --days 30 --resource searches
  ```
- **`monitor diff`** — Compare two synced monitor runs and see exactly which URLs are new, gone, or unchanged.

  _Reach for this to see what a scheduled search found since the last run instead of re-reading entire run outputs._

  ```bash
  exa-pp-cli monitor diff <monitor-id>
  ```
- **`entity report`** — Build a first-seen / last-seen / mention-count timeline for any company or person across your synced searches and webset items.

  _Reach for this to track when an entity first appeared and how often it shows up across all your research surfaces._

  ```bash
  exa-pp-cli entity report "Acme Corp" --type company --since 30d
  ```
- **`webset new`** — List items added to a live webset since your last sync, so you only see what changed.

  _Reach for this for the weekly what's-new sweep over a curated set instead of re-listing every item._

  ```bash
  exa-pp-cli webset new <webset-id> --since 7d
  ```

## Command Reference

**agent** — Manage agent

- `exa-pp-cli agent cancel-run` — Cancel a queued or running Agent run.
- `exa-pp-cli agent create-run` — Create an asynchronous Agent run. By default, the API returns the run object immediately.
- `exa-pp-cli agent delete-run` — Delete a stored Agent run.
- `exa-pp-cli agent get-run` — Retrieve a single Agent run by ID.
- `exa-pp-cli agent list-run-events` — List stored events for an Agent run. Set `Accept: text/event-stream` to replay stored events as server-sent events.
- `exa-pp-cli agent list-runs` — List Agent runs for your team, ordered from newest to oldest.

**answer** — Manage answer

- `exa-pp-cli answer` — Performs a search based on the query and generates either a direct answer or a detailed summary with citations

**contents** — Manage contents

- `exa-pp-cli contents` — Contents

**events** — Manage events

- `exa-pp-cli events get` — Get a single Event by id. You can subscribe to Events by creating a Webhook.
- `exa-pp-cli events list` — List all events that have occurred in the system. You can paginate through the results using the `cursor` parameter.

**find-similar** — Manage find similar

- `exa-pp-cli find-similar` — Find links similar to the provided URL and optionally retrieve their contents.

**imports** — Manage imports

- `exa-pp-cli imports create` — Creates a new import to upload your data into Websets.
- `exa-pp-cli imports delete` — Deletes a import.
- `exa-pp-cli imports get` — Gets a specific import.
- `exa-pp-cli imports list` — Lists all imports for the Webset.
- `exa-pp-cli imports update` — Updates a import configuration.

**monitors** — Manage monitors

- `exa-pp-cli monitors batch` — Perform a batch action on monitors matching the provided filters.
- `exa-pp-cli monitors create` — Creates a new Monitor to run recurring Exa searches on a schedule.
- `exa-pp-cli monitors create-endpoint` — Creates a new `Monitor` to continuously keep your Websets updated with fresh data.
- `exa-pp-cli monitors delete` — Deletes a monitor. This cannot be undone.
- `exa-pp-cli monitors delete-id` — Deletes a monitor.
- `exa-pp-cli monitors get` — Retrieves a single monitor by its ID.
- `exa-pp-cli monitors get-id` — Gets a specific monitor.
- `exa-pp-cli monitors list` — Lists all monitors for the authenticated team. Supports filtering by status and cursor-based pagination.
- `exa-pp-cli monitors list-endpoint` — Lists all monitors for the Webset.
- `exa-pp-cli monitors update` — Updates an existing monitor. All fields are optional.
- `exa-pp-cli monitors update-id` — Updates a monitor configuration.

**teams** — Manage teams

- `exa-pp-cli teams` — Returns information about the authenticated team, including current concurrency usage and limits.

**webhooks** — Manage webhooks

- `exa-pp-cli webhooks create` — Create a Webhook
- `exa-pp-cli webhooks delete` — Delete a Webhook
- `exa-pp-cli webhooks get` — Get a Webhook
- `exa-pp-cli webhooks list` — List webhooks
- `exa-pp-cli webhooks update` — Update a Webhook

**websearch** — Manage websearch

- `exa-pp-cli websearch` — Perform a search with an Exa prompt-engineered query and retrieve a list of relevant results. Optionally get contents.

**websets** — Manage websets

- `exa-pp-cli websets create` — Creates a new Webset with optional search, import, and enrichment configurations.
- `exa-pp-cli websets delete` — Deletes a Webset. Once deleted, the Webset and all its Items will no longer be available.
- `exa-pp-cli websets get` — Get a Webset
- `exa-pp-cli websets list` — Returns a list of Websets. You can paginate through the results using the `cursor` parameter.
- `exa-pp-cli websets preview` — Preview how a search query will be decomposed before creating a webset.
- `exa-pp-cli websets update` — Update a Webset


### Finding the right command

When you know what you want to do but not which command does it, ask the CLI directly:

```bash
exa-pp-cli which "<capability in your own words>"
```

`which` resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code `0` means at least one match; exit code `2` means no confident match — fall back to `--help` or use a narrower query.

## Recipes

### Deep research with structured output

```bash
exa-pp-cli websearch --query "compare the latest frontier AI model releases" --type deep --output-schema '{"type":"object","required":["models"],"properties":{"models":{"type":"array","items":{"type":"object"}}}}'
```

Run deep search and get a synthesized, schema-shaped answer with grounding.

### Monitor a competitor's news

```bash
exa-pp-cli monitors create --name "competitor-news" --search-query "new product launches" --search-contents-highlights true
```

Schedule a recurring search, then run `sync` to persist runs locally.

### What changed since last monitor run

```bash
exa-pp-cli monitor diff <monitor-id>
```

Schedule a recurring search, then run `sync` to persist runs locally.

### Track an entity over time

```bash
exa-pp-cli entity report "Acme Corp" --since 30d --agent --select entity,mentionCount
```

Get an agent-shaped first-seen/last-seen timeline for a company across all synced research.

### Narrow a search response with --select

```bash
exa-pp-cli websearch --query "AI regulation policy updates" --category news --num-results 5 --select results.title,results.url
```

Request only the high-gravity fields so agent context is not flooded with full result payloads.

## Auth Setup

Run `exa-pp-cli auth setup` for the URL and steps to obtain a token (add `--launch` to open the URL). Then store it:

```bash
exa-pp-cli auth set-token YOUR_TOKEN_HERE
```

Or set `EXA_API_KEY` as an environment variable.

Run `exa-pp-cli doctor` to verify setup.

## Agent Mode

Add `--agent` to any command. Expands to: `--json --compact --no-input --no-color --yes`.

- **Pipeable** — JSON on stdout, errors on stderr
- **Filterable** — `--select` keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:

  ```bash
  exa-pp-cli contents --agent --select id,name,status
  ```
- **Previewable** — `--dry-run` shows the request without sending
- **Offline-friendly** — sync/search commands can use the local SQLite store when available
- **Non-interactive** — never prompts, every input is a flag
- **Explicit retries** — use `--idempotent` only when an already-existing create should count as success, and use `--ignore-missing` only when a missing delete target should count as success

### Response envelope

Commands that read from the local store or the API wrap output in a provenance envelope:

```json
{
  "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."},
  "results": <data>
}
```

Parse `.results` for data and `.meta.source` to know whether it's live or local. A human-readable `N results (live)` summary is printed to stderr only when stdout is a terminal AND no machine-format flag (`--json`, `--csv`, `--compact`, `--quiet`, `--plain`, `--select`) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout.

## Paths and state

Agents should treat the CLI's path resolver as part of the runtime contract:

- Use `--home <dir>` for one invocation, or set `EXA_HOME=<dir>` to relocate all four path kinds under one root.
- Use per-kind env vars only when a specific kind must diverge: `EXA_CONFIG_DIR`, `EXA_DATA_DIR`, `EXA_STATE_DIR`, `EXA_CACHE_DIR`.
- Resolution order is per-kind env var, `--home`, `EXA_HOME`, XDG (`XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`), then platform defaults.
- `config` contains settings like `config.toml` and profiles. `data` contains `credentials.toml`, `data.db`, and auth sidecars. `state` contains persisted queries, jobs, and `teach.log`. `cache` contains regenerable HTTP/cache files.
- Stored secrets live in `credentials.toml` under the data dir. Existing legacy `config.toml` secrets are read for compatibility and leave `config.toml` on the first auth write.
- Run `exa-pp-cli doctor --fail-on warn` to surface path and credential-location warnings. `agent-context` exposes a schema v4 `paths` block for agents that need the resolved dirs.
- For MCP, pass relocation through the MCP host config. The MCP binary does not inherit CLI flags:

  ```json
  {
    "mcpServers": {
      "exa": {
        "command": "exa-pp-mcp",
        "env": {
          "EXA_HOME": "/srv/exa"
        }
      }
    }
  }
  ```

Fleet precedence: an inherited per-kind env var overrides an explicit `--home` for that kind. Use `EXA_HOME` or per-kind vars as durable fleet levers, and use `--home` only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing `EXA_HOME`, or `doctor` will not find credentials left under the former root.

## Automatic learning

This CLI ships a self-capturing learning loop. The CLI does its own bookkeeping: every invocation is journaled locally, a failed flag followed by a corrected retry auto-derives a `flag_alias` candidate, and a `teach` on a query family without a playbook auto-synthesizes a `playbook_candidate` from the session's journal. Your job is judgment only: `recall` first, act on surfaced candidates, `teach` the final answer, `playbook amend` when you observe a correction. You never record failures by hand.

### Step 1: `recall` before any discovery

Before list/search/drill commands on a new user question, run:

```bash
exa-pp-cli recall "<user's question>" --agent
```

The response envelope:

```json
{
  "query": "...",
  "normalized": "<normalized form>",
  "query_entities": ["..."],
  "found": true | false,
  "match_score": 0.0,
  "results": [
    { "resource_id": "...", "resource_type": "...", "venue": "...",
      "confidence": 2, "entity_match": "exact|partial|unknown",
      "source": "taught|preseed|pattern", "warnings": ["..."] }
  ],
  "mismatches": [ /* only when --debug-mismatches */ ],
  "warnings": [ /* top-level */ ],
  "candidates": [
    { "id": 12, "class": "flag_alias | playbook_candidate",
      "summary": "...", "sightings": 3, "last_seen": "...",
      "rationale": "...",
      "next_action": ["<trial command>", "exa-pp-cli learnings confirm 12"] }
  ],
  "playbook": {
    "query_family": "...",
    "playbook": {
      "steps": [ { "cmd": "<command with {slot} substitution>", "purpose": "..." } ],
      "entity_slots": ["$ENTITY"],
      "expected_tool_calls": 3
    },
    "slots_resolved": { "$ENTITY": { "token": "<live token>", "canonical": "<canonical>" } },
    "notes": "<workarounds + gotchas for this query family>"
  },
  "notes": "<duplicate surface for non-playbook callers>"
}
```

Empty-store short-circuit: if the store has no learnings, playbooks, or candidates yet (recall finds nothing and `learnings list` and `learnings candidates` are both empty), skip recall for the rest of this session instead of taxing every query; resume recall-first once something has been taught.

### Step 2: decision tree

Read `candidates`, `playbook`, `notes`, `results[0]`, and warnings in that order:

```
if Candidates present (warnings include "candidates_present"):
    -> candidates are try-then-confirm, never facts. Follow each candidate's
       two-step next_action verbatim: run the trial command first, then run
       `learnings confirm <id>` only after the trial verified the behavior.
       Reject a wrong candidate with `learnings reject <id>`.
    -> NEVER re-teach something recall surfaced as a candidate; confirm or
       reject that candidate instead of teaching a duplicate.
    -> candidates ride alongside playbooks and resource hits, not instead of
       them; continue with the branches below after acting on them.

if Playbook present:
    -> READ Playbook.notes verbatim FIRST (workarounds + gotchas the CLI surface doesn't expose)
    -> replay Playbook.steps in order, substituting Playbook.slots_resolved entries
       for the entity slot tokens. If a step's slot is unresolved, fall back to
       discovery for that step only.
    -> the Playbook's expected_tool_calls is a budget; if you find yourself running
       materially more, record the divergence via `exa-pp-cli playbook amend`
       at end-of-session.

elif Notes present (no Playbook):
    -> read Notes verbatim before any discovery step; they carry known gotchas
       for this query family even when no structured choreography exists yet.

elif Found AND Results[0].EntityMatch == "exact" AND Results[0].Confidence >= 2:
    -> skip discovery; fetch live data for Results[*].ResourceID in parallel

elif Found AND Results[0].EntityMatch == "partial":
    -> candidate hint, NOT a hit; read the resource title to validate before trusting

elif (any row in Mismatches[] when --debug-mismatches was passed):
    -> treat as cold start; the stored learning is for a different entity
       (different canonical resolved from query_entities)

else:  // Found == false, no playbook, no notes
    -> cold start; run discovery normally; teach the answer afterward (Step 4).
       If the family has no playbook yet, that teach auto-synthesizes a
       playbook candidate from this session's journal - you do not need to
       record one by hand.
```

Playbook and Notes are orthogonal to the per-resource path. A recall response can carry both a Playbook AND a `Results[]` hit - use both: the Playbook tells you which choreography to run; the resource hits short-circuit specific steps. Default to skipping `mismatches`; pass `--debug-mismatches` only when investigating cold-start surprises.

Candidate judgment details: `learnings confirm <id>` prints the candidate's full payload before materializing it - check that the printed payload matches the behavior you verified. `learnings reject <id>` tombstones the derivation signature so the same candidate does not resurface. The envelope carries only the few candidates worth acting on now; `exa-pp-cli learnings candidates` lists the full open set.

Graceful degradation: if `learnings confirm` is an unknown command, you are driving an older binary - ignore the candidates guidance and follow the rest of the protocol.

### Step 3: always read `warnings`

- `low_confidence`: row exists at `confidence<2`. Treat as a hint, not a skip-discovery hit.
- `resource_not_in_store`: the local store doesn't have the resource the learning points at. The match validator couldn't classify entities — direct-fetch and re-evaluate.
- `cross_alias_match` (per-result): the row was taught under a different alias and matched the live query's canonical via `entity_lookups` (e.g., a "USA" teach satisfying a "United States" recall). Trust the resource_id.
- `similar_shape_different_entity:<canonical>` (top-level): a structurally matching row exists but its canonical entity differs from the live query's. Treated as cold start; the warning carries the conflicting canonical as a hint, but the row is NOT promoted into Results.
- `ambiguous_alias` (top-level): a single query entity resolved to multiple canonicals (e.g., "Cards" → Arizona Cardinals + St. Louis Cardinals). Surface the ambiguity from context before committing to a resource.
- `candidates_present` (top-level): the envelope carries a `candidates` section. Handle it via the candidates branch in Step 2 before anything else.
- `lookup_refresh_available` (top-level): an entity in the query has no lookup row yet, but synced data could provide one. Run `exa-pp-cli sync` to refresh entity lookups.
- Top-level `no_learnings_for_query_family`: the table had no rows above the Jaccard floor. Pure cold start.

### Step 4: `teach &` after finalizing your response - always

Teaching is unconditional. After resolving a query the store could not answer, background-teach the final resource mapping - no call-count threshold, no judging whether it was "worth" learning. The teach is the anchor of the loop: it triggers playbook synthesis for a family without a playbook, and same-referent phrasings fold int