---
name: pp-sarvam
description: "Every Sarvam AI model on your terminal — translate, speak, transcribe, chat, and extract documents in 22 Indian languages, with a local history that compounds. Trigger phrases: `translate this to Hindi`, `generate speech audio in Tamil`, `transcribe this audio file`, `what language is this text`, `extract fields from this document`, `use sarvam`, `run sarvam`."
author: "Som Samantray"
license: "Apache-2.0"
argument-hint: "<command> [args] | install cli|mcp"
allowed-tools: "Read Bash"
metadata:
  openclaw:
    requires:
      bins:
        - sarvam-pp-cli
    install:
      - kind: go
        bins: [sarvam-pp-cli]
        module: github.com/mvanhorn/printing-press-library/library/ai/sarvam/cmd/sarvam-pp-cli
---

# Sarvam AI — Printing Press CLI

## Prerequisites: Install the CLI

This skill drives the `sarvam-pp-cli` binary. **You must verify the CLI is installed before invoking any command from this skill.** If it is missing, install it first:

1. Install via the Printing Press installer. It defaults binaries to `$HOME/.local/bin` on macOS/Linux and `%LOCALAPPDATA%\Programs\PrintingPress\bin` on Windows:
   ```bash
   npx -y @mvanhorn/printing-press-library install sarvam --cli-only
   ```
2. Verify: `sarvam-pp-cli --version`
3. Ensure the reported install directory is on `$PATH` for the agent/runtime that will invoke this skill.

If the `npx` install fails before this CLI has a public-library category, install Node or use the category-specific Go fallback after publish.

If `--version` reports "command not found" after install, the runtime cannot see the binary directory on `$PATH`. Do not proceed with skill commands until verification succeeds.

Sarvam AI's official SDKs and MCP server are great for code, but nothing offers offline capability: no local history of translations, TTS generations, transcriptions, or chat threads. sarvam-pp-cli adds a local SQLite store, voice auditioning, conversation resume, batch job retry/report, pronunciation spot-checks, subtitle export, and a doc-ai extraction schema library — all with --json, --dry-run, and typed exit codes for agents and scripts.

## When to Use This CLI

Use sarvam-pp-cli when you need to translate, transliterate, transcribe, synthesize speech, chat, or extract document fields through Sarvam AI's Indic-language models — especially when you want a record of what you did, offline search across past work, batch orchestration, or scriptable/agent-friendly output. It is the right tool for voice-agent prompt engineering, call-center transcription QA, content localization, and document-intelligence pipelines.

## Anti-triggers

Do not use this CLI for:
- Do not use this CLI for the Voice Agents platform (apps.sarvam.ai) — deployments, campaigns, cohorts, analytics, and instant outbound use a separate X-API-Key auth system with org/workspace scoping and are not implemented here.
- Do not use this CLI for real-time WebSocket streaming (realtime STT, TTS WS) — those are AsyncAPI surfaces not covered by the REST CLI.
- Do not use this CLI to replace a full speech-recognition model benchmark — the API is a paid service and the CLI does not add model-evaluation tooling.

## Unique Capabilities

These capabilities aren't available in any other tool for this API.

### Speech workflow
- **`voices preview`** — Generate one sample sentence across every TTS speaker and hear them all side by side

  _Use when choosing a TTS voice for a new language or prompt set without burning manual API calls_

  ```bash
  sarvam-pp-cli voices preview --lang hi-IN --sample "नमस्ते, स्वागत है" --speakers shubh,ritu,priya --output ./voices
  ```
- **`pron-check`** — Verify a term's TTS pronunciation via a speech round-trip (TTS then STT)

  _Use to confirm a pronunciation dictionary edit took effect before shipping IVR prompts_

  ```bash
  sarvam-pp-cli pron-check "SarvamPay" --lang hi-IN
  ```

### Local state that compounds
- **`chat resume`** — Continue a past chat thread from local history with full context

  _Use to continue an assistant session without losing context, offline from the original thread_

  ```bash
  sarvam-pp-cli chat resume 20260814_2d09e061 "what was our conclusion?"
  ```
- **`subs`** — Emit .srt/.vtt subtitles from timestamped transcriptions in local history

  _Use to turn a timestamped transcription into subtitles without a throwaway script_

  ```bash
  sarvam-pp-cli subs --from last --format srt --output subtitles.srt
  ```

### Job orchestration
- **`stt-job retry`** — Re-run only the failed files of a batch STT job with one command

  _Use when a batch job partially fails and you need to reprocess just the failures_

  ```bash
  sarvam-pp-cli stt-job retry 20260707_9f1c2b3a-4d5e-6f70-8a9b-c0d1e2f3a4b5 --failed-only --dir ./audio/
  ```
- **`stt-job report`** — Per-file digest of a batch STT job with typed exit codes for cron alerting

  _Use in cron to alert when a batch transcription job degrades_

  ```bash
  sarvam-pp-cli stt-job report 20260707_9f1c2b3a-4d5e-6f70-8a9b-c0d1e2f3a4b5 --json
  ```

### Document intelligence
- **`docai schema list`** — Save, list, and diff doc-ai extraction schemas locally

  _Use to version extraction schemas so schema changes never silently break extraction runs_

  ```bash
  sarvam-pp-cli docai schema list
  ```
- **`docai batch`** — Run a saved extraction schema over a folder of documents with job pacing

  _Use for weekly batch document extraction (KYC, invoices) without writing orchestration code_

  ```bash
  sarvam-pp-cli docai batch --schema invoice-v1 --dir ./docs/ --out ./results/
  ```

## Command Reference

**chat** — Manage chat

- `sarvam-pp-cli chat` — Creates a model response for the given chat conversation. Serves sarvam-105b and sarvam-105b-conversations models.

**doc-ai** — Manage doc ai

- `sarvam-pp-cli doc-ai digitise` — Creates and starts a digitisation job from files or pre-uploaded handles.
- `sarvam-pp-cli doc-ai download-url` — Returns a presigned URL to download the output of a completed doc-ai job.
- `sarvam-pp-cli doc-ai extract` — Creates and starts an extract job from files or pre-uploaded handles.
- `sarvam-pp-cli doc-ai results` — Fetches the results of a completed doc-ai job, including extracted fields and annotations with confidence scores.
- `sarvam-pp-cli doc-ai status` — Polls the status of a doc-ai job until a terminal status (completed, partially_completed, failed, rejected).
- `sarvam-pp-cli doc-ai upload` — Creates a presigned URL to upload a document for doc-ai processing.

**models** — Manage models

- `sarvam-pp-cli models` — Lists the model IDs this deployment currently serves.

**speech-to-text** — Manage speech to text

- `sarvam-pp-cli speech-to-text speech-to-text` — Transcribes speech to text in multiple Indian languages and English. Accepts an audio file via multipart form-data.
- `sarvam-pp-cli speech-to-text stt-job-download` — Returns presigned download URLs for the output files of a completed batch speech-to-text job.
- `sarvam-pp-cli speech-to-text stt-job-initiate` — Creates a new speech-to-text bulk job and returns a job UUID and storage details for processing multiple audio files.
- `sarvam-pp-cli speech-to-text stt-job-start` — Starts processing a speech-to-text bulk job after all audio files have been uploaded.
- `sarvam-pp-cli speech-to-text stt-job-status` — Returns the status of a batch speech-to-text job including per-file details and download information.
- `sarvam-pp-cli speech-to-text stt-job-upload` — Generates presigned upload URLs for audio files that will be processed in a speech-to-text bulk job.

**text-lid** — Manage text lid

- `sarvam-pp-cli text-lid` — Identifies the language (e.g. hi-IN) and script (e.g.

**text-to-speech** — Manage text to speech

- `sarvam-pp-cli text-to-speech create-pronunciation-dictionary` — Uploads a .json file to create a new pronunciation dictionary. Only supported by bulbul:v3.
- `sarvam-pp-cli text-to-speech delete-pronunciation-dictionary` — Deletes a pronunciation dictionary by ID.
- `sarvam-pp-cli text-to-speech get-pronunciation-dictionary` — Fetches a single pronunciation dictionary by ID.
- `sarvam-pp-cli text-to-speech list-pronunciation-dictionaries` — Lists all pronunciation dictionaries for the user. Dictionaries define custom word pronunciations used by bulbul:v3 TTS.
- `sarvam-pp-cli text-to-speech stream` — Converts the input text into a streamed spoken audio response using the specified output codec (e.g. MP3).
- `sarvam-pp-cli text-to-speech text-to-speech` — Converts text into spoken audio.
- `sarvam-pp-cli text-to-speech update-pronunciation-dictionary` — Updates an existing pronunciation dictionary with a new JSON file.

**translate** — Manage translate

- `sarvam-pp-cli translate` — Converts text from one language to another while preserving meaning. Supports 22 Indic languages plus English.

**transliterate** — Manage transliterate

- `sarvam-pp-cli transliterate` — Transliterates text from one script to another (e.g.


### Finding the right command

When you know what you want to do but not which command does it, ask the CLI directly:

```bash
sarvam-pp-cli which "<capability in your own words>"
```

`which` resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code `0` means at least one match; exit code `2` means no confident match — fall back to `--help` or use a narrower query.

## Recipes

### Translate a support reply

```bash
sarvam-pp-cli translate --input "Your EMI of Rs. 3000 is pending" --source-language-code en-IN --target-language-code hi-IN --mode formal
```

Translate a single message with formal tone for customer communications

### Audition voices for a new prompt set

```bash
sarvam-pp-cli voices preview --lang hi-IN --sample "नमस्ते, स्वागत है" --speakers shubh,ritu,priya
```

Hear 3 voices on the same sample before committing to a speaker

### Subtitle a promo video

```bash
sarvam-pp-cli subs --from last --format srt --output subtitles.srt
```

Turn the last timestamped transcription into SRT subtitles

### Reprocess failed batch files

```bash
sarvam-pp-cli stt-job retry 20260707_9f1c2b3a-4d5e-6f70-8a9b-c0d1e2f3a4b5 --failed-only --dir ./audio/
```

Re-run only the files that failed in a batch transcription job

### Check a pronunciation dictionary

```bash
sarvam-pp-cli pron-check "SarvamPay" --lang hi-IN --dict p_5cb7faa6
```

Verify a custom pronunciation actually changed how the term sounds

### Extract fields from a folder of documents

```bash
sarvam-pp-cli docai batch --schema invoice-v1 --dir ./docs/ --out ./results/
```

Run a saved extraction schema over every document with pacing

### Continue a past chat session

```bash
sarvam-pp-cli chat resume 20260814_2d09e061 "what was our conclusion?"
```

Pick up an assistant conversation where it left off, with full context

## Auth Setup

Authentication uses the Sarvam AI API subscription key (sk_ format). Set it with `export SARVAM_API_KEY=sk_...` or `sarvam-pp-cli auth set-token`. The key goes in the `api-subscription-key` header (or `Authorization: Bearer`). Note: an invalid key returns HTTP 403 with `invalid_api_key_error`, not 401 — treat 403 as the auth-failure signal. The separate Voice Agents platform (apps.sarvam.ai) uses a different X-API-Key system and is out of scope for this CLI.

Run `sarvam-pp-cli doctor` to verify setup.

## Agent Mode

Add `--agent` to any command. Expands to: `--json --compact --no-input --no-color --yes`.

- **Pipeable** — JSON on stdout, errors on stderr
- **Filterable** — `--select` keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:

  ```bash
  sarvam-pp-cli models --agent --select created,id,object
  ```
- **Previewable** — `--dry-run` shows the request without sending
- **Offline-friendly** — sync/search commands can use the local SQLite store when available
- **Non-interactive** — never prompts, every input is a flag
- **Explicit retries** — use `--idempotent` only when an already-existing create should count as success, and use `--ignore-missing` only when a missing delete target should count as success

### Response envelope

Commands that read from the local store or the API wrap output in a provenance envelope:

```json
{
  "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."},
  "results": <data>
}
```

Parse `.results` for data and `.meta.source` to know whether it's live or local. A human-readable `N results (live)` summary is printed to stderr only when stdout is a terminal AND no machine-format flag (`--json`, `--csv`, `--compact`, `--quiet`, `--plain`, `--select`) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout.

## Paths and state

Agents should treat the CLI's path resolver as part of the runtime contract:

- Use `--home <dir>` for one invocation, or set `SARVAM_HOME=<dir>` to relocate all four path kinds under one root.
- Use per-kind env vars only when a specific kind must diverge: `SARVAM_CONFIG_DIR`, `SARVAM_DATA_DIR`, `SARVAM_STATE_DIR`, `SARVAM_CACHE_DIR`.
- Resolution order is per-kind env var, `--home`, `SARVAM_HOME`, XDG (`XDG_CONFIG_HOME`, `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME`), then platform defaults.
- `config` contains settings like `config.toml` and profiles. `data` contains `credentials.toml`, `data.db`, cookies, and auth sidecars. `state` contains persisted queries, jobs, and `teach.log`. `cache` contains regenerable HTTP/cache files.
- Stored secrets live in `credentials.toml` under the data dir. Existing legacy `config.toml` secrets are read for compatibility and leave `config.toml` on the first auth write.
- Run `sarvam-pp-cli doctor --fail-on warn` to surface path and credential-location warnings. `agent-context` exposes a schema v4 `paths` block for agents that need the resolved dirs.
- For MCP, pass relocation through the MCP host config. The MCP binary does not inherit CLI flags:

  ```json
  {
    "mcpServers": {
      "sarvam": {
        "command": "sarvam-pp-mcp",
        "env": {
          "SARVAM_HOME": "/srv/sarvam"
        }
      }
    }
  }
  ```

Fleet precedence: an inherited per-kind env var overrides an explicit `--home` for that kind. Use `SARVAM_HOME` or per-kind vars as durable fleet levers, and use `--home` only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing `SARVAM_HOME`, or `doctor` will not find credentials left under the former root.

## Automatic learning

This CLI ships a self-capturing learning loop. The CLI does its own bookkeeping: every invocation is journaled locally, a failed flag followed by a corrected retry auto-derives a `flag_alias` candidate, and a `teach` on a query family without a playbook auto-synthesizes a `playbook_candidate` from the session's journal. Your job is judgment only: `recall` first, act on surfaced candidates, `teach` the final answer, `playbook amend` when you observe a correction. You never record failures by hand.

### Step 1: `recall` before any discovery

Before list/search/drill commands on a new user question, run:

```bash
sarvam-pp-cli recall "<user's question>" --agent
```

The response envelope:

```json
{
  "query": "...",
  "normalized": "<normalized form>",
  "query_entities": ["..."],
  "found": true | false,
  "match_score": 0.0,
  "results": [
    { "resource_id": "...", "resource_type": "...", "venue": "...",
      "confidence": 2, "entity_match": "exact|partial|unknown",
      "source": "taught|preseed|pattern", "warnings": ["..."] }
  ],
  "mismatches": [ /* only when --debug-mismatches */ ],
  "warnings": [ /* top-level */ ],
  "candidates": [
    { "id": 12, "class": "flag_alias | playbook_candidate",
      "summary": "...", "sightings": 3, "last_seen": "...",
      "rationale": "...",
      "next_action": ["<trial command>", "sarvam-pp-cli learnings confirm 12"] }
  ],
  "playbook": {
    "query_family": "...",
    "playbook": {
      "steps": [ { "cmd": "<command with {slot} substitution>", "purpose": "..." } ],
      "entity_slots": ["$ENTITY"],
      "expected_tool_calls": 3
    },
    "slots_resolved": { "$ENTITY": { "token": "<live token>", "canonical": "<canonical>" } },
    "notes": "<workarounds + gotchas for this query family>"
  },
  "notes": "<duplicate surface for non-playbook callers>"
}
```

Empty-store short-circuit: if the store has no learnings, playbooks, or candidates yet (recall finds nothing and `learnings list` and `learnings candidates` are both empty), skip recall for the rest of this session instead of taxing every query; resume recall-first once something has been taught.

### Step 2: decision tree

Read `candidates`, `playbook`, `notes`, `results[0]`, and warnings in that order:

```
if Candidates present (warnings include "candidates_present"):
    -> candidates are try-then-confirm, never facts. Follow each candidate's
       two-step next_action verbatim: run the trial command first, then run
       `learnings confirm <id>` only after the trial verified the behavior.
       Reject a wrong candidate with `learnings reject <id>`.
    -> NEVER re-teach something recall surfaced as a candidate; confirm or
       reject that candidate instead of teaching a duplicate.
    -> candidates ride alongside playbooks and resource hits, not instead of
       them; continue with the branches below after acting on them.

if Playbook present:
    -> READ Playbook.notes verbatim FIRST (workarounds + gotchas the CLI surface doesn't expose)
    -> replay Playbook.steps in order, substituting Playbook.slots_resolved entries
       for the entity slot tokens. If a step's slot is unresolved, fall back to
       discovery for that step only.
    -> the Playbook's expected_tool_calls is a budget; if you find yourself running
       materially more, record the divergence via `sarvam-pp-cli playbook amend`
       at end-of-session.

elif Notes present (no Playbook):
    -> read Notes verbatim before any discovery step; they carry known gotchas
       for this query family even when no structured choreography exists yet.

elif Found AND Results[0].EntityMatch == "exact" AND Results[0].Confidence >= 2:
    -> skip discovery; fetch live data for Results[*].ResourceID in parallel

elif Found AND Results[0].EntityMatch == "partial":
    -> candidate hint, NOT a hit; read the resource title to validate before trusting

elif (any row in Mismatches[] when --debug-mismatches was passed):
    -> treat as cold start; the stored learning is for a different entity
       (different canonical resolved from query_entities)

else:  // Found == false, no playbook, no notes
    -> cold start; run discovery normally; teach the answer afterward (Step 4).
       If the family has no playbook yet, that teach auto-synthesizes a
       playbook candidate from this session's journal - you do not need to
       record one by hand.
```

Playbook and Notes are orthogonal to the per-resource path. A recall response can carry both a Playbook AND a `Results[]` hit - use both: the Playbook tells you which choreography to run; the resource hits short-circuit specific steps. Default to skipping `mismatches`; pass `--debug-mismatches` only when investigating cold-start surprises.

Candidate judgment details: `learnings confirm <id>` prints the candidate's full payload before materializing it - check that the printed payload matches the behavior you verified. `learnings reject <id>` tombstones the derivation signature so the same candidate does not resurface. The envelope carries only the few candidates worth acting on now; `sarvam-pp-cli learnings candidates` lists the full open set.

Graceful degradation: if `learnings confirm` is an unknown command, you are driving an older binary - ignore the candidates guidance and follow the rest of the protocol.

### Step 3: always read `