# https://compresr.ai/docs/sdks/python

> Human-readable page: https://compresr.ai/docs/sdks/python

The `compresr` package is the official Python client. It wraps the REST API with typed methods, handles auth, and ships both sync and async variants of every call. Python 3.9+.

## 1. Install

Install from PyPI: `pip`, `poetry`, and `uv` all work.

pip:

```bash
pip install compresr
```

poetry:

```bash
poetry add compresr
```

uv:

```bash
uv add compresr
```

> **The agent client ships in the base install**
> As of `compresr 2.8.2` the [agent client](#6-agent-client) layer (`client.messages.create`, `client.chat.completions.create`, `client.run`, `WebSearchTool`) is part of the base install: `pip install compresr` is enough. LangChain + provider chat-model + Tavily/Brave deps are pulled in automatically. Old `compresr[agents]` / `compresr[agents-all]` brackets still work as no-op aliases.

## 2. Initialize the client

Construct `CompressionClient` once at module scope and reuse it: the client keeps an internal `httpx` connection pool. Read the key from env, never hardcode.

The constructor takes `api_key` (required), plus optional `base_url` (defaults to `https://api.compresr.ai`) and `timeout` (seconds; default uses the SDK's built-in timeout). Override `base_url` only for regional or self-hosted endpoints.

```python
import os
from compresr import CompressionClient

# Minimal (recommended)
client = CompressionClient(api_key=os.environ["COMPRESR_API_KEY"])

# Explicit overrides
client = CompressionClient(
    api_key=os.environ["COMPRESR_API_KEY"],
    base_url="https://api.compresr.ai",  # optional
    timeout=300,                          # optional, seconds
)
```

**TypeScript**

```typescript
import { CompressionClient } from '@compresr/sdk';

const client = new CompressionClient({
  apiKey: process.env.COMPRESR_API_KEY!,
});
```

**cURL**

```bash
# No client object - set the key and base URL once.
export COMPRESR_API_KEY="cmp_your_api_key"
export COMPRESR_BASE_URL="https://api.compresr.ai"
```

See [Authentication](/docs/authentication) for key rotation, budgets, and rules.

#### Constructor options

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `api_key` | str \| None | no | If omitted, resolved from `COMPRESR_API_KEY`, then the `[default]` profile in `~/.compresr/credentials` (INI format, populated by `compresr-sdk login`). Select a different profile with `COMPRESR_PROFILE`. |
| `base_url` | str \| None | no | API endpoint override. Non-HTTPS URLs are refused unless `COMPRESR_ALLOW_INSECURE=1` is set (raises `CompresrError("insecure_base_url")`). |
| `timeout` | int \| None | no | Request timeout in seconds (per HTTP call). |
| `retry_config` | RetryConfig \| None | no | Override the retry policy. Default retries `429` and `503` with exponential backoff (respects `Retry-After`). Import `RetryConfig` from `compresr`: `RetryConfig(max_retries=..., retry_on_status=...)`. |
| `llm` | str \| None | no | Provider for the agent surface (e.g. `"anthropic"`, `"openai:gpt-4o-mini"`). Required for `client.messages` / `client.chat` / `client.run` / `client.research`. See [Section 6](#6-agent-client). |
| `llm_api_key` | str \| None | no | API key for the LLM provider. Falls back to `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GOOGLE_API_KEY` depending on `llm`. |
| `llm_http_client` | httpx.Client \| None | no | Custom sync `httpx.Client` the SDK uses for the downstream LLM call. Escape hatch for corporate proxies, custom CA bundles, and mTLS: `httpx.Client(verify="/etc/ssl/corp-ca.pem", proxies="http://proxy:3128")`. |
| `llm_http_async_client` | httpx.AsyncClient \| None | no | Async twin of `llm_http_client`, used by `acreate` / `arun`. |
| `compression` | dict \| CompressionPolicy \| None | no | Middleware compression policy applied to every tool output. See [Compression knobs](#compression-knobs-compression). |
| `enable_prompt_cache` | bool | no | Enable provider-side prompt caching (Anthropic `cache_control`, OpenAI `prompt_cache_key`). No-op for Gemini (implicit caching is always on server-side). |
| `prompt_cache_ttl` | "5m" \| "1h" | no | Anthropic cache TTL. Longer TTL costs more per cache write but survives longer between calls. On OpenAI, `"1h"` maps to `prompt_cache_retention: "24h"`. |
| `prompt_cache_min_messages` | int | no | Skip caching for very short conversations (avoids paying the cache-write premium on trivial prompts). |
| `openai_prompt_cache_key` | str \| None | no | Explicit OpenAI `prompt_cache_key`; when omitted the SDK does not set `prompt_cache_key` on the OpenAI call and provider defaults apply. |

#### Environment variables

| Variable | Purpose |
|---|---|
| `COMPRESR_API_KEY` | Fallback for `api_key`. Also read by the cURL examples. |
| `COMPRESR_BASE_URL` | Fallback for `base_url`. |
| `COMPRESR_ALLOW_INSECURE` | Set to `1` to allow non-HTTPS `base_url` (local dev / self-hosted gateway). Client refuses to start otherwise. |
| `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GOOGLE_API_KEY` | Fallback for `llm_api_key` when the matching provider is used. |

> **CLI authentication**
> Instead of passing `api_key=` explicitly, run **`compresr-sdk login`** to write credentials to the `[default]` profile in `~/.compresr/credentials` (INI format, no extension); the SDK picks them up automatically on the next `CompressionClient()` call. Set `COMPRESR_PROFILE` to select a different profile. **`compresr-sdk logout`** clears them. Both are also exposed programmatically as `login()` / `logout()` from `compresr`.

## 3. compress

Synchronous single-request compression. Pass `context`, `query`, and `compression_model_name="{{DEFAULT_MODEL}}"`; the model keeps the spans that matter for the query. For many chunks against one query, see [`compress_batch`](#5-batch); for incremental output, [`compress_stream`](#4-stream).

```python
result = client.compress(
    context=(
        "The James Webb Space Telescope (JWST) is a space telescope designed "
        "primarily to conduct infrared astronomy. Its 6.5-metre primary mirror "
        "is composed of 18 gold-coated hexagonal beryllium segments. JWST orbits "
        "the Sun near the Sun-Earth L2 Lagrange point, about 1.5 million "
        "kilometres from Earth, where its sunshield keeps the instruments "
        "below 50 K. Launched on 25 December 2021, it is operated jointly by "
        "NASA, ESA, and the Canadian Space Agency."
    ),
    query="What is the diameter of JWST's primary mirror?",
    compression_model_name="{{DEFAULT_MODEL}}",
    target_compression_ratio=0.5,
)

print(result.data.compressed_context)
print(
    f"{result.data.original_tokens} → {result.data.compressed_tokens} tokens "
    f"({result.data.actual_compression_ratio:.2f}x)"
)
```

**TypeScript**

```typescript
const result = await client.compress({
  context:
    "The James Webb Space Telescope (JWST) is a space telescope designed " +
    "primarily to conduct infrared astronomy. Its 6.5-metre primary mirror " +
    "is composed of 18 gold-coated hexagonal beryllium segments. JWST orbits " +
    "the Sun near the Sun-Earth L2 Lagrange point, about 1.5 million " +
    "kilometres from Earth, where its sunshield keeps the instruments " +
    "below 50 K. Launched on 25 December 2021, it is operated jointly by " +
    "NASA, ESA, and the Canadian Space Agency.",
  query: "What is the diameter of JWST's primary mirror?",
  compressionModelName: '{{DEFAULT_MODEL}}',
  targetCompressionRatio: 0.5,
});

console.log(result.data.compressed_context);
```

**cURL**

```bash
curl -X POST https://api.compresr.ai/api/compress/question-specific/ \
  -H "X-API-Key: $COMPRESR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "context": "The James Webb Space Telescope (JWST) is a space telescope...",
    "query": "What is the diameter of JWST'"'"'s primary mirror?",
    "compression_model_name": "{{DEFAULT_MODEL}}",
    "target_compression_ratio": 0.5
  }'
```

### Parameters

`latte_v2` accepts every parameter `latte_v1` accepts, **plus** three `latte_v2`-only knobs for dynamic compression-ratio selection. See the [Models reference](/docs/api-reference/models) for the canonical decision guide and the at-a-glance support matrix.

#### Shared parameters (both models)

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `context` | string | yes | The long text to compress: RAG chunks, document body, chat history. |
| `query` | str \| None | no | The question the compressed context must still answer. **Required for `latte_v1`**; optional for `latte_v2`. Backend validates. |
| `compression_model_name` | "latte_v1" \| "latte_v2" | no | Routes the call. SDK default is `{{SDK_DEFAULT_MODEL}}` for stability; pass `"{{DEFAULT_MODEL}}"` to opt into the newer backbone. See the [Models](/docs/api-reference/models) reference. |
| `target_compression_ratio` | float \| None | no | Removal strength when `0 < r ≤ 1`, or Nx target when `r > 1` (e.g. `60` = 60×). Server hard-caps at `200`. Ignored on `latte_v2` when `dynamic=True`. See [Models › target_compression_ratio](/docs/api-reference/models#target_compression_ratio). |
| `coarse` | boolean \| None | no | `None` = backend default (paragraph-level); `True` locks paragraph-level; `False` opts into token-level precision. |
| `heuristic_chunking` | boolean \| None | no | Heuristic splitter (paragraphs, code blocks) instead of fixed-size chunks. |
| `disable_placeholders` | boolean \| None | no | Skip the `[...]` placeholders inserted where content was dropped. |

#### `latte_v2`-only parameters

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `dynamic` | bool \| None | no | `None` = server default; `True` picks the compression ratio per-input automatically inside `[dynamic_min_ratio, dynamic_max_ratio]` and overrides `target_compression_ratio`; `False` explicitly forces the fixed-ratio path. Rejected on `latte_v1` with `ValidationError`. |
| `dynamic_min_ratio` | float \| None | no | Floor on the chosen Nx ratio when `dynamic=True`. Must be `≥ 1.0`. Only consulted when `dynamic=True`. |
| `dynamic_max_ratio` | float \| None | no | Ceiling on the chosen Nx ratio when `dynamic=True`. Must be `≥ 1.0`. Only consulted when `dynamic=True`. |

### Response

`compress()` returns a typed object; access fields as attributes (`result.data.compressed_context`). Response field names stay snake_case across every SDK.

| Field | Type | Description |
| --- | --- | --- |
| `data` | object |  |
| `data.compressed_context` | string | The compressed text, ready to drop into your prompt. |
| `data.original_tokens` | integer | Token count of the input context (tiktoken cl100k). |
| `data.compressed_tokens` | integer | Token count of the compressed output. |
| `data.tokens_saved` | integer | original_tokens − compressed_tokens. |
| `data.actual_compression_ratio` | number | Fraction of input tokens removed (0..1) when target_compression_ratio was 0..1, or the achieved Nx factor when the Nx form was requested. Mirrors the input regime. |
| `data.duration_ms` | integer | Server-side wall-clock time for the compression pass. |

## 4. Stream

`client.compress_stream(...)` returns an iterator yielding `{content, done}` chunks as the model produces them; the final chunk has `done=True` and empty `content`. Use it anywhere time-to-first-token matters (UIs, agent loops); for one-shot calls stick with [`compress`](#3-compress).

```python
for chunk in client.compress_stream(
    context=long_document,
    query="What was the project's Q3 churn rate?",
    compression_model_name="{{DEFAULT_MODEL}}",
):
    print(chunk.content, end="", flush=True)
    if chunk.done:
        break
```

**TypeScript**

```typescript
for await (const chunk of client.compressStream({
  context: longDocument,
  query: "What was the project's Q3 churn rate?",
  compressionModelName: '{{DEFAULT_MODEL}}',
})) {
  process.stdout.write(chunk.content);
  if (chunk.done) break;
}
```

**cURL**

```bash
curl -N -X POST https://api.compresr.ai/api/compress/question-specific/stream \
  -H "X-API-Key: $COMPRESR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "context": "<your long document>",
    "query": "What was the project'"'"'s Q3 churn rate?",
    "compression_model_name": "{{DEFAULT_MODEL}}"
  }'
```

The iterator is a normal generator: wrap it in `itertools.islice`, push chunks through a queue, or consume from a worker thread. Same `context` / `query` / `compression_model_name` rules as `compress()`.

## 5. Batch

`client.compress_batch(...)` compresses many contexts in one request. Pass `contexts: list[str]` plus either a single `queries: str` (applied to every context) or a `queries: list[str]` matching `contexts` in length. Cheaper than firing N concurrent `compress()` calls, and ideal for RAG re-ranking or bulk document processing.

```python
# One query against many candidate contexts (typical RAG re-ranking shape).
result = client.compress_batch(
    contexts=[chunk_1, chunk_2, chunk_3, chunk_4],
    queries="What did the customer cite as the reason for churn?",
    compression_model_name="{{DEFAULT_MODEL}}",
)

# Or per-context queries: one query per context, same length.
result = client.compress_batch(
    contexts=[doc_a, doc_b, doc_c],
    queries=[
        "Who signed the contract?",
        "When was the renewal date?",
        "What was the agreed unit price?",
    ],
    compression_model_name="{{DEFAULT_MODEL}}",
)

for item in result.data.results:
    print(item.compressed_context)
```

**TypeScript**

```typescript
// One query against many candidate contexts.
const result = await client.compressBatch({
  contexts: [chunk1, chunk2, chunk3, chunk4],
  queries: 'What did the customer cite as the reason for churn?',
  compressionModelName: '{{DEFAULT_MODEL}}',
});

// Or per-context queries: same length as contexts.
const perItem = await client.compressBatch({
  contexts: [docA, docB, docC],
  queries: [
    'Who signed the contract?',
    'When was the renewal date?',
    'What was the agreed unit price?',
  ],
  compressionModelName: '{{DEFAULT_MODEL}}',
});

for (const item of result.data.results) {
  console.log(item.compressed_context);
}
```

**cURL**

```bash
curl -X POST https://api.compresr.ai/api/compress/question-specific/batch \
  -H "X-API-Key: $COMPRESR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contexts": ["<chunk 1>", "<chunk 2>", "<chunk 3>", "<chunk 4>"],
    "queries": "What did the customer cite as the reason for churn?",
    "compression_model_name": "{{DEFAULT_MODEL}}"
  }'
```

`queries` is either a string (applied to every context) or a list matching `contexts` in length; mixing the two raises `ValidationError`. Per-item results carry the same fields as a single `compress()` call **except** `target_compression_ratio` (request-level only). The envelope also exposes aggregates: `result.data.count`, `total_original_tokens`, `total_compressed_tokens`, `total_tokens_saved`, `average_compression_ratio`.

#### Alternate form: `inputs=[{context, query}, ...]`

The wire format is a list of `{context, query}` pairs. Pass `inputs=` instead of `contexts=`/`queries=` when it matches your data shape more naturally (queues, streaming pipelines, per-item queries). **Exactly one of `inputs` OR `contexts` is required**; passing both — or neither — raises `ValidationError`.

```python
result = client.compress_batch(
    inputs=[
        {"context": doc_a, "query": "Who signed the contract?"},
        {"context": doc_b, "query": "When was the renewal date?"},
        {"context": doc_c, "query": "What was the agreed unit price?"},
    ],
    compression_model_name="{{DEFAULT_MODEL}}",
)
```

**TypeScript**

```typescript
const result = await client.compressBatch({
  inputs: [
    { context: docA, query: 'Who signed the contract?' },
    { context: docB, query: 'When was the renewal date?' },
    { context: docC, query: 'What was the agreed unit price?' },
  ],
  compressionModelName: '{{DEFAULT_MODEL}}',
});
```

## 6. Agent client

Construct `CompressionClient` with `llm=` and you get an **agent surface**: three call-shapes (Anthropic-style `messages.create`, OpenAI-style `chat.completions.create`, native `run`) that auto-compress every tool output above `min_tokens` before the LLM sees it. Behind all three sits LangChain 1.0's `create_agent` + the SDK's `CompresrToolMiddleware`. Use it as a drop-in for `anthropic.Anthropic()` / `openai.OpenAI()`; for raw `(context, query)` calls stick with [`compress`](#3-compress).

> These surfaces are SDK-shaped and have no direct cURL equivalent. The underlying compression is still the same `/api/compress/question-specific/` endpoint; it's what the middleware fires whenever a tool returns.

### Construct with `llm=`

Provider lives on the client; **model lives at the call site**. Swap providers by changing one string: same tools, same code:

```python
import os
from compresr import CompressionClient, WebSearchTool

client = CompressionClient(
    api_key=os.environ["COMPRESR_API_KEY"],
    llm="anthropic",                                # or "openai", "google_genai"
    llm_api_key=os.environ["ANTHROPIC_API_KEY"],
    compression={"target_compression_ratio": 0.5, "min_tokens": 300},
)
```

**TypeScript**

```typescript
import { CompressionClient, WebSearchTool } from '@compresr/sdk';

const client = new CompressionClient({
  apiKey: process.env.COMPRESR_API_KEY!,
  llm: 'anthropic',                                 // or 'openai', 'google_genai'
  llmApiKey: process.env.ANTHROPIC_API_KEY!,
  compression: { targetCompressionRatio: 0.5, minTokens: 300 },
});
```

The `llm` string accepts `"anthropic"` (provider only, every call must pass `model="..."`), `"anthropic:claude-haiku-4-5"` (default model, overridable at call site), or `"anthropic/claude-haiku-4-5"` (Vercel AI SDK convention; both separators accepted). If neither provides a model, the SDK raises `CompresrError("model is required …")`.

### Three call shapes

`messages.create` duck-types `anthropic.types.Message`, `chat.completions.create` duck-types `openai.types.chat.ChatCompletion`, and `run` returns a native `NormalizedResult` (`.text`, `.tool_uses`, `.citations`, `.stop_reason`, `.usage`).

```python
tavily = WebSearchTool.tavily(api_key=os.environ["TAVILY_API_KEY"], max_results=3)
messages = [{"role": "user", "content": "What's the latest AI news?"}]

# Anthropic shape
msg = client.messages.create(model="claude-haiku-4-5", max_tokens=512, messages=messages, tools=[tavily])
# OpenAI shape
completion = client.chat.completions.create(model="gpt-4o-mini", messages=messages, tools=[tavily])
# Native
result = client.run(prompt="What's the latest AI news?", model="claude-haiku-4-5", tools=[tavily], max_tokens=512)
```

**TypeScript**

```typescript
const search = await WebSearchTool.tavily({ apiKey: process.env.TAVILY_API_KEY!, maxResults: 3 });
const messages = [{ role: 'user' as const, content: "What's the latest AI news?" }];

// Anthropic shape
const msg = await client.messages.create({ model: 'claude-haiku-4-5', maxTokens: 512, messages, tools: [search] });
// OpenAI shape
const completion = await client.chat.completions.create({ model: 'gpt-4o-mini', messages, tools: [search] });
// Native
const result = await client.run({ prompt: "What's the latest AI news?", model: 'claude-haiku-4-5', tools: [search], maxTokens: 512 });
```

Python also exposes async variants: `acreate`, `arun`. TypeScript is async by default.

> **run() and arun() are keyword-only**
> `client.run(...)` and `client.arun(...)` accept **only** keyword arguments — `client.run("question")` raises `TypeError`. Always pass `prompt=...`, `model=...`, `tools=...` by name. This matches how `messages.create` and `chat.completions.create` are called.

### Web search: `WebSearchTool`

Three providers ship in the box: **Tavily**, **Brave**, and **AgentCore** (Amazon Bedrock via MCP). All three return a real LangChain `BaseTool`; their output flows through `CompresrToolMiddleware` automatically.

```python
from compresr import WebSearchTool

# Tavily — domain filtering supported natively.
tavily = WebSearchTool.tavily(
    api_key=os.environ["TAVILY_API_KEY"],
    max_results=5,
    allowed_domains=["nytimes.com"],   # optional
    blocked_domains=["example.com"],   # optional
)

# Brave — reads BRAVE_SEARCH_API_KEY (preferred) or BRAVE_API_KEY.
brave = WebSearchTool.brave(max_results=5)

# AgentCore — reads five env vars (see table below).
agentcore = WebSearchTool.agentcore(max_results=5)
```

**TypeScript**

```typescript
import { WebSearchTool } from '@compresr/sdk';

const tavily = await WebSearchTool.tavily({
  apiKey: process.env.TAVILY_API_KEY!,
  maxResults: 5,
  allowedDomains: ['nytimes.com'],
});

// Brave: BRAVE_SEARCH_API_KEY (preferred) or BRAVE_API_KEY.
const brave = await WebSearchTool.brave({ maxResults: 5 });

// AgentCore: reads five env vars (see table below).
const agentcore = await WebSearchTool.agentcore({ maxResults: 5 });
```

> **Why not Anthropic / OpenAI / Gemini server search?**
> Provider-native server search tools (`web_search_20250305`, `web_search_preview`, `google_search`) execute server-side and return opaque/encrypted content that Compresr cannot read or compress. Use Tavily, Brave, or AgentCore so the result is plaintext. See the [Web search guide](/docs/guides/web-search).

#### Provider reference

**Tavily** (`WebSearchTool.tavily`) — reads `api_key=`, then `TAVILY_API_KEY` env var. Raises `ValueError` if neither is set. Supports `allowed_domains` / `blocked_domains` natively.

**Brave** (`WebSearchTool.brave`) — reads `api_key=`, then `BRAVE_SEARCH_API_KEY`, then `BRAVE_API_KEY`. Raises `ValueError` if none of the three are set. `allowed_domains` / `blocked_domains` are **not** supported (Brave uses Goggles for filtering, out of scope); passing them emits a `UserWarning`.

**AgentCore** (`WebSearchTool.agentcore`) — install with `pip install compresr[agentcore]`. Talks to an Amazon Bedrock AgentCore gateway over MCP streamable-HTTP, authenticated via a Cognito OAuth 2.0 client-credentials handshake. Bearer tokens are cached; a 401 triggers one automatic re-mint. `max_results` is clamped to 1..25; responses larger than 1 MB are rejected. `allowed_domains` / `blocked_domains` are accepted for signature parity but emit a `UserWarning` — use Tavily if you need domain filtering.

AgentCore config resolves per field with precedence **explicit arg → AgentCore-namespaced env → short env**. If any field is unresolved, `WebSearchTool.agentcore(...)` raises `ValueError` listing every missing field:

| Argument | Env var (primary) | Env var (fallback) |
|---|---|---|
| `gateway_url` | `AGENTCORE_GATEWAY_MCP_URL` | `GATEWAY_MCP_URL` |
| `cognito_token_url` | `AGENTCORE_COGNITO_TOKEN_URL` | `COGNITO_TOKEN_URL` |
| `client_id` | `AGENTCORE_COGNITO_CLIENT_ID` | `COGNITO_CLIENT_ID` |
| `client_secret` | `AGENTCORE_COGNITO_CLIENT_SECRET` | `COGNITO_CLIENT_SECRET` |
| `scope` | `AGENTCORE_COGNITO_SCOPE` | `COGNITO_SCOPE` |

### Bring your own tool

Any LangChain `@tool`-decorated function works. The string return value is compressed before the LLM sees it.

```python
from langchain_core.tools import tool

@tool
def kb_lookup(topic: str) -> str:
    """Look up the internal policy on the given topic."""
    return INTERNAL_KB.get(topic, "Not found.")

client.messages.create(model="claude-haiku-4-5", max_tokens=256,
    messages=[{"role": "user", "content": "Refund policy?"}], tools=[kb_lookup])
```

**TypeScript**

```typescript
import { tool } from '@langchain/core/tools';
import { z } from 'zod';

const kbLookup = tool(async ({ topic }) => INTERNAL_KB[topic] ?? 'Not found.', {
  name: 'kb_lookup',
  description: 'Look up the internal policy on the given topic.',
  schema: z.object({ topic: z.string() }),
});

await client.messages.create({ model: 'claude-haiku-4-5', maxTokens: 256,
  messages: [{ role: 'user', content: 'Refund policy?' }], tools: [kbLookup] });
```

Streaming isn't on the agent layer yet: the Python facades expose only `.create` / `.acreate`, so `client.messages.stream(...)` / `client.chat.completions.stream(...)` raise `AttributeError` (not a typed `CompresrError`). The compression-API stream ([`compress_stream`](#4-stream)) is unaffected.

### Research: `client.research`

When constructed with `llm=`, the client also exposes `client.research`, a multi-step search-and-summarize loop that runs a web-search tool for you, compresses each snippet before it enters the LLM's context, and returns a structured result with citations. `client.research.run(question)` runs the full loop (up to `max_steps`); `client.research.search(question)` is the same loop capped at 2 steps for quick lookups.

Accessing `client.research` when `llm=` was not passed raises `CompresrError` — the facade needs a chat model to reason about search results.

```python
from compresr import CompressionClient

client = CompressionClient(
    api_key=os.environ["COMPRESR_API_KEY"],
    llm="anthropic:claude-haiku-4-5",
    llm_api_key=os.environ["ANTHROPIC_API_KEY"],
)

result = client.research.run(
    "What did the JWST discover about early galaxies in 2025?",
    search="tavily",           # or "brave", or a preconstructed WebSearchTool
    max_steps=10,
    compress_snippets=True,
    compression_model="{{SDK_DEFAULT_MODEL}}",
)

print(result.answer)
for c in result.citations:
    print(f"- {c.title or c.url}\n  {c.url}")
```

**TypeScript**

```typescript
import { CompressionClient } from '@compresr/sdk';

const client = new CompressionClient({
  apiKey: process.env.COMPRESR_API_KEY!,
  llm: 'anthropic:claude-haiku-4-5',
  llmApiKey: process.env.ANTHROPIC_API_KEY!,
});

const result = await client.research.run(
  'What did the JWST discover about early galaxies in 2025?',
  {
    search: 'tavily',          // or 'brave', or a preconstructed WebSearchTool
    maxSteps: 10,
    compressSnippets: true,
    compressionModel: '{{SDK_DEFAULT_MODEL}}',
  },
);

console.log(result.answer);
for (const c of result.citations) {
  console.log(`- ${c.title ?? c.url}\n  ${c.url}`);
}
```

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `search` | "tavily" \| "brave" \| BaseTool | no | Provider string (uses env-var fallbacks) or a preconstructed `WebSearchTool`. |
| `max_steps` | int | no | Upper bound on search / synthesize iterations. `.search()` overrides this to 2. |
| `model` | str \| None | no | Override the client-level model for this call. |
| `compress_snippets` | bool | no | Route each search snippet through the compression API before it enters the LLM context. |
| `compression_model` | str | no | Which model runs the snippet compression. |
| `min_compress_tokens` | int | no | Skip compression for snippets shorter than this many tokens. |
| `max_context_tokens` | int | no | Hard ceiling on total tokens across all compressed snippets before synthesis. |
| `system_prompt` | str \| None | no | Override the built-in system prompt (see DEFAULT_RESEARCH_SYSTEM_PROMPT). |

`ResearchResult` fields: `answer: str`, `explanation: str`, `confidence: float | None`, `text: str`, `citations: list[Citation]`, `trajectory: list[Step]`, `usage: ResearchUsage`, `raw: Any` (defaults to `None`; usually the provider's raw response object). `ResearchUsage` has int counters `input_tokens`, `output_tokens`, `cache_read_tokens`, `cache_creation_tokens`, `calls`, `search_calls`. Each `Citation` has `url: str`, `title: str | None`, `snippet: str | None`.

### Per-call LLM knobs

Forwarded to the underlying chat model: `temperature, top_p, top_k, max_tokens, max_output_tokens, stop, stop_sequences, presence_penalty, frequency_penalty, seed, logprobs, top_logprobs`. Anything else is silently dropped.

```python
client.messages.create(model="claude-haiku-4-5", max_tokens=512,
    temperature=0.2, top_p=0.9, messages=[...], tools=[...])
```

**TypeScript**

```typescript
await client.messages.create({ model: 'claude-haiku-4-5', maxTokens: 512,
  temperature: 0.2, topP: 0.9, messages: [...], tools: [...] });
```

> **Gemini aliasing**
> When `provider == "google_genai"` the SDK renames `max_tokens` → `max_output_tokens` automatically. Pass `max_tokens` from any provider; the SDK will do the right thing.

### Compression knobs: `compression={...}`

Set at client construction. Applies to every tool-output compression the middleware fires. The model-routing keys mirror [`compress()`](#3-compress) — `compression_model_name` picks the backbone, and the compression-shaping keys forward through to the same `/compress/question-specific/` endpoint.

**Shared keys (accepted regardless of `compression_model_name`):**

| Key | Default | Effect |
|---|---|---|
| `compression_model_name` | `"{{SDK_DEFAULT_MODEL}}"` | Backend validates; `"latte_v1"` and `"latte_v2"` are both public. See [Models](/docs/api-reference/models). |
| `target_compression_ratio` | `0.5` | 0–1 removal strength; `>1` = Nx factor (same as [`compress`](#3-compress) arg). Ignored on `latte_v2` when `dynamic=True`. |
| `min_tokens` | `200` | Tool outputs shorter than this skip compression. Middleware-side gate; not forwarded to the API. |
| `coarse` | server default (`True`) | Paragraph-level vs token-level. |
| `allow_tools` | `None` | Whitelist of tool names to compress. |
| `ignore_tools` | `None` | Blacklist of tool names to leave untouched. |
| `on_error` | `"passthrough"` | `"raise"` to fail loudly on backend errors instead of returning the original tool output. |

The middleware policy doesn't expose the `dynamic*` `latte_v2`-only knobs. If you need adaptive ratio selection on tool outputs, call `client.compress(...)` directly with `dynamic=True` instead of routing through the middleware.

## 7. Async

`compress_async` and `compress_batch_async` are the async twins of `compress` and `compress_batch`: same params, return awaitables. Streaming is sync-only (no `compress_stream_async`). Call `await client.aclose()` when done to release the `httpx` pool, or use the client as an async context manager (`async with CompressionClient(...) as client:`). Use these inside event loops (FastAPI handlers, Discord bots, agent runtimes); for scripts the sync methods are simpler.

```python
import asyncio
import os
from compresr import CompressionClient

async def main():
    client = CompressionClient(api_key=os.environ["COMPRESR_API_KEY"])
    try:
        result = await client.compress_async(
            context=long_document,
            query="What was the project's Q3 churn rate?",
            compression_model_name="{{DEFAULT_MODEL}}",
        )
        print(result.data.compressed_context)
    finally:
        await client.aclose()

asyncio.run(main())
```

**TypeScript**

```typescript
// The TypeScript SDK is already async. compress, compressStream,
// and compressBatch return Promise or AsyncGenerator - no _async variants.
const result = await client.compress({
  context: longDocument,
  query: "What was the project's Q3 churn rate?",
  compressionModelName: '{{DEFAULT_MODEL}}',
});
console.log(result.data.compressed_context);
```

**cURL**

```bash
# HTTP is already request/response - there is no async pattern specific to cURL.
# For incremental output, use the /stream endpoint (Section 4).
# For concurrency, fire requests in parallel with shell jobs:
curl ... &
curl ... &
wait
```

## 8. Errors & types

Every Compresr error inherits from `CompresrError`. Catch the base for a single handler; catch subclasses when recovery differs. Every subclass carries a stable `code` string and, where relevant, structured attributes you can branch on (e.g. `err.retry_after`, `err.credits_remaining`, `err.available_models`) instead of parsing prose.

| Exception | HTTP | `code` | Structured attributes |
|---|---|---|---|
| `AuthenticationError` | 401 | `authentication_error` | — |
| `ScopeError` | 403 | `scope_error` | `required_scope` |
| `NotFoundError` | 404 | `not_found` | `resource: str \| None` |
| `RateLimitError` | 429 | `rate_limit_exceeded` | `retry_after: int \| None` |
| `ValidationError` | 400 / 422 | `validation_error` | `field: str \| None` |
| `InsufficientCreditsError` | 402 | `insufficient_credits` | `credits_required`, `credits_remaining` |
| `BudgetLimitError` | 402 | `budget_limit_reached` | `current_budget`, `budget_used` |
| `ApiKeyBudgetError` | 402 | `api_key_budget_exceeded` | `api_key_budget`, `api_key_used` |
| `DailyLimitError` | 429 | `daily_limit_exceeded` | `daily_limit`, `requests_used` |
| `ModelNotFoundError` | 404 | `model_not_found` | `model_name`, `available_models` |
| `ContextWindowExceededError` | 413 | `context_window_exceeded` | `max_tokens`, `actual_tokens` |
| `ContentPolicyError` | 400 | `content_policy_violation` | `provider` |
| `TargetAuthenticationError` | 401 | `target_authentication_error` | `provider` |
| `ServiceUnavailableError` | 503 | `service_unavailable` | `service`, `retry_after: int \| None` |
| `ServerError` | 5xx | `server_error` | — |
| `CompresrTimeoutError` | — | `timeout` | `timeout_seconds: int \| None`. Reserved; the transport currently maps HTTP timeouts to `CompresrConnectionError("Request timed out")` — catch that class instead. |
| `CompresrConnectionError` | — | `connection_error` | `service: str \| None` |
| `CompresrError` | — | (varies) | base class; catch it as a fallback |

> **Fractional Retry-After**
> Servers occasionally emit `Retry-After: 1.5`. The SDK stores `retry_after` as `Optional[int]`, so fractional values are dropped (attribute reads as `None`). Guard with `time.sleep(err.retry_after or 1)`.

```python
import time
import logging
from compresr import (
    CompresrError,
    AuthenticationError,
    RateLimitError,
    ValidationError,
)

logger = logging.getLogger(__name__)

try:
    result = client.compress(
        context=long_document,
        query="What was the project's Q3 churn rate?",
        compression_model_name="{{DEFAULT_MODEL}}",
    )
except RateLimitError as e:
    time.sleep(e.retry_after or 1)
    # ...retry
except AuthenticationError:
    raise RuntimeError("COMPRESR_API_KEY is invalid or revoked")
except ValidationError as e:
    raise ValueError(f"Bad request: {e}") from e
except CompresrError as e:
    logger.exception("Compression failed: %s", e)
    raise
```

**TypeScript**

```typescript
import {
  CompresrError,
  AuthenticationError,
  RateLimitError,
  ValidationError,
} from '@compresr/sdk';

try {
  const result = await client.compress({
    context: longDocument,
    query: "What was the project's Q3 churn rate?",
    compressionModelName: '{{DEFAULT_MODEL}}',
  });
} catch (error: unknown) {
  if (error instanceof RateLimitError) {
    await new Promise((r) => setTimeout(r, (error.retryAfter ?? 1) * 1000));
    // ...retry
  } else if (error instanceof AuthenticationError) {
    throw new Error('COMPRESR_API_KEY is invalid or revoked');
  } else if (error instanceof ValidationError) {
    throw new Error(`Bad request: ${error.message}`);
  } else if (error instanceof CompresrError) {
    throw error;
  } else {
    throw error;
  }
}
```

**cURL**

```bash
# Branch on HTTP status. The response body always has a stable error envelope:
# { "success": false, "error": "...", "code": "..." }
status=$(curl -s -o body.json -D headers.txt -w "%{http_code}" \
  -X POST https://api.compresr.ai/api/compress/question-specific/ \
  -H "X-API-Key: $COMPRESR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"context":"...","query":"...","compression_model_name":"{{DEFAULT_MODEL}}"}')

if [ "$status" = "429" ]; then
  retry_after=$(grep -i '^Retry-After:' headers.txt | awk '{print $2}' | tr -d '\r')
  sleep "${retry_after:-1}"
  # ...retry
elif [ "$status" -ge 400 ]; then
  echo "Request failed ($status):" >&2
  cat body.json >&2
  exit 1
fi
```

> **Always handle 429**
> The default tier has tight per-minute limits. A retry loop with exponential backoff (respecting `retry_after`) is the single most important piece of error handling for production.
