Files
yvgude--lean-ctx/docs/guides/compress-sdk.md
T
wehub-resource-sync 26382a7ac6
CI / Clippy (push) Failing after 15m13s
CI / Test (ubuntu-latest) (push) Failing after 16m1s
CI / Test (macos-latest) (push) Has been cancelled
CI / Test (windows-latest) (push) Has been cancelled
CI / Build (no embeddings / no ORT) (push) Has been cancelled
CI / Format (push) Has been cancelled
CI / Cookbook (Node) (push) Has been cancelled
CI / Pi Extension (Node) (push) Has been cancelled
CI / Rust SDK (lean-ctx-client) (push) Has been cancelled
CI / Embed SDK (lean-ctx-sdk) (push) Has been cancelled
CI / Python SDK (leanctx) (push) Has been cancelled
CI / Hermes Plugin (Python) (push) Has been cancelled
CI / SDK Conformance Matrix (push) Has been cancelled
CI / Coverage (push) Has been cancelled
CI / cargo-deny (push) Has been cancelled
CI / Adversarial Safety (push) Has been cancelled
CI / Benchmarks (push) Has been cancelled
CI / Output-Quality Gate (eval A/B) (push) Has been cancelled
CI / Documentation (push) Has been cancelled
CI / CI Green (push) Has been cancelled
JetBrains Plugin / Actionlint (push) Has been cancelled
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (rust) (push) Has been cancelled
JetBrains Plugin / Validation (push) Has been cancelled
JetBrains Plugin / Build (push) Has been cancelled
JetBrains Plugin / Test (push) Has been cancelled
Security Check / Security Scan (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:35:30 +08:00

209 lines
6.8 KiB
Markdown

# compress() SDK Cookbook (Python + TypeScript)
> Drop-in context compression for any LLM app. `compress(messages, model)` sends
> a chat-style array to the local lean-ctx daemon's deterministic
> [`POST /v1/compress`](../contracts/http-mcp-contract-v1.md) endpoint and returns
> the rewritten messages — byte-stable, so provider prompt caching keeps working.
Only **text payloads** are rewritten through lean-ctx's deterministic funnel;
images, `tool_use`/`tool_call` blocks and ids pass through untouched. lean-ctx's
own `ctx_*` tool results are left verbatim (they are already compressed).
## Install
```bash
pip install lean-ctx-sdk # Python ≥ 3.9
npm install lean-ctx-sdk # Node ≥ 18
```
Both SDKs talk to a running daemon — start it once with `lean-ctx proxy enable`.
## 1. Drop-in compress
```python
# Python
from lean_ctx import compress
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": large_log_or_file_dump},
]
messages = compress(messages, model="claude-sonnet-4")
# → send `messages` to your provider as usual
```
```ts
// TypeScript
import { compress } from "lean-ctx-sdk";
let messages = [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: largeLogOrFileDump },
];
messages = await compress(messages, { model: "claude-sonnet-4" });
```
## 2. Token-savings stats
Use the client directly to read the savings reported by the daemon:
```python
from lean_ctx import ProxyClient
result = ProxyClient().compress(messages, model="gpt-4o")
print(result.saved_tokens, result.saved_pct) # e.g. 11979 17.2
messages = result.messages
```
```ts
import { ProxyClient } from "lean-ctx-sdk";
const result = await new ProxyClient().compress(messages, "gpt-4o");
console.log(result.stats.saved_tokens, result.stats.saved_pct);
messages = result.messages;
```
## 3. Vercel AI SDK middleware (TypeScript)
Compress every prompt automatically — no per-call changes:
```ts
import { wrapLanguageModel } from "ai";
import { openai } from "@ai-sdk/openai";
import { leanCtxMiddleware } from "lean-ctx-sdk";
const model = wrapLanguageModel({
model: openai("gpt-4o"),
middleware: leanCtxMiddleware({ model: "gpt-4o" }),
});
// generateText / streamText now send compressed prompts
```
`withLeanCtx(openai("gpt-4o"))` is a one-line shortcut. A compaction failure
(proxy down, auth, malformed) never breaks the generation — the original,
uncompressed prompt is sent instead.
## 4. LiteLLM (Python)
```python
import litellm
from lean_ctx import LeanCtxLiteLLMHandler
litellm.callbacks = [LeanCtxLiteLLMHandler(model="gpt-4o")]
# every completion now sends compressed messages
```
For non-LiteLLM code, `compress_request_data(data)` rewrites the `messages` of
any OpenAI-style request dict in place.
### LiteLLM proxy guardrail (zero-code, gateway-side)
LiteLLM ≥ v1.92 ships a native prompt-compression guardrail that calls a
sidecar's `POST {api_base}/v1/compress` during `pre_call` and swaps in the
returned `messages`. lean-ctx's `/v1/compress` speaks that wire contract
(request `{"messages": [...], "model": "..."}`; response `messages` +
`tokens_before`/`tokens_after`/`compression_ratio`, #700), so the lean-ctx
daemon can be the compression sidecar — no client change, works for every
model behind the gateway, including Claude Code via `ANTHROPIC_BASE_URL`:
```yaml
# litellm config.yaml
guardrails:
- guardrail_name: prompt-compression
litellm_params:
guardrail: headroom # LiteLLM's generic compress-sidecar hook
mode: pre_call
api_base: http://127.0.0.1:<lean-ctx-proxy-port>
api_key: <lean-ctx proxy token> # sent as Bearer; see `lean-ctx proxy token`
default_on: true
```
The guardrail's **CCR agentic loop** works against lean-ctx too (#702): when a
rewrite is lossy, the compressed text carries a `hash=<24-hex>` retrieval
marker (the exact `hash=([a-f0-9]{24})` shape LiteLLM scans for). LiteLLM then
injects its retrieve tool, and when the model asks for the original, LiteLLM
calls `GET {api_base}/v1/retrieve/{hash}` — lean-ctx resolves the hash against
the same content-addressed tee store that backs `ctx_expand`, and returns
`{"original_content": "..."}`. Compression stays reversible end-to-end through
the gateway, with no lean-ctx-specific client code.
Because lean-ctx's output is deterministic (#498), the compressed prefix stays
byte-stable across turns — provider prompt caching keeps working even behind
the gateway. Attach the guardrail to a virtual key to A/B compression per
developer; the `x-litellm-applied-guardrails` response header confirms it ran.
## 5. LangChain (Python)
```python
from langchain_core.messages import HumanMessage, SystemMessage
from lean_ctx import compress_messages
messages = compress_messages(
[
SystemMessage(content="You are a helpful assistant."),
HumanMessage(content=large_log_or_file_dump),
],
model="gpt-4o",
)
```
Message types and metadata are preserved (only `content` is rewritten).
## 6. Reference retrieval (reversibility)
When lean-ctx omits an oversized payload it leaves a durable reference id. Fetch
the original back on demand:
```python
from lean_ctx import ProxyClient
original = ProxyClient().resolve_reference("ref_abc123")
```
```ts
import { ProxyClient } from "lean-ctx-sdk";
const original = await new ProxyClient().resolveReference("ref_abc123");
```
## Configuration
The endpoint and session token are auto-discovered from the running daemon.
Every step is overridable:
| Setting | Env var | Default |
| --- | --- | --- |
| Proxy URL | `LEAN_CTX_PROXY_URL` | `http://127.0.0.1:<port>` |
| Proxy port | `LEAN_CTX_PROXY_PORT` | `config.toml` `proxy_port`, else UID-derived |
| Session token | `LEAN_CTX_PROXY_TOKEN` | `<data_dir>/session_token` |
```python
compress(messages, base_url="http://127.0.0.1:4444", token="…")
```
```ts
await compress(messages, { baseUrl: "http://127.0.0.1:4444", token: "…" });
```
If the daemon is down, `compress()` raises/rejects with `LeanCtxConnectionError`;
an unauthenticated request raises `LeanCtxAuthError`. Both extend `LeanCtxError`.
## Determinism (#498)
`/v1/compress` output is a pure function of `(messages, model)` — the same input
yields byte-identical output. Savings are reported in `stats`, never injected
into message bodies, so compressed prompts stay friendly to provider prompt
caching (Anthropic 90% / OpenAI 50% cached-token discounts). This is guarded by a
regression test (`proxy::compress_api::tests::determinism_regression_full_conversation_498`).
## Benchmark
Reproduce a head-to-head ratio + latency report (lean-ctx vs Headroom) over a
real corpus — see [`bench/compress/README.md`](../../bench/compress/README.md):
```bash
python bench/compress/benchmark.py --json
```
See also: [lean-ctx vs Headroom](../comparisons/vs-headroom.md).