0ef5fcb1c5
Security / Dependency audit (pip-audit) (push) Has been cancelled
Security / CodeQL (javascript-typescript) (push) Has been cancelled
Security / CodeQL (python) (push) Has been cancelled
Security / Secret scan (gitleaks) (push) Has been cancelled
rust / test (ubuntu) (push) Has been cancelled
rust / simulator e2e (macos-latest) (push) Has been cancelled
rust / simulator e2e (ubuntu-latest) (push) Has been cancelled
rust / simulator e2e (windows-latest) (push) Has been cancelled
rust / wheels (aarch64-apple-darwin) (push) Has been cancelled
rust / wheels (x86_64-unknown-linux-gnu) (push) Has been cancelled
rust / wheels (x86_64-apple-darwin) (push) Has been cancelled
rust / audit (push) Has been cancelled
rust / parity (nightly, allowed to fail during Phase 0) (push) Has been cancelled
CI / commitlint (push) Has been skipped
Dev Containers / validate (.devcontainer/devcontainer.json, default) (push) Failing after 0s
Dev Containers / validate (.devcontainer/memory-stack/devcontainer.json, memory-stack) (push) Failing after 0s
Dev Containers / validate-worktree (push) Failing after 0s
CI / changes (push) Failing after 4s
Deploy Documentation / validate (push) Has been skipped
Deploy Documentation / deploy (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, claude) (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, codex) (push) Failing after 1s
Install Native E2E / install-native (ubuntu-latest) (push) Failing after 1s
OpenCode Plugin / typecheck + build + test (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, copilot) (push) Failing after 1s
Release Please / release-please (push) Failing after 1s
Wrap E2E / docker-wrap-e2e (push) Failing after 1s
Wrap Native E2E / wrap-native (ubuntu-latest) (push) Failing after 1s
Init E2E / docker-init-e2e (push) Failing after 4s
Merge Conflicts / merge-conflicts (push) Failing after 4s
CI / lint (push) Has been cancelled
CI / build-wheel (push) Has been cancelled
CI / build-wheel-windows (push) Has been cancelled
CI / prefetch-model (push) Has been cancelled
CI / test-dashboard-ui (push) Has been cancelled
CI / test (1) (push) Has been cancelled
CI / test (2) (push) Has been cancelled
CI / test (3) (push) Has been cancelled
CI / test (4) (push) Has been cancelled
CI / test-extras (push) Has been cancelled
CI / test-agno (push) Has been cancelled
CI / build (push) Has been cancelled
CI / workflow-validation (push) Has been cancelled
CI / docker-native-e2e (push) Has been cancelled
CI / windows-native-wrapper (push) Has been cancelled
CI / macos-native-wrapper (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / promote-latest (push) Has been cancelled
Init Native E2E / init-native (macos-latest, claude) (push) Has been cancelled
Init Native E2E / init-native (macos-latest, codex) (push) Has been cancelled
Init Native E2E / init-native (macos-latest, copilot) (push) Has been cancelled
Install Native E2E / install-native (macos-latest) (push) Has been cancelled
Wrap Native E2E / wrap-native (macos-latest) (push) Has been cancelled
319 lines
7.5 KiB
Markdown
319 lines
7.5 KiB
Markdown
# Image Compression
|
||
|
||
Headroom automatically compresses images in your LLM requests, reducing token usage by **40-90%** while maintaining answer accuracy.
|
||
|
||
## Overview
|
||
|
||
Vision models charge by the token, and images are expensive:
|
||
- A 1024x1024 image costs ~765 tokens (OpenAI)
|
||
- A 2048x2048 image costs ~2,900 tokens
|
||
|
||
Headroom's image compression uses a **trained ML router** to analyze your query and automatically select the optimal compression technique:
|
||
|
||
| Technique | Savings | When Used |
|
||
|-----------|---------|-----------|
|
||
| `full_low` | ~87% | General questions ("What is this?") |
|
||
| `preserve` | 0% | Fine details needed ("Count the whiskers") |
|
||
| `crop` | 50-90% | Region-specific ("What's in the corner?") |
|
||
| `transcode` | ~99% | Text extraction ("Read the sign") |
|
||
|
||
## How It Works
|
||
|
||
```
|
||
User uploads image + asks question
|
||
↓
|
||
[Query Analysis]
|
||
TrainedRouter (MiniLM from HuggingFace)
|
||
Classifies: "What animal is this?" → full_low
|
||
↓
|
||
[Image Analysis]
|
||
SigLIP analyzes image properties
|
||
(has text? complex? fine details?)
|
||
↓
|
||
[Apply Compression]
|
||
OpenAI: detail="low"
|
||
Anthropic: Resize to 512px
|
||
Google: Resize to 768px
|
||
↓
|
||
Compressed request to LLM
|
||
```
|
||
|
||
## Quick Start
|
||
|
||
### With Headroom Proxy (Zero Code Changes)
|
||
|
||
```bash
|
||
# Start the proxy
|
||
headroom proxy --port 8787
|
||
|
||
# Connect your client
|
||
ANTHROPIC_BASE_URL=http://localhost:8787 claude
|
||
```
|
||
|
||
Images are automatically compressed based on your queries.
|
||
|
||
### With HeadroomClient
|
||
|
||
```python
|
||
from headroom import HeadroomClient
|
||
|
||
client = HeadroomClient(provider="openai")
|
||
|
||
response = client.chat.completions.create(
|
||
model="gpt-4o",
|
||
messages=[{
|
||
"role": "user",
|
||
"content": [
|
||
{"type": "text", "text": "What animal is this?"},
|
||
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
|
||
]
|
||
}]
|
||
)
|
||
# Image automatically compressed with detail="low" (87% savings)
|
||
```
|
||
|
||
### Direct API
|
||
|
||
```python
|
||
from headroom.image import ImageCompressor
|
||
|
||
compressor = ImageCompressor()
|
||
|
||
# Compress images in messages
|
||
compressed_messages = compressor.compress(messages, provider="openai")
|
||
|
||
# Check savings
|
||
print(f"Saved {compressor.last_savings:.0f}% tokens")
|
||
print(f"Technique: {compressor.last_result.technique.value}")
|
||
```
|
||
|
||
## Configuration
|
||
|
||
### Proxy Configuration
|
||
|
||
```bash
|
||
# Enable image compression (default: true)
|
||
headroom proxy --image-optimize
|
||
|
||
# Disable image compression
|
||
headroom proxy --no-image-optimize
|
||
```
|
||
|
||
### Programmatic Configuration
|
||
|
||
```python
|
||
from headroom.image import ImageCompressor
|
||
|
||
compressor = ImageCompressor(
|
||
model_id="chopratejas/technique-router", # HuggingFace model
|
||
use_siglip=True, # Enable image analysis
|
||
device="cuda", # Use GPU if available
|
||
)
|
||
```
|
||
|
||
## Provider Support
|
||
|
||
| Provider | Detection | Compression Method |
|
||
|----------|-----------|-------------------|
|
||
| **OpenAI** | `image_url` | Sets `detail="low"` |
|
||
| **Anthropic** | `image` with `source` | Resizes to 512px |
|
||
| **Google** | `inlineData` | Resizes to 768px (tile-optimized) |
|
||
|
||
### OpenAI
|
||
|
||
Uses the native `detail` parameter:
|
||
```python
|
||
# Before
|
||
{"type": "image_url", "image_url": {"url": "data:..."}}
|
||
|
||
# After (full_low technique)
|
||
{"type": "image_url", "image_url": {"url": "data:...", "detail": "low"}}
|
||
```
|
||
|
||
### Anthropic
|
||
|
||
Resizes the image using PIL:
|
||
```python
|
||
# Before: 1024x1024 image (~1,398 tokens)
|
||
# After: 512x512 image (~349 tokens) - 75% savings
|
||
```
|
||
|
||
### Google Gemini
|
||
|
||
Resizes to 768px (optimal for Gemini's 768x768 tile system):
|
||
```python
|
||
# Before: 1536x1536 image (4 tiles × 258 = 1,032 tokens)
|
||
# After: 768x768 image (1 tile × 258 = 258 tokens) - 75% savings
|
||
```
|
||
|
||
## Techniques Explained
|
||
|
||
### `full_low` (87% savings)
|
||
|
||
Best for general understanding questions:
|
||
- "What is this?"
|
||
- "Describe the scene"
|
||
- "Is this indoors or outdoors?"
|
||
|
||
The model doesn't need fine details to answer these questions.
|
||
|
||
### `preserve` (0% savings)
|
||
|
||
Required when fine details matter:
|
||
- "Count the whiskers"
|
||
- "What brand is shown?"
|
||
- "Read the serial number"
|
||
- "What time does the clock show?"
|
||
|
||
### `crop` (50-90% savings)
|
||
|
||
For region-specific queries:
|
||
- "What's in the top-right corner?"
|
||
- "Focus on the background"
|
||
- "Zoom into the left side"
|
||
|
||
*Note: Currently implemented as resize. True cropping coming soon.*
|
||
|
||
### `transcode` (99% savings)
|
||
|
||
For text extraction (converts image to text):
|
||
- "Read the sign"
|
||
- "What does it say?"
|
||
- "Transcribe the document"
|
||
|
||
*Note: Requires vision model call. Currently falls back to preserve.*
|
||
|
||
## The Trained Router
|
||
|
||
The routing decision is made by a fine-tuned **MiniLM** classifier:
|
||
|
||
- **Model**: `chopratejas/technique-router` on HuggingFace
|
||
- **Size**: ~128MB
|
||
- **Accuracy**: 93.7% on validation set
|
||
- **Training data**: 1,157 examples across 4 techniques
|
||
|
||
The model is downloaded automatically on first use and cached locally.
|
||
|
||
### Training Data Examples
|
||
|
||
| Query | Technique |
|
||
|-------|-----------|
|
||
| "What animal is this?" | `full_low` |
|
||
| "Count the spots" | `preserve` |
|
||
| "Read the text on the sign" | `transcode` |
|
||
| "What's in the corner?" | `crop` |
|
||
|
||
## Performance
|
||
|
||
### Token Savings by Query Type
|
||
|
||
| Query Type | Before | After | Savings |
|
||
|------------|--------|-------|---------|
|
||
| General ("What is this?") | 765 | 85 | 89% |
|
||
| Detail ("Count items") | 765 | 765 | 0% |
|
||
| Region ("Top corner?") | 765 | 85 | 89% |
|
||
| Text ("Read the sign") | 765 | 85 | 89% |
|
||
|
||
### Latency
|
||
|
||
- Router inference: ~10ms (CPU), ~2ms (GPU)
|
||
- Image resize: ~5-20ms depending on size
|
||
- First request: +2-3s (model download, cached after)
|
||
|
||
## Troubleshooting
|
||
|
||
### Model Download Issues
|
||
|
||
The HuggingFace model downloads on first use:
|
||
|
||
```python
|
||
# Force a specific cache directory
|
||
import os
|
||
os.environ["HF_HOME"] = "/path/to/cache"
|
||
|
||
from headroom.image import ImageCompressor
|
||
compressor = ImageCompressor()
|
||
```
|
||
|
||
### GPU Memory
|
||
|
||
SigLIP requires ~400MB GPU memory. To use CPU only:
|
||
|
||
```python
|
||
compressor = ImageCompressor(device="cpu")
|
||
```
|
||
|
||
### Disable Image Compression
|
||
|
||
```python
|
||
# Proxy
|
||
headroom proxy --no-image-optimize
|
||
|
||
# Direct
|
||
# Simply don't call compress()
|
||
```
|
||
|
||
## API Reference
|
||
|
||
### `ImageCompressor`
|
||
|
||
```python
|
||
class ImageCompressor:
|
||
def __init__(
|
||
self,
|
||
model_id: str = "chopratejas/technique-router",
|
||
use_siglip: bool = True,
|
||
device: str | None = None,
|
||
): ...
|
||
|
||
def has_images(self, messages: list[dict]) -> bool:
|
||
"""Check if messages contain images."""
|
||
|
||
def compress(
|
||
self,
|
||
messages: list[dict],
|
||
provider: str = "openai",
|
||
) -> list[dict]:
|
||
"""Compress images in messages."""
|
||
|
||
@property
|
||
def last_result(self) -> CompressionResult | None:
|
||
"""Result of last compression."""
|
||
|
||
@property
|
||
def last_savings(self) -> float:
|
||
"""Savings percentage from last compression."""
|
||
```
|
||
|
||
### `CompressionResult`
|
||
|
||
```python
|
||
@dataclass
|
||
class CompressionResult:
|
||
technique: Technique # full_low, preserve, crop, transcode
|
||
original_tokens: int # Estimated tokens before
|
||
compressed_tokens: int # Estimated tokens after
|
||
confidence: float # Router confidence (0-1)
|
||
|
||
@property
|
||
def savings_percent(self) -> float:
|
||
"""Percentage of tokens saved."""
|
||
```
|
||
|
||
### `Technique`
|
||
|
||
```python
|
||
class Technique(Enum):
|
||
FULL_LOW = "full_low" # 87% savings
|
||
PRESERVE = "preserve" # 0% savings
|
||
CROP = "crop" # 50-90% savings
|
||
TRANSCODE = "transcode" # 99% savings
|
||
```
|
||
|
||
## See Also
|
||
|
||
- [Compression Guide](compression.md) - Text compression techniques
|
||
- [CCR Guide](ccr.md) - Reversible compression with retrieval
|
||
- [Proxy Guide](proxy.md) - Zero-code deployment
|
||
- [Architecture](ARCHITECTURE.md) - System design
|