Files
chopratejas--headroom/wiki/image-compression.md
T
wehub-resource-sync 0ef5fcb1c5
Security / Dependency audit (pip-audit) (push) Has been cancelled
Security / CodeQL (javascript-typescript) (push) Has been cancelled
Security / CodeQL (python) (push) Has been cancelled
Security / Secret scan (gitleaks) (push) Has been cancelled
rust / test (ubuntu) (push) Has been cancelled
rust / simulator e2e (macos-latest) (push) Has been cancelled
rust / simulator e2e (ubuntu-latest) (push) Has been cancelled
rust / simulator e2e (windows-latest) (push) Has been cancelled
rust / wheels (aarch64-apple-darwin) (push) Has been cancelled
rust / wheels (x86_64-unknown-linux-gnu) (push) Has been cancelled
rust / wheels (x86_64-apple-darwin) (push) Has been cancelled
rust / audit (push) Has been cancelled
rust / parity (nightly, allowed to fail during Phase 0) (push) Has been cancelled
CI / commitlint (push) Has been skipped
Dev Containers / validate (.devcontainer/devcontainer.json, default) (push) Failing after 0s
Dev Containers / validate (.devcontainer/memory-stack/devcontainer.json, memory-stack) (push) Failing after 0s
Dev Containers / validate-worktree (push) Failing after 0s
CI / changes (push) Failing after 4s
Deploy Documentation / validate (push) Has been skipped
Deploy Documentation / deploy (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, claude) (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, codex) (push) Failing after 1s
Install Native E2E / install-native (ubuntu-latest) (push) Failing after 1s
OpenCode Plugin / typecheck + build + test (push) Failing after 1s
Init Native E2E / init-native (ubuntu-latest, copilot) (push) Failing after 1s
Release Please / release-please (push) Failing after 1s
Wrap E2E / docker-wrap-e2e (push) Failing after 1s
Wrap Native E2E / wrap-native (ubuntu-latest) (push) Failing after 1s
Init E2E / docker-init-e2e (push) Failing after 4s
Merge Conflicts / merge-conflicts (push) Failing after 4s
CI / lint (push) Has been cancelled
CI / build-wheel (push) Has been cancelled
CI / build-wheel-windows (push) Has been cancelled
CI / prefetch-model (push) Has been cancelled
CI / test-dashboard-ui (push) Has been cancelled
CI / test (1) (push) Has been cancelled
CI / test (2) (push) Has been cancelled
CI / test (3) (push) Has been cancelled
CI / test (4) (push) Has been cancelled
CI / test-extras (push) Has been cancelled
CI / test-agno (push) Has been cancelled
CI / build (push) Has been cancelled
CI / workflow-validation (push) Has been cancelled
CI / docker-native-e2e (push) Has been cancelled
CI / windows-native-wrapper (push) Has been cancelled
CI / macos-native-wrapper (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-manifest (map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-nonroot name:code-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-slim name:code-slim]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-code-slim-nonroot name:code-slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-nonroot name:nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-slim name:slim]) (push) Has been cancelled
Docker / docker-build (map[name:amd64 platform:linux/amd64 runs_on:ubuntu-24.04], map[bake_target:runtime-slim-nonroot name:slim-nonroot]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime name:]) (push) Has been cancelled
Docker / docker-build (map[name:arm64 platform:linux/arm64 runs_on:ubuntu-24.04-arm], map[bake_target:runtime-code name:code]) (push) Has been cancelled
Docker / promote-latest (push) Has been cancelled
Init Native E2E / init-native (macos-latest, claude) (push) Has been cancelled
Init Native E2E / init-native (macos-latest, codex) (push) Has been cancelled
Init Native E2E / init-native (macos-latest, copilot) (push) Has been cancelled
Install Native E2E / install-native (macos-latest) (push) Has been cancelled
Wrap Native E2E / wrap-native (macos-latest) (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:03:20 +08:00

319 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Image Compression
Headroom automatically compresses images in your LLM requests, reducing token usage by **40-90%** while maintaining answer accuracy.
## Overview
Vision models charge by the token, and images are expensive:
- A 1024x1024 image costs ~765 tokens (OpenAI)
- A 2048x2048 image costs ~2,900 tokens
Headroom's image compression uses a **trained ML router** to analyze your query and automatically select the optimal compression technique:
| Technique | Savings | When Used |
|-----------|---------|-----------|
| `full_low` | ~87% | General questions ("What is this?") |
| `preserve` | 0% | Fine details needed ("Count the whiskers") |
| `crop` | 50-90% | Region-specific ("What's in the corner?") |
| `transcode` | ~99% | Text extraction ("Read the sign") |
## How It Works
```
User uploads image + asks question
[Query Analysis]
TrainedRouter (MiniLM from HuggingFace)
Classifies: "What animal is this?" → full_low
[Image Analysis]
SigLIP analyzes image properties
(has text? complex? fine details?)
[Apply Compression]
OpenAI: detail="low"
Anthropic: Resize to 512px
Google: Resize to 768px
Compressed request to LLM
```
## Quick Start
### With Headroom Proxy (Zero Code Changes)
```bash
# Start the proxy
headroom proxy --port 8787
# Connect your client
ANTHROPIC_BASE_URL=http://localhost:8787 claude
```
Images are automatically compressed based on your queries.
### With HeadroomClient
```python
from headroom import HeadroomClient
client = HeadroomClient(provider="openai")
response = client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What animal is this?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
]
}]
)
# Image automatically compressed with detail="low" (87% savings)
```
### Direct API
```python
from headroom.image import ImageCompressor
compressor = ImageCompressor()
# Compress images in messages
compressed_messages = compressor.compress(messages, provider="openai")
# Check savings
print(f"Saved {compressor.last_savings:.0f}% tokens")
print(f"Technique: {compressor.last_result.technique.value}")
```
## Configuration
### Proxy Configuration
```bash
# Enable image compression (default: true)
headroom proxy --image-optimize
# Disable image compression
headroom proxy --no-image-optimize
```
### Programmatic Configuration
```python
from headroom.image import ImageCompressor
compressor = ImageCompressor(
model_id="chopratejas/technique-router", # HuggingFace model
use_siglip=True, # Enable image analysis
device="cuda", # Use GPU if available
)
```
## Provider Support
| Provider | Detection | Compression Method |
|----------|-----------|-------------------|
| **OpenAI** | `image_url` | Sets `detail="low"` |
| **Anthropic** | `image` with `source` | Resizes to 512px |
| **Google** | `inlineData` | Resizes to 768px (tile-optimized) |
### OpenAI
Uses the native `detail` parameter:
```python
# Before
{"type": "image_url", "image_url": {"url": "data:..."}}
# After (full_low technique)
{"type": "image_url", "image_url": {"url": "data:...", "detail": "low"}}
```
### Anthropic
Resizes the image using PIL:
```python
# Before: 1024x1024 image (~1,398 tokens)
# After: 512x512 image (~349 tokens) - 75% savings
```
### Google Gemini
Resizes to 768px (optimal for Gemini's 768x768 tile system):
```python
# Before: 1536x1536 image (4 tiles × 258 = 1,032 tokens)
# After: 768x768 image (1 tile × 258 = 258 tokens) - 75% savings
```
## Techniques Explained
### `full_low` (87% savings)
Best for general understanding questions:
- "What is this?"
- "Describe the scene"
- "Is this indoors or outdoors?"
The model doesn't need fine details to answer these questions.
### `preserve` (0% savings)
Required when fine details matter:
- "Count the whiskers"
- "What brand is shown?"
- "Read the serial number"
- "What time does the clock show?"
### `crop` (50-90% savings)
For region-specific queries:
- "What's in the top-right corner?"
- "Focus on the background"
- "Zoom into the left side"
*Note: Currently implemented as resize. True cropping coming soon.*
### `transcode` (99% savings)
For text extraction (converts image to text):
- "Read the sign"
- "What does it say?"
- "Transcribe the document"
*Note: Requires vision model call. Currently falls back to preserve.*
## The Trained Router
The routing decision is made by a fine-tuned **MiniLM** classifier:
- **Model**: `chopratejas/technique-router` on HuggingFace
- **Size**: ~128MB
- **Accuracy**: 93.7% on validation set
- **Training data**: 1,157 examples across 4 techniques
The model is downloaded automatically on first use and cached locally.
### Training Data Examples
| Query | Technique |
|-------|-----------|
| "What animal is this?" | `full_low` |
| "Count the spots" | `preserve` |
| "Read the text on the sign" | `transcode` |
| "What's in the corner?" | `crop` |
## Performance
### Token Savings by Query Type
| Query Type | Before | After | Savings |
|------------|--------|-------|---------|
| General ("What is this?") | 765 | 85 | 89% |
| Detail ("Count items") | 765 | 765 | 0% |
| Region ("Top corner?") | 765 | 85 | 89% |
| Text ("Read the sign") | 765 | 85 | 89% |
### Latency
- Router inference: ~10ms (CPU), ~2ms (GPU)
- Image resize: ~5-20ms depending on size
- First request: +2-3s (model download, cached after)
## Troubleshooting
### Model Download Issues
The HuggingFace model downloads on first use:
```python
# Force a specific cache directory
import os
os.environ["HF_HOME"] = "/path/to/cache"
from headroom.image import ImageCompressor
compressor = ImageCompressor()
```
### GPU Memory
SigLIP requires ~400MB GPU memory. To use CPU only:
```python
compressor = ImageCompressor(device="cpu")
```
### Disable Image Compression
```python
# Proxy
headroom proxy --no-image-optimize
# Direct
# Simply don't call compress()
```
## API Reference
### `ImageCompressor`
```python
class ImageCompressor:
def __init__(
self,
model_id: str = "chopratejas/technique-router",
use_siglip: bool = True,
device: str | None = None,
): ...
def has_images(self, messages: list[dict]) -> bool:
"""Check if messages contain images."""
def compress(
self,
messages: list[dict],
provider: str = "openai",
) -> list[dict]:
"""Compress images in messages."""
@property
def last_result(self) -> CompressionResult | None:
"""Result of last compression."""
@property
def last_savings(self) -> float:
"""Savings percentage from last compression."""
```
### `CompressionResult`
```python
@dataclass
class CompressionResult:
technique: Technique # full_low, preserve, crop, transcode
original_tokens: int # Estimated tokens before
compressed_tokens: int # Estimated tokens after
confidence: float # Router confidence (0-1)
@property
def savings_percent(self) -> float:
"""Percentage of tokens saved."""
```
### `Technique`
```python
class Technique(Enum):
FULL_LOW = "full_low" # 87% savings
PRESERVE = "preserve" # 0% savings
CROP = "crop" # 50-90% savings
TRANSCODE = "transcode" # 99% savings
```
## See Also
- [Compression Guide](compression.md) - Text compression techniques
- [CCR Guide](ccr.md) - Reversible compression with retrieval
- [Proxy Guide](proxy.md) - Zero-code deployment
- [Architecture](ARCHITECTURE.md) - System design