7.9 KiB
Embedded Local Tiny-Model Experiments
This document summarizes the experiments behind the optional local tiny-model paths for
session-title generation (providers.tinyModel), Mnemopi memory extraction/consolidation
(providers.memoryModel), and the auto thinking-level difficulty classifier
(providers.autoThinkingModel, which reuses the memory-model registry). It is a factual engineering
record for maintainers: what we measured, which recipes won, and which models we shipped. All three
settings default to online, so existing users incur no downloads or on-device inference cost unless
they opt in.
Runtime / environment findings
- Stack:
@huggingface/transformers(transformers.js) v4 running under Bun. In Bun the library loads the nativeonnxruntime-nodebackend (not the WASM build). - Device policy: local tiny models default to CPU-only inference and retry once on CPU if an
explicit accelerated provider cannot initialize.
- Pick a provider persistently with the
providers.tinyModelDevicesetting (defaultkeeps CPU), or per-run with thePI_TINY_DEVICEenv var (which overrides the setting). - Accepted values are
cpu,gpu,metal/webgpu,auto,cuda,dml,coreml,wasm,webnn,webnn-gpu,webnn-cpu, andwebnn-npu. - Direct
coremlremains opt-in viaPI_TINY_DEVICE=coreml; it is not part of the default because cached decoder-LLM ONNX loads can fail during session initialization. - WebGPU/Metal works for the single-process eval harness, but the production worker forces
Darwin
gpu/webgpu/autorequests back to CPU because ONNX Runtime/Bun currently hard-crashes on worker teardown after WebGPU inference. - Use
providers.tinyModelDeviceorPI_TINY_DEVICEonly when explicitly opting out of the CPU default.
- Pick a provider persistently with the
- Quantization: q4 is the sweet spot — smaller on disk, faster to load, and fast at inference.
q8/int8 loads slower and infers slower on CPU. Every shipped model defaults to
q4; override the precision persistently with theproviders.tinyModelDtypesetting (defaultkeepsq4, e.g.fp16for higher fidelity), or per-run withPI_TINY_DTYPE(which overrides the setting). Acceptsauto,fp32,fp16,q8,int8,uint8,q4,bnb4,q4f16,q2,q2f16,q1,q1f16; an unrecognized value fails loudly at worker startup. - Load-time correction (important). An earlier belief that "q4 >=1B models take minutes to load"
was a measurement artifact caused by running ~5 multi-GB HuggingFace downloads in parallel
(I/O saturation). Clean, isolated warm loads are all sub-3s:
- TinyLlama-1.1B q4: ~0.5s
- Llama-3.2-1B q4: ~2.8s (
graphOpt=all) / ~0.5s (disabled) - LFM2-1.2B q4: ~0.36s
- Qwen2.5-1.5B q4: ~1.5s
- Qwen3-1.7B q4: ~1.6s
- gemma-3-1b q4: ~1.1s
- Conclusion: 1B–1.7B models are viable on CPU.
session_options.graphOptimizationLeveltrades load vs inference speed:disabled= fastest load, slightly slower inference;all= default.- First run downloads weights from the HF Hub to a cache dir (q4 weights ~200MB–1.1GB depending on model); subsequent warm loads are sub-second to ~3s. Inference is async and background-friendly for memory tasks; titles are semi-interactive.
Task 1: Session title generation (providers.tinyModel)
Task: turn the first user message into a 3–6 word title. Tiny models (sub-1B) suffice.
Winning recipe:
- Plain system prompt (no few-shot).
- Prefill the assistant turn with
<title>and stop at</title>, then take the first line. - Greedy decoding (
do_sample:false),enable_thinking:falsein the chat template.
What we learned:
- Few-shot examples HURT sub-0.6B models for titles; the tag-prefill rescues even 270M models.
- Token biasing (
bad_words_ids) is a confirmed no-op here — the prefill already controls the opener.
Leaderboard (tag trick, CPU, warm):
| Model | Verdict |
|---|---|
| LFM2-350M | Best speed/quality balance (~212MB) |
| Qwen3-0.6B | Most robust |
| gemma-3-270m | Smallest viable |
| Qwen2.5-0.5B | Acceptable |
| SmolLM2-135M | Too small |
| flan-t5-small | Rejected — just echoes the input |
Shipped local options: lfm2-350m, qwen3-0.6b, gemma-270m, qwen2.5-0.5b, lfm2-700m.
Default: online (pi/smol).
Task 2: Mnemopi memory (providers.memoryModel)
Mnemopi runs two small-LLM tasks:
- Extraction — pull durable, structured items from a single message.
- Consolidation — summarize a list of memories into 1–3 faithful sentences.
These need bigger models than titles: 1B–1.7B. We tested LFM2-1.2B, Qwen2.5-1.5B, Qwen3-1.7B, and gemma-3-1b (q4, CPU) via four parallel agents each running 27–31 experiments.
Extraction findings
The stock 5-category JSON prompt fails on small models in two ways:
- The all-empty example
{"facts":[],...}gets copied verbatim → 0 facts extracted. - Capable models emit JSON objects inside arrays, which Mnemopi's
String(item)coerces into the literal string[object Object].
The robust fix is a one-item-per-line output format (consumed by Mnemopi's parser line-fallback) or a flat JSON array of strings. Every model also over-extracts pure small talk; an explicit chit-chat → NONE example is the best mitigation.
Technique polarity flips vs titles
- At 1B+, few-shot is the dominant quality lever: e.g. Qwen2.5-1.5B extraction F1 0.52 → 0.83 going 1 → 3 shots; gemma recall 0.65 → 0.92 with 2 shots.
- Prefill HURTS extraction — it forces output on small talk, producing false positives.
- System-split (instructions in the system role) helps models that have a system role.
- Greedy >= temperature for both tasks.
- Token biasing is again a no-op.
Per-model verdicts (head-to-head, 16-fixture set)
- Qwen3-1.7B — most disciplined extraction: returns empty on small talk, no buried-fact leak, preserves language, clean flat JSON. Weaknesses: coarse granularity, missed a multi-turn value update.
- Qwen2.5-1.5B — best extraction granularity (atomic facts), caught the value update, zero small-talk leakage. Weaknesses: weakest consolidation (run-on, no dedup) and one degenerate buried-fact output.
- gemma-3-1b — best consolidation (dedup works, faithful, clean single-memory). Weaknesses: leaks small talk and translated German.
- LFM2-1.2B — solid and fastest to load. Weaknesses:
Label: valuenoise, small-talk + buried leaks, a fluffy single-memory summary.
Recommendation
Extraction favors precision (do not pollute long-term memory) → Qwen3-1.7B is the best single pick (its consolidation is good enough). If running a second model for consolidation, gemma-3-1b wins that task.
Shipped local options: llama3.2:3b, qwen3-1.7b (recommended), gemma-3-1b, qwen2.5-1.5b, lfm2-1.2b.
Default: online (the configured smol model).
Known Mnemopi parser bugs (surfaced by these experiments)
String(item)produces[object Object]on object array items.- The line-fallback drops items
<=10chars, so a correct short fact likeName: Canis discarded.
Integration notes
providers.tinyModel,providers.memoryModel, andproviders.autoThinkingModeldefault toonline, so existing users get no downloads or on-device inference cost unless they opt in.- Local inference runs in a worker (off the main thread); models are cached on disk and downloaded on first use.
- The memory local path applies the refined recipes (line-format + small-talk-guarded extraction prompt, hardened consolidation prompt) via Mnemopi prompt overrides; the online path is unchanged.
providers.autoThinkingModeluses the same shipped local options asproviders.memoryModel.