Files
elizaos--eliza/plugins/plugin-xai/AGENTS.md
T
wehub-resource-sync 426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:43:05 +08:00

9.6 KiB

@elizaos/plugin-xai

xAI Grok models for text generation and embeddings in elizaOS.

Purpose / Role

Registers xAI's Grok family as the active model handlers for TEXT_SMALL, TEXT_LARGE, and TEXT_EMBEDDING within an Eliza agent runtime. It is opt-in and auto-enabled — the elizaOS plugin loader enables it automatically when XAI_API_KEY or GROK_API_KEY is present in the environment (via auto-enable.ts). No actions, providers, services, or evaluators are added; this plugin is purely a model-handler registration. For X (Twitter) social interactions, use @elizaos/plugin-x instead.

Plugin Surface

This plugin registers no actions, providers, services, or evaluators. It registers three model handlers:

ModelType Handler Default model
TEXT_SMALL handleTextSmall grok-3-mini
TEXT_LARGE handleTextLarge grok-3
TEXT_EMBEDDING handleTextEmbedding grok-embedding

All handlers POST to the xAI OpenAI-compatible REST API (/chat/completions, /embeddings). Streaming (onStreamChunk) is supported for both text handlers. Tool-call plumbing (tools, toolChoice, responseSchema, messages) is handled natively via generateText in models/grok.ts; callers that pass those fields receive an XaiNativeTextResult shape rather than a plain string. Token usage is emitted via EventType.MODEL_USED after every call; if the API does not return usage data, it is estimated from character count.

Layout

plugins/plugin-xai/
  index.ts              Plugin definition (XAIPlugin export, model handler wiring)
  index.node.ts         Node/Bun entry — re-exports index.ts
  index.browser.ts      Browser entry — re-exports index.ts
  auto-enable.ts        elizaos.plugin.autoEnableModule — lightweight env check
  models/
    grok.ts             All Grok logic: getConfig, generateText, createEmbedding,
                        handleTextSmall, handleTextLarge, handleTextEmbedding,
                        listModels, isGrokConfigured, tool-call normalization
    index.ts            Re-exports grok.ts
  __tests__/
    native-plumbing.shape.test.ts   Unit tests for tool-call normalization shapes
    error-policy.shape.test.ts      #12182 failure-path surfaces (typed errors, no fabricated completions)
    plugin.live.test.ts             Live API integration test (requires XAI_API_KEY)
  build.ts              Bun.build script (produces node ESM, browser ESM, CJS)
  vitest.config.ts      Vitest config

Commands

All scripts from this plugin's package.json:

bun run --cwd plugins/plugin-xai build        # compile to dist/
bun run --cwd plugins/plugin-xai dev          # watch build (bun --hot)
bun run --cwd plugins/plugin-xai test         # vitest run
bun run --cwd plugins/plugin-xai typecheck    # tsgo --noEmit
bun run --cwd plugins/plugin-xai lint         # biome check
bun run --cwd plugins/plugin-xai format       # biome format --write
bun run --cwd plugins/plugin-xai format:check # biome format (check only)
bun run --cwd plugins/plugin-xai clean        # rm -rf dist .turbo

Config / Env Vars

Variable Required Default Description
XAI_API_KEY one-of xAI API key. Checked first; falls back to GROK_API_KEY.
GROK_API_KEY one-of Alias accepted by both auto-enable and getConfig; used if XAI_API_KEY is not set.
XAI_MODEL no grok-3 Large text model. Also aliased as XAI_LARGE_MODEL.
XAI_SMALL_MODEL no grok-3-mini Small text model.
XAI_EMBEDDING_MODEL no grok-embedding Embedding model.
XAI_BASE_URL no https://api.x.ai/v1 API base URL (useful for proxies).

*At least one of XAI_API_KEY / GROK_API_KEY is required. Both keys are accepted by getConfig in models/grok.ts (XAI_API_KEY ?? GROK_API_KEY), so either key is sufficient for model calls.

Read via runtime.getSetting(key) — not process.env directly.

How to Extend

Add a new model type handler (e.g. TEXT_TOKENIZE):

  1. Implement the handler function in models/grok.ts following the pattern of handleTextSmall / handleTextLarge.
  2. Export it from models/index.ts.
  3. Register it in the models map in index.ts:
    [ModelType.TEXT_TOKENIZE]: handleTextTokenize,
    
  4. Add the corresponding capability string to the elizaos.plugin.capabilities array in package.json if a standard name exists.

Add a provider or action — this plugin intentionally has none. If you need runtime context exposure or conversation-level actions, consider adding a separate plugin rather than bloating this model-only plugin.

Conventions / Gotchas

  • XAI_API_KEY and GROK_API_KEY are both accepted. getConfig in models/grok.ts reads XAI_API_KEY ?? GROK_API_KEY, so either key works for model calls as well as auto-enable.
  • Native vs. string return. generateText returns XaiNativeTextResult when messages, tools, toolChoice, or responseSchema are passed. The handler's TypeScript signature says Promise<string | TextStreamResult> to satisfy the elizaOS ModelHandler type, but callers passing those fields receive the native shape. Do not widen the return type — the elizaOS plugin contract does not support it.
  • Streaming. Pass stream: true and onStreamChunk together. The handler returns the accumulated fullText string, not a stream object.
  • No native SDK dependency. This plugin calls the xAI REST API directly with fetch — there is no openai or xai-sdk npm package involved. This keeps the bundle small and browser-compatible.
  • Dual build targets. The package exports separate browser and node builds (see exports in package.json). Both resolve to the same index.ts logic since fetch is available in both environments.
  • For repo-wide logger rules, naming, and architecture commandments, see the root AGENTS.md.

NON-NEGOTIABLE — evidence, trajectories & real end-to-end tests

The binding, repo-wide standard is AGENTS.md. Read it. Nothing in this package is done until it is proven done — a reviewer must confirm it works without reading the code, from the artifacts you attach. This applies to every feature, fix, refactor, and chore here. "Tests pass" is not proof; "CI is green" is not proof.

  • Record AND read model trajectories. Capture the actual inputs and outputs of the model from a live LLM — not the deterministic proxy, not a mock: the prompt, the providers/context, the raw model output, every tool/action call, and the result. Then open the trajectory and review it by hand. A captured-but-unread trajectory is not evidence (packages/scenario-runner/bin/eliza-scenarios run <scenario> --report <out>).
  • Real, full-featured E2E — no larp. Every feature ships detailed end-to-end tests that drive the real path end to end. Not the happy "front door" only: cover error paths, edge/empty/invalid input, concurrency, roles/permissions, and adversarial input. A test that asserts against a mock/stub/fixture standing in for the thing under test does not count. If the real model/device/chain/connector/account is hard to reach, make it reachable — that is the work, not an excuse to mock. If the existing tests here are shallow or mocked, fixing them is part of your change.
  • Screenshots + logs at every phase, plus a complete walkthrough video/run-through of the entire feature or view, start to finish (bun run test:e2e:record).
  • Manually review every artifact the change touches — never just the green check: client logs (console + network), server logs ([ClassName] …), the model trajectories in and out, before/after full-page screenshots, and the domain artifacts listed below for this package.
  • No residuals. No shortcuts. The goal is not "done" — it is everything done. Clear every blocker by the hard path: build the real architecture, stand up the real model/device/service, actually test it. Never leave a TODO, a stub, a stepping-stone, or a "follow-up." When unsure, research thoroughly, weigh the options, and ship the best, highest-effort, production-ready version. Keep going until every possibility is exhausted.

Artifacts → attached inline in the PR (MP4 video, JPG screenshots, logs in <details>); attach each evidence type or explicitly mark it N/A with a reason — never leave it blank. If develop moved and changed behavior, re-capture evidence; stale proof is worse than none.

Capture & manually review for this package — model provider:

  • A trajectory from a live call to this provider (not the proxy, not a mock): full request, raw response, token usage, finish reason, and streamed chunks.
  • Proof of tool/function-calling and structured-output parsing against the real model.
  • The error paths exercised: bad key, model-not-found, oversized context, timeout, rate-limit, mid-stream disconnect — plus latency and cost from the real call.
  • If no key is available in CI, attach the documented live-run transcript as evidence — never a mocked client passed off as a pass.