Compare commits

..

54 Commits

Author SHA1 Message Date
Codex 9301b8f751 docs(priorities): refresh architect queue
Harness (E2E) / Harnesses (mock LLM) (push) Waiting to run
Harness (E2E) / Provider harnesses (live LLM conformance) (push) Waiting to run
Lint / golangci-lint (push) Waiting to run
Run Tests / Unit Tests (push) Waiting to run
Run Tests / Etcd Integration Tests (push) Waiting to run
2026-07-01 11:09:49 +00:00
Asim Aslam c110774dc7 Add opt-in retries for agent tool calls (#3535)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 10:57:33 +01:00
Asim Aslam c0a5775fb5 docs(priorities): refresh architect queue (#3533)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 10:12:45 +01:00
Asim Aslam f844a23bb2 docs: align public AI harness facts (#3531)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 09:40:28 +01:00
Asim Aslam 04e8759d41 test agent checkpoint resume after restart (#3529)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 09:06:46 +01:00
Asim Aslam 78725135aa docs(priorities): refresh architect queue (#3527)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 08:33:01 +01:00
Asim Aslam 2ff64ff0b2 Document canonical 0-to-hero reference path (#3522)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 07:12:17 +01:00
Asim Aslam 7a8e7cd9ae docs(priorities): advance queue to 0-to-hero reference (#3520)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 06:37:15 +01:00
Asim Aslam b58eed1698 Add retrieval-backed agent memory (#3518)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 04:59:16 +01:00
Asim Aslam c92c9cc244 docs(priorities): refresh architect queue (#3516)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 04:06:59 +01:00
Asim Aslam a2bf43e9ef test: broaden stream provider conformance (#3512)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 03:37:26 +01:00
Asim Aslam 49bac7e4a8 trace scheduled flow dispatch metadata (#3510)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 02:42:31 +01:00
Asim Aslam 5aac7e3ca0 docs(priorities): refresh architect queue after scheduling (#3507)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 01:51:21 +01:00
Asim Aslam d86585bf5c Add scheduled flow agent harness (#3505)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 01:02:29 +01:00
Asim Aslam 88e2b58711 docs(priorities): refresh architect queue (#3503)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 00:38:51 +01:00
Asim Aslam d259383645 Harden agent terminal failure statuses (#3499)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 00:03:21 +01:00
Asim Aslam 2d7ee300a4 docs(priorities): refresh architect queue after conformance (#3497)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 23:40:19 +01:00
Asim Aslam 57fa4e3b7a Add mock provider conformance target (#3495)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 23:09:19 +01:00
Asim Aslam 0f1917f26b docs(priorities): refresh architect queue after verification (#3493)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 22:40:05 +01:00
Asim Aslam 6e9c5e87e9 Add flow step verification loop (#3489)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 22:18:01 +01:00
Asim Aslam d8bb892425 docs(priorities): refresh architect queue after a2a continuity (#3487)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 21:37:56 +01:00
Asim Aslam 064a112c6b a2a: expose resubscribe and input-required support (#3484)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 21:23:43 +01:00
Asim Aslam 24d103c658 docs(priorities): refresh architect queue after memory (#3482)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 20:47:39 +01:00
Asim Aslam 2d8e3da943 blog: expand /blog/32 into a field guide on agent frameworks (#3478)
* blog+docs: drop the word "bet" from the tRPC-Agent-Go comparison

Reword "two bets" / "opposite bet" / "the bet is" to approaches / premise /
principle across blog/32 and the comparison guide.

* blog: expand /blog/32 into a field guide on agent frameworks

Roughly double the length with deeper context: the first wave (LangChain &
co.), the two layers of a harness (intra-agent vs operational), loop
engineering and the move to scheduled/looping/work-performing agents, a
survey of where the frameworks are going (LangGraph, CrewAI, AutoGen, ADK,
tRPC-Agent-Go), then Go Micro's "an agent is a service" position, the honest
tRPC-Agent-Go contrast, and MCP/A2A interop. Drops the word "bet" throughout.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-30 20:20:14 +01:00
Asim Aslam 3081dab246 Add agent memory summarizer hook (#3479)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 20:18:23 +01:00
Asim Aslam fdc422c16e blog+docs: position go-micro vs tRPC-Agent-Go (agent = service) (#3475)
Add a fair, honest positioning piece on the architectural fork with
tRPC-Agent-Go (an agent SDK alongside your services / graph DSL) vs Go Micro
(one runtime where an agent is a service, every endpoint a tool, durable
flows not a graph DSL). New blog post /blog/32 + a parallel section in the
existing comparison guide; honest about where tRPC-Agent-Go is ahead
(eval, self-evolution, RAG) and that they interoperate over MCP/A2A.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-30 19:45:52 +01:00
Asim Aslam 5e3f8db193 docs(priorities): refresh architect queue for memory (#3476)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 19:41:17 +01:00
Asim Aslam f5bf5f7987 Wire A2A streaming through agent StreamAsk (#3471)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 19:18:23 +01:00
Asim Aslam 3ec12b7c72 Trace agent checkpoint resume events (#3468)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 18:22:00 +01:00
Asim Aslam fa4d4b7f6a docs(priorities): refresh architect queue after failure hardening (#3466)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 17:50:23 +01:00
Asim Aslam 43beaafab5 flow: classify workflow failure kinds (#3464)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 17:30:45 +01:00
Asim Aslam 21a20005cb docs(priorities): refresh architect queue (#3462)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 16:51:07 +01:00
Asim Aslam e1b3c587aa Add configurable provider conformance dispatch (#3459)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 16:31:49 +01:00
Asim Aslam da2bbab80c docs(priorities): refresh architect queue (#3457)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 15:55:52 +01:00
Asim Aslam 5a59e1ece4 agent: verify durable resume example (#3452)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 15:26:32 +01:00
Asim Aslam aea03ba9e7 priorities: queue the flow verification/grader loop (#3435) (#3436)
Add the verification loop as a ranked priority — the one missing layer from
the four-loop framing (agent / verification / event-driven / hill-climbing):
flow.Verify + flow.LLMGrader to grade a step's output against a rubric and
retry with feedback. The architect will re-rank on its next pass.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-30 15:17:18 +01:00
Asim Aslam 54bc05e48f docs(priorities): advance agent durability queue (#3450)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 14:40:14 +01:00
Asim Aslam 3ef0d98a16 flow: analyze run traces for optimization (#3447)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 14:25:15 +01:00
Asim Aslam 5a6c8d8b30 docs(priorities): advance architect queue (#3445)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 13:41:58 +01:00
Asim Aslam 1cd918c2b9 flow: add verification grader loop (#3443)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 13:28:53 +01:00
Asim Aslam e96d4a67bc docs(priorities): refresh architect queue (#3441)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 13:09:37 +01:00
Asim Aslam 8994fd03f6 Add A2A fallback provider conformance harness (#3438)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 12:37:40 +01:00
Asim Aslam 0595130f16 run: surface MCP tools in the micro run banner by default (#3434)
The gateway already serves /mcp/tools on :8080 unconditionally (every
endpoint is an AI-callable tool), but the startup banner only printed an MCP
line when --mcp-address was set — so the live `micro run` experience hid the
harness's signature feature even though it was running, and didn't match the
README. Always advertise MCP Tools on the gateway address; keep the optional
standalone MCP-protocol server (--mcp-address) as a clearly separate line.
No behavior change — banner output only.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-30 11:21:26 +01:00
Asim Aslam 37eccc425e docs(priorities): advance provider conformance queue (#3433)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 11:14:39 +01:00
Asim Aslam 010e0fe57c Classify agent run failure summaries (#3430)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 10:55:14 +01:00
Asim Aslam 113f268268 docs: align DevRel public surface facts (#3428)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 09:26:20 +01:00
Asim Aslam c10697f08d docs(priorities): advance resilience queue (#3427)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 09:25:32 +01:00
Asim Aslam f61f3dc04d test: add AI stream provider conformance (#3423)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 08:56:12 +01:00
Asim Aslam d9e6c68938 docs(priorities): refresh architect queue (#3421)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 07:40:48 +01:00
Asim Aslam 8847668c83 Improve agent telemetry error classification (#3418)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 07:06:38 +01:00
Asim Aslam e45bf3d114 docs(priorities): remove completed durable resume item (#3416)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 05:50:51 +01:00
Asim Aslam 98d1f58cd6 Add durable agent resume example (#3414)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 04:56:14 +01:00
Asim Aslam 51158ed2a7 docs(priorities): refresh architect queue (#3412)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 04:04:11 +01:00
Asim Aslam 09bf82d4f6 fix a2a stream fallback (#3410)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 03:26:42 +01:00
58 changed files with 3197 additions and 162 deletions
+29 -3
View File
@@ -14,6 +14,20 @@ on:
schedule:
- cron: "17 6 * * *" # daily, so the world is exercised even without changes
workflow_dispatch:
inputs:
providers:
description: "Comma-separated providers for live conformance (default: all supported)"
required: false
default: "anthropic,openai,gemini,groq,mistral,together,atlascloud"
harnesses:
description: "Comma-separated harnesses for live conformance"
required: false
default: "agent,universe,agent-flow,plan-delegate,a2a-stream-fallback"
require_configured:
description: "Fail selected live providers that do not have repository secrets"
required: false
type: boolean
default: false
jobs:
harness:
@@ -54,10 +68,22 @@ jobs:
TOGETHER_API_KEY: ${{ secrets.TOGETHER_API_KEY }}
ATLASCLOUD_API_KEY: ${{ secrets.ATLASCLOUD_API_KEY }}
run: |
go run ./internal/harness/provider-conformance \
-summary-json provider-conformance-summary.json \
-summary-markdown provider-conformance-summary.md \
PROVIDERS="${{ github.event.inputs.providers || 'anthropic,openai,gemini,groq,mistral,together,atlascloud' }}"
HARNESSES="${{ github.event.inputs.harnesses || 'agent,universe,agent-flow,plan-delegate,a2a-stream-fallback' }}"
REQUIRE_CONFIGURED="${{ github.event.inputs.require_configured || 'false' }}"
args=(
-providers "$PROVIDERS"
-harnesses "$HARNESSES"
-summary-json provider-conformance-summary.json
-summary-markdown provider-conformance-summary.md
-capabilities-markdown provider-capabilities.md
)
if [ "$REQUIRE_CONFIGURED" = "true" ]; then
args+=( -require-configured )
fi
go run ./internal/harness/provider-conformance "${args[@]}"
- name: Publish provider conformance summary
if: always()
run: |
+8 -1
View File
@@ -8,7 +8,7 @@ LDFLAGS = -X $(GIT_IMPORT).BuildDate=$(BUILD_DATE) -X $(GIT_IMPORT).GitCommit=$(
# GORELEASER_DOCKER_IMAGE = ghcr.io/goreleaser/goreleaser-cross:v1.25.7
GORELEASER_DOCKER_IMAGE = ghcr.io/goreleaser/goreleaser:latest
.PHONY: test test-race test-coverage harness provider-conformance lint fmt install-tools proto clean help gorelease-dry-run gorelease-dry-run-docker
.PHONY: test test-race test-coverage harness provider-conformance-mock provider-conformance lint fmt install-tools proto clean help gorelease-dry-run gorelease-dry-run-docker
# Default target
help:
@@ -19,6 +19,7 @@ help:
@echo " make test-coverage - Run tests with coverage"
@echo " make lint - Run linter"
@echo " make harness - Run deterministic getting-started and end-to-end harnesses"
@echo " make provider-conformance-mock - Run cross-provider harness with deterministic mock provider"
@echo " make provider-conformance - Run harnesses against configured live providers"
@echo " make fmt - Format code"
@echo " make install-tools - Install development tools"
@@ -50,6 +51,12 @@ harness:
go test ./cmd/micro/cli/new -run TestZeroToOne -count=1
./internal/harness/zero-to-hero-ci/run.sh
go run ./internal/harness/agent-flow
$(MAKE) provider-conformance-mock
# Run the shared provider conformance contract with the deterministic mock
# provider. This is the no-secret path used by CI and local dogfooding to keep
# provider-facing agent/tool semantics covered on every machine.
provider-conformance-mock:
go run ./internal/harness/provider-conformance -providers mock
# Run the same harnesses against every configured live provider. Providers
+4 -3
View File
@@ -65,7 +65,7 @@ curl -X POST http://localhost:8080/api/helloworld/Helloworld.Call \
```
This scaffold → run → call path is covered by the no-secret CI harness. To run
the same local contract (including the 0→hero services → agents → workflows path,
the same local contract (including the [0→hero services → agents → workflows path](internal/website/docs/guides/zero-to-hero.md),
chat/inspect CLI boundaries, and deploy dry-run), use:
```bash
@@ -401,8 +401,8 @@ Swap providers with a single import — same interface everywhere:
| Google Gemini | `gemini-2.5-flash` |
| Groq | `llama-3.3-70b-versatile` |
| Mistral | `mistral-large-latest` |
| Together AI | `Llama-3.3-70B-Instruct-Turbo` |
| Atlas Cloud | `llama-3.3-70b` |
| Together AI | `meta-llama/Llama-3.3-70B-Instruct-Turbo` |
| Atlas Cloud | `deepseek-ai/DeepSeek-V3-0324` |
```go
m := ai.New("anthropic", ai.WithAPIKey(key))
@@ -423,6 +423,7 @@ See [all examples](examples/README.md).
- [Getting Started](internal/website/docs/getting-started.md)
- [AI Integration](internal/website/docs/ai-integration.md)
- [0→hero Reference](internal/website/docs/guides/zero-to-hero.md)
- [Agents and Workflows](internal/website/docs/guides/agents-and-workflows.md)
- [Agent Design](internal/docs/AGENT_DESIGN.md)
- [Plan & Delegate](internal/website/docs/guides/plan-delegate.md)
+102
View File
@@ -0,0 +1,102 @@
package agent
import (
"bytes"
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/gateway/a2a"
)
func TestA2AStreamUsesAgentChatPathWithTools(t *testing.T) {
var sawTool bool
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
if opts.ToolHandler == nil {
t.Fatal("model was not wired with agent tool handler")
}
result := opts.ToolHandler(ctx, ai.ToolCall{
ID: "call-1",
Name: "echo",
Input: map[string]any{"value": "a2a-stream"},
})
if !strings.Contains(result.Content, "a2a-stream-ok") {
t.Fatalf("tool result = %q, want marker", result.Content)
}
return &ai.Response{Answer: "streamed " + result.Content}, nil
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("stream-agent"), WithTool("echo", "echo text", nil, func(ctx context.Context, input map[string]any) (string, error) {
sawTool = true
if info, ok := ai.RunInfoFrom(ctx); !ok || info.RunID == "" || info.Agent != "stream-agent" {
t.Fatalf("RunInfo = %+v ok=%v, want stream-agent run", info, ok)
}
if input["value"] != "a2a-stream" {
t.Fatalf("tool input = %+v, want a2a-stream", input)
}
return "a2a-stream-ok", nil
}))
h := a2a.NewAgentStreamHandler(
a2a.Card("stream-agent", "http://example.invalid/stream-agent", "", nil),
func(ctx context.Context, text string) (string, error) {
resp, err := a.Ask(ctx, text)
if err != nil {
return "", err
}
return resp.Reply, nil
},
a.streamAskAI,
)
body := []byte(`{"jsonrpc":"2.0","id":1,"method":"message/stream","params":{"message":{"role":"user","parts":[{"kind":"text","text":"run stream tool"}],"kind":"message"}}}`)
req := httptest.NewRequest(http.MethodPost, "/", bytes.NewReader(body))
rr := httptest.NewRecorder()
h.ServeHTTP(rr, req)
if !sawTool {
t.Fatal("A2A stream did not execute the agent tool path")
}
if ct := rr.Result().Header.Get("Content-Type"); !strings.HasPrefix(ct, "text/event-stream") {
t.Fatalf("content-type = %q, want text/event-stream", ct)
}
if !strings.Contains(rr.Body.String(), "a2a-stream-ok") {
t.Fatalf("stream body missing tool marker: %s", rr.Body.String())
}
var final struct {
Result struct {
Status struct {
State string `json:"state"`
} `json:"status"`
Artifacts []struct {
Parts []struct {
Text string `json:"text"`
} `json:"parts"`
} `json:"artifacts"`
} `json:"result"`
Error any `json:"error"`
}
for _, line := range strings.Split(strings.TrimSpace(rr.Body.String()), "\n") {
line = strings.TrimSpace(strings.TrimPrefix(strings.TrimSpace(line), "data: "))
if line == "" {
continue
}
if err := json.Unmarshal([]byte(line), &final); err != nil {
t.Fatalf("decode event %q: %v", line, err)
}
}
if final.Error != nil {
t.Fatalf("final event error: %+v", final.Error)
}
if final.Result.Status.State != "completed" {
t.Fatalf("final state = %q, want completed", final.Result.Status.State)
}
if len(final.Result.Artifacts) != 1 || len(final.Result.Artifacts[0].Parts) != 1 || !strings.Contains(final.Result.Artifacts[0].Parts[0].Text, "a2a-stream-ok") {
t.Fatalf("final artifacts = %+v, want tool marker", final.Result.Artifacts)
}
}
+14 -5
View File
@@ -19,6 +19,7 @@ import (
"net/http"
"strings"
"sync"
"time"
"github.com/google/uuid"
pb "go-micro.dev/v6/agent/proto"
@@ -177,7 +178,9 @@ func (a *agentImpl) setupWithToolHandler(handler ai.ToolHandler) {
case a.ephemeral:
a.mem = NewInMemory(a.opts.HistoryLimit)
case a.opts.MemoryCompaction.MaxMessages > 0:
a.mem = NewCompactingMemory(a.stateStore(), "history", a.opts.MemoryCompaction.MaxMessages, a.opts.MemoryCompaction.KeepRecent)
a.mem = NewCompactingMemoryWithOptions(a.stateStore(), "history", a.opts.MemoryCompaction)
case a.opts.MemoryRetrievalLimit > 0:
a.mem = NewRetrievalMemory(a.stateStore(), "history", a.opts.MemoryRetrievalLimit)
default:
a.mem = NewMemory(a.stateStore(), "history", a.opts.HistoryLimit)
}
@@ -271,6 +274,9 @@ func (a *agentImpl) askLocked(ctx context.Context, runID, message, parentRunID s
return nil, err
}
ctx, endRun := a.startRun(ctx, message)
if existing != nil {
a.recordTimelineEvent(ctx, RunEvent{Time: time.Now(), RunID: runID, ParentID: parentRunID, Agent: a.opts.Name, Kind: "resume", Name: run.State.Stage})
}
defer func() { endRun(err) }()
messages := a.mem.Messages()
@@ -294,12 +300,15 @@ func (a *agentImpl) askLocked(ctx context.Context, runID, message, parentRunID s
Backoff: a.opts.ModelRetryBackoff,
})
if err != nil {
run.Status = "failed"
run.Steps[0].Status = "failed"
run.Steps[0].Error = err.Error()
run.Status = agentRunFailureStatus(err)
if a.currentRun != nil {
run.Steps = a.currentRun.Steps
}
if len(run.Steps) == 0 {
run.Steps = []flow.StepRecord{{Name: agentAskStep}}
}
run.Steps[0].Status = run.Status
run.Steps[0].Error = err.Error()
_ = a.saveRun(ctx, run)
return nil, err
}
@@ -419,7 +428,7 @@ func (a *agentImpl) Run() error {
return "", err
}
return resp.Reply, nil
}, a.Stream)
}, a.streamAskAI)
go func() {
if err := http.ListenAndServe(a.opts.A2AAddress, handler); err != nil {
fmt.Printf("agent %s A2A server: %v\n", a.opts.Name, err)
+96
View File
@@ -5,6 +5,7 @@ import (
"encoding/json"
"fmt"
"strings"
"time"
"go-micro.dev/v6/ai"
codecBytes "go-micro.dev/v6/codec/bytes"
@@ -121,6 +122,7 @@ func (a *agentImpl) toolHandler() ai.ToolHandler {
// so the result runs plan → step → loop → approve → checkpoint → base.
h := a.baseHandler()
h = a.toolTimeoutWrap(h)
h = a.toolRetryWrap(h)
h = a.checkpointToolWrap(h)
h = a.approveWrap(h)
h = a.loopWrap(h)
@@ -164,6 +166,99 @@ func (a *agentImpl) toolTimeoutWrap(next ai.ToolHandler) ai.ToolHandler {
}
}
// toolRetryWrap retries transient tool failures with bounded backoff. It is
// opt-in because tools can have side effects; guardrail refusals and caller
// cancellation are never retried.
func (a *agentImpl) toolRetryWrap(next ai.ToolHandler) ai.ToolHandler {
return func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
maxAttempts := a.opts.ToolMaxAttempts
if maxAttempts <= 0 {
maxAttempts = 1
}
var res ai.ToolResult
for attempt := 1; attempt <= maxAttempts; attempt++ {
if err := ctx.Err(); err != nil {
return errResult(call.ID, err.Error())
}
res = next(ctx, call)
if !retryableToolResult(res) || attempt == maxAttempts || ctx.Err() != nil {
return annotateToolAttempts(res, attempt)
}
t := time.NewTimer(toolRetryBackoff(attempt, a.opts.ToolRetryBackoff))
select {
case <-ctx.Done():
if !t.Stop() {
<-t.C
}
return errResult(call.ID, ctx.Err().Error())
case <-t.C:
}
}
return annotateToolAttempts(res, maxAttempts)
}
}
func retryableToolResult(res ai.ToolResult) bool {
if res.Refused != "" {
return false
}
msg := toolErrorMessage(res)
if msg == "" {
return false
}
return ai.IsTransientError(fmt.Errorf("%s", msg))
}
func toolErrorMessage(res ai.ToolResult) string {
if m, ok := res.Value.(map[string]string); ok {
return m["error"]
}
if m, ok := res.Value.(map[string]any); ok {
if v, ok := m["error"].(string); ok {
return v
}
}
var decoded map[string]string
if err := json.Unmarshal([]byte(res.Content), &decoded); err == nil {
return decoded["error"]
}
return ""
}
func annotateToolAttempts(res ai.ToolResult, attempts int) ai.ToolResult {
if attempts <= 1 {
return res
}
res.Attempts = attempts
if m, ok := res.Value.(map[string]string); ok {
cp := map[string]any{}
for k, v := range m {
cp[k] = v
}
cp["attempts"] = attempts
res.Value = cp
if b, err := json.Marshal(cp); err == nil {
res.Content = string(b)
}
}
return res
}
func toolRetryBackoff(attempt int, base time.Duration) time.Duration {
if base <= 0 {
base = 200 * time.Millisecond
}
if shift := attempt - 1; shift > 0 {
base <<= shift
}
if base > 30*time.Second {
return 30 * time.Second
}
return base
}
// baseHandler executes a tool call: a developer custom tool, the built-in
// delegate, or an RPC to the service. It is the innermost handler.
func (a *agentImpl) baseHandler() ai.ToolHandler {
@@ -338,6 +433,7 @@ func (a *agentImpl) handleDelegate(ctx context.Context, call ai.ToolCall) ai.Too
ModelCallTimeout(a.opts.ModelTimeout),
ModelRetry(a.opts.ModelMaxAttempts, a.opts.ModelRetryBackoff),
ToolCallTimeout(a.opts.ToolTimeout),
ToolRetry(a.opts.ToolMaxAttempts, a.opts.ToolRetryBackoff),
TraceProvider(a.opts.TraceProvider),
)
// Record lineage so the sub-agent's tool calls carry this run as parent.
+20 -1
View File
@@ -49,6 +49,12 @@ func (a *agentImpl) saveRun(ctx context.Context, run flow.Run) error {
if err := a.opts.Checkpoint.Save(ctx, run); err != nil {
return fmt.Errorf("agent %s checkpoint save: %w", a.opts.Name, err)
}
if info, ok := ai.RunInfoFrom(ctx); ok {
a.recordTimelineEvent(ctx, RunEvent{
Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent,
Kind: "checkpoint", Name: run.State.Stage, Status: run.Status,
})
}
return nil
}
@@ -165,13 +171,26 @@ func (a *agentImpl) pending(ctx context.Context) ([]flow.Run, error) {
func terminalAgentRunStatus(status string) bool {
switch status {
case "done", "canceled", "expired":
case "done", "canceled", "timeout", "rate_limited", "expired":
return true
default:
return false
}
}
func agentRunFailureStatus(err error) string {
switch ai.ClassifyError(err) {
case ai.ErrorKindCanceled:
return "canceled"
case ai.ErrorKindTimeout:
return "timeout"
case ai.ErrorKindRateLimited:
return "rate_limited"
default:
return "failed"
}
}
func (a *agentImpl) checkpointToolWrap(next ai.ToolHandler) ai.ToolHandler {
return func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
if a.opts.Checkpoint == nil || a.currentRun == nil {
+76 -7
View File
@@ -13,7 +13,7 @@ import (
func TestResumeCompletedCheckpointDoesNotReplayModel(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "durable-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "durable-agent")
calls := 0
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
calls++
@@ -52,7 +52,7 @@ func TestResumeCompletedCheckpointDoesNotReplayModel(t *testing.T) {
func TestResumeFailedCheckpointDoesNotReplayCompletedTool(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "tool-resume-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "tool-resume-agent")
toolRuns := 0
first := true
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
@@ -102,6 +102,75 @@ func TestResumeFailedCheckpointDoesNotReplayCompletedTool(t *testing.T) {
}
}
func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "restart-resume-agent")
toolRuns := 0
modelCalls := 0
failFirst := true
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
modelCalls++
if opts.ToolHandler != nil {
res := opts.ToolHandler(ctx, ai.ToolCall{ID: "call-1", Name: "external.provision", Input: map[string]any{"service": "api"}})
if res.Content != "provisioned" {
t.Fatalf("tool result = %q, want provisioned", res.Content)
}
}
if failFirst {
failFirst = false
return nil, errors.New("process stopped after tool checkpoint")
}
return &ai.Response{Reply: "resumed after restart"}, nil
}
defer func() { fakeGen = nil }()
newAgent := func() *agentImpl {
return newTestAgent(Name("restart-resume-agent"), WithCheckpoint(cp),
WithTool("external.provision", "provision service once", nil, func(context.Context, map[string]any) (string, error) {
toolRuns++
return "provisioned", nil
}))
}
first := newAgent()
_, err := first.Ask(ctx, "provision api")
if err == nil {
t.Fatal("Ask succeeded, want simulated process stop")
}
if toolRuns != 1 {
t.Fatalf("tool executions after failed Ask = %d, want 1", toolRuns)
}
runs, err := Pending(ctx, first)
if err != nil {
t.Fatalf("Pending before restart: %v", err)
}
if len(runs) != 1 {
t.Fatalf("Pending before restart returned %d runs, want 1", len(runs))
}
restarted := newAgent()
resp, err := Resume(ctx, restarted, runs[0].ID)
if err != nil {
t.Fatalf("Resume after restart: %v", err)
}
if resp.Reply != "resumed after restart" || resp.RunID != runs[0].ID {
t.Fatalf("response = %#v, want resumed reply on original run id", resp)
}
if toolRuns != 1 {
t.Fatalf("tool executions after restart resume = %d, want checkpointed tool not replayed", toolRuns)
}
if modelCalls != 2 {
t.Fatalf("model calls = %d, want initial call plus resumed call", modelCalls)
}
loaded, ok, err := cp.Load(ctx, runs[0].ID)
if err != nil || !ok {
t.Fatalf("Load resumed run ok=%v err=%v", ok, err)
}
if loaded.Status != "done" || loaded.ParentID != runs[0].ParentID {
t.Fatalf("loaded run status/parent = %s/%s, want done/%s", loaded.Status, loaded.ParentID, runs[0].ParentID)
}
}
func TestResumeFailedCheckpointDoesNotDuplicateCompactedMemory(t *testing.T) {
ctx := context.Background()
st := store.NewMemoryStore()
@@ -170,7 +239,7 @@ func countMemoryContent(messages []ai.Message, needle string) int {
func TestPendingReturnsUnfinishedAgentRuns(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "pending-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "pending-agent")
run := flow.Run{ID: "run-1", Flow: "pending-agent", Status: "failed", State: flow.State{Stage: agentAskStep, Data: []byte("retry me")}}
if err := cp.Save(ctx, run); err != nil {
t.Fatalf("Save: %v", err)
@@ -187,7 +256,7 @@ func TestPendingReturnsUnfinishedAgentRuns(t *testing.T) {
func TestPendingSkipsTerminalCanceledAndExpiredAgentRuns(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "terminal-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "terminal-agent")
for _, run := range []flow.Run{
{ID: "active", Flow: "terminal-agent", Status: "failed", State: flow.State{Stage: agentAskStep, Data: []byte("retry me")}},
{ID: "done", Flow: "terminal-agent", Status: "done", State: flow.State{Stage: agentAskStep, Data: []byte("done")}},
@@ -216,7 +285,7 @@ func TestPendingSkipsTerminalCanceledAndExpiredAgentRuns(t *testing.T) {
func TestHumanInputPauseResumesSameRunWithInput(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "input-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "input-agent")
calls := 0
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
calls++
@@ -274,7 +343,7 @@ func TestHumanInputPauseResumesSameRunWithInput(t *testing.T) {
func TestHumanInputResumeHonorsCanceledContextAndLeavesRunPending(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "input-cancel-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "input-cancel-agent")
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
if opts.ToolHandler != nil {
opts.ToolHandler(ctx, ai.ToolCall{ID: "input-1", Name: toolHumanInput, Input: map[string]any{"prompt": "Approve deploy?"}})
@@ -319,7 +388,7 @@ func TestHumanInputResumeHonorsCanceledContextAndLeavesRunPending(t *testing.T)
func TestApprovalDenialPausesCheckpointedRunAndResumeContinues(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewStore(), "approval-agent")
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "approval-agent")
calls := 0
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
calls++
+55 -9
View File
@@ -24,6 +24,12 @@ type Memory interface {
Clear()
}
// MemorySummaryFunc turns older conversation messages into a compact
// replacement message for active context. It is called while the default
// memory is locked, so implementations should be deterministic and avoid
// calling back into the same memory instance.
type MemorySummaryFunc func([]ai.Message) ai.Message
// MemoryCompaction configures deterministic, store-backed context compaction
// for the default memory implementation. When the retained conversation grows
// past MaxMessages, older turns are collapsed into a summary message while the
@@ -31,6 +37,7 @@ type Memory interface {
type MemoryCompaction struct {
MaxMessages int
KeepRecent int
Summarize MemorySummaryFunc
}
// MemoryRecall is implemented by memory backends that can retrieve durable
@@ -49,11 +56,29 @@ func NewMemory(s store.Store, key string, limit int) Memory {
return m
}
// NewRetrievalMemory returns store-backed memory that keeps a bounded active
// conversation and archives every turn for retrieval. It is useful when callers
// want relevant durable recall without summary compaction in the active context.
// A nil store or empty key keeps only the active in-process buffer.
func NewRetrievalMemory(s store.Store, key string, activeLimit int) Memory {
m := &storeMemory{store: s, key: key, hist: ai.NewHistory(activeLimit), retrieveAll: true}
m.load()
return m
}
// NewCompactingMemory returns store-backed memory with explicit compaction and
// retrieval controls. It keeps all messages in the backing store, compacts older
// turns into a deterministic summary when the conversation exceeds maxMessages,
// and lets callers recall relevant prior turns with Recall.
func NewCompactingMemory(s store.Store, key string, maxMessages, keepRecent int) Memory {
return NewCompactingMemoryWithOptions(s, key, MemoryCompaction{MaxMessages: maxMessages, KeepRecent: keepRecent})
}
// NewCompactingMemoryWithOptions returns store-backed memory configured with
// explicit compaction options, including an optional summarization hook.
func NewCompactingMemoryWithOptions(s store.Store, key string, compaction MemoryCompaction) Memory {
maxMessages := compaction.MaxMessages
keepRecent := compaction.KeepRecent
if keepRecent <= 0 {
keepRecent = maxMessages / 2
}
@@ -69,6 +94,7 @@ func NewCompactingMemory(s store.Store, key string, maxMessages, keepRecent int)
compaction: MemoryCompaction{
MaxMessages: maxMessages,
KeepRecent: keepRecent,
Summarize: compaction.Summarize,
},
}
m.load()
@@ -84,16 +110,20 @@ func NewInMemory(limit int) Memory {
// storeMemory is the default Memory: an ai.History buffer optionally
// persisted to a store.
type storeMemory struct {
mu sync.Mutex
store store.Store
key string
hist *ai.History
compaction MemoryCompaction
archive []ai.Message
mu sync.Mutex
store store.Store
key string
hist *ai.History
compaction MemoryCompaction
archive []ai.Message
retrieveAll bool
}
func (m *storeMemory) Add(role, content string) {
m.mu.Lock()
if m.retrieveAll {
m.archive = append(m.archive, ai.Message{Role: role, Content: content})
}
m.hist.Add(role, content)
m.mu.Unlock()
m.compact()
@@ -117,6 +147,8 @@ func (m *storeMemory) Clear() {
// Recall returns archived messages whose content contains words from query.
// It is deterministic and provider-neutral: no embeddings or model calls are
// required, but semantic/vector stores can replace Memory for richer retrieval.
// When created with NewRetrievalMemory the archive contains every persisted
// turn; when created with NewCompactingMemory it contains compacted older turns.
func (m *storeMemory) Recall(query string, limit int) []ai.Message {
m.mu.Lock()
defer m.mu.Unlock()
@@ -170,6 +202,9 @@ func (m *storeMemory) load() {
}
m.mu.Lock()
m.archive = state.Archive
if m.retrieveAll && len(m.archive) == 0 {
m.archive = append(m.archive, state.Messages...)
}
for _, msg := range state.Messages {
m.hist.Add(msg.Role, msg.Content)
}
@@ -213,9 +248,13 @@ func (m *storeMemory) compact() {
older := msgs[:cut]
recent := msgs[cut:]
m.archive = append(m.archive, older...)
summary := ai.Message{
Role: "system",
Content: fmt.Sprintf("Conversation memory summary: %s", summarizeMessages(older)),
summarize := m.compaction.Summarize
if summarize == nil {
summarize = defaultMemorySummary
}
summary := summarize(older)
if summary.Role == "" {
summary.Role = "system"
}
m.hist.Reset()
m.hist.Add(summary.Role, summary.Content)
@@ -224,6 +263,13 @@ func (m *storeMemory) compact() {
}
}
func defaultMemorySummary(msgs []ai.Message) ai.Message {
return ai.Message{
Role: "system",
Content: fmt.Sprintf("Conversation memory summary: %s", summarizeMessages(msgs)),
}
}
func summarizeMessages(msgs []ai.Message) string {
var b strings.Builder
for i, msg := range msgs {
+75
View File
@@ -3,9 +3,11 @@ package agent
import (
"context"
"errors"
"strconv"
"strings"
"testing"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/registry"
"go-micro.dev/v6/store"
)
@@ -62,6 +64,49 @@ func TestWithMemoryUsed(t *testing.T) {
}
}
func TestRetrievalMemoryArchivesAllTurnsAndRanksRelevant(t *testing.T) {
st := store.NewMemoryStore()
m := NewRetrievalMemory(st, "agent/retrieval/history", 2)
m.Add("user", "alpha budget is 42")
m.Add("assistant", "noted")
m.Add("user", "beta owner is lee")
m.Add("assistant", "tracked")
m.Add("user", "alpha owner is sam")
if got := len(m.Messages()); got != 2 {
t.Fatalf("active messages = %d, want bounded history of 2", got)
}
recall, ok := m.(MemoryRecall)
if !ok {
t.Fatal("retrieval memory should support recall")
}
recalled := recall.Recall("alpha budget", 2)
if len(recalled) == 0 {
t.Fatal("expected relevant recalled turns")
}
if got := recalled[0].Content.(string); !strings.Contains(got, "alpha budget is 42") {
t.Fatalf("top recall = %q, want archived alpha budget turn", got)
}
}
func TestRetrievalMemoryPersistsArchiveAcrossReload(t *testing.T) {
st := store.NewMemoryStore()
m := NewRetrievalMemory(st, "agent/retrieval/reload", 1)
m.Add("user", "alpha budget is 42")
m.Add("assistant", "noted")
m.Add("user", "beta budget is 7")
reloaded := NewRetrievalMemory(st, "agent/retrieval/reload", 1)
recalled := reloaded.(MemoryRecall).Recall("alpha budget", 1)
if len(recalled) != 1 {
t.Fatalf("recalled %d messages, want 1", len(recalled))
}
if got := recalled[0].Content.(string); !strings.Contains(got, "alpha budget is 42") {
t.Fatalf("reloaded recall = %q, want alpha budget", got)
}
}
func TestCompactingMemoryRecallRanksSpecificMatches(t *testing.T) {
m := NewCompactingMemory(store.NewMemoryStore(), "agent/rank/history", 3, 1).(MemoryRecall)
writer := m.(Memory)
@@ -102,6 +147,36 @@ func TestCompactingMemoryArchivePersistsAndReloads(t *testing.T) {
}
}
func TestCompactingMemoryUsesCustomSummarizerAndReloadsRecall(t *testing.T) {
st := store.NewMemoryStore()
m := NewCompactingMemoryWithOptions(st, "agent/custom/history", MemoryCompaction{
MaxMessages: 3,
KeepRecent: 1,
Summarize: func(msgs []ai.Message) ai.Message {
return ai.Message{Role: "system", Content: "custom summary count=" + strconv.Itoa(len(msgs))}
},
})
m.Add("user", "alpha budget is 42")
m.Add("assistant", "noted")
m.Add("user", "beta budget is 7")
m.Add("assistant", "noted")
msgs := m.Messages()
if len(msgs) == 0 || msgs[0].Content != "custom summary count=3" {
t.Fatalf("summary = %#v, want custom summarizer output", msgs)
}
reloaded := NewCompactingMemoryWithOptions(st, "agent/custom/history", MemoryCompaction{MaxMessages: 3, KeepRecent: 1})
recall := reloaded.(MemoryRecall)
recalled := recall.Recall("alpha budget", 1)
if len(recalled) != 1 {
t.Fatalf("recalled %d messages, want 1", len(recalled))
}
if got := recalled[0].Content.(string); !strings.Contains(got, "alpha budget is 42") {
t.Fatalf("reloaded recall = %q, want alpha budget", got)
}
}
// A custom tool is offered to the model and dispatched to its handler.
func TestWithToolExposedAndDispatched(t *testing.T) {
var got map[string]any
+43 -1
View File
@@ -60,10 +60,19 @@ type Options struct {
// applied before custom tools, delegate, and service RPC calls so context
// deadlines propagate consistently through the agent loop.
ToolTimeout time.Duration
// ToolMaxAttempts bounds tool execution attempts including the first call.
// Default 1; retries are opt-in because tools can have side effects.
ToolMaxAttempts int
// ToolRetryBackoff is the base delay between transient tool failures.
ToolRetryBackoff time.Duration
// Memory is the agent's conversation memory. Nil = the default
// store-backed memory (durable across restarts).
Memory Memory
// MemoryRetrievalLimit enables retrieval-backed default memory without
// compaction. The active conversation stays bounded to this many messages
// while every turn is archived for deterministic recall.
MemoryRetrievalLimit int
// MemoryCompaction enables deterministic compaction/retrieval on the
// default store-backed memory. Custom Memory implementations can expose
// retrieval by implementing MemoryRecall.
@@ -117,6 +126,8 @@ func newOptions(opts ...Option) Options {
ModelMaxAttempts: 1, // retries opt-in via ModelRetry (see field doc)
ModelRetryBackoff: 100 * time.Millisecond,
ToolTimeout: 30 * time.Second,
ToolMaxAttempts: 1,
ToolRetryBackoff: 100 * time.Millisecond,
// On by default and lenient: identical repeated calls are a
// no-progress loop, never useful. Set LoopLimit(0) to disable.
LoopLimit: 3,
@@ -225,6 +236,16 @@ func ModelRetry(maxAttempts int, backoff time.Duration) Option {
}
}
// ToolRetry sets the tool retry budget and backoff for transient failures.
// Attempts include the first call. Retries are opt-in because tools may have
// side effects; keep handlers idempotent before enabling this.
func ToolRetry(maxAttempts int, backoff time.Duration) Option {
return func(o *Options) {
o.ToolMaxAttempts = maxAttempts
o.ToolRetryBackoff = backoff
}
}
// WithA2A makes Run serve the agent over the A2A protocol on addr (e.g.
// ":4000"), so other agents can reach it directly by URL without a
// separate gateway. The agent stays a normal go-micro service as well;
@@ -240,19 +261,40 @@ func WithMemory(m Memory) Option {
return func(o *Options) { o.Memory = m }
}
// RetrievalMemory enables deterministic, store-backed retrieval memory for
// the default agent memory without compaction. Active context is capped at
// activeLimit messages while every turn is archived in the store for Recall.
func RetrievalMemory(activeLimit int) Option {
return func(o *Options) {
o.MemoryRetrievalLimit = activeLimit
if o.MemoryRecallLimit == 0 {
o.MemoryRecallLimit = 5
}
}
}
// CompactMemory enables deterministic, store-backed memory compaction for the
// default agent memory. Older turns are summarized once active context exceeds
// maxMessages, keepRecent newest turns remain verbatim, and recalled archived
// turns are injected into matching future asks.
func CompactMemory(maxMessages, keepRecent int) Option {
return func(o *Options) {
o.MemoryCompaction = MemoryCompaction{MaxMessages: maxMessages, KeepRecent: keepRecent}
o.MemoryCompaction.MaxMessages = maxMessages
o.MemoryCompaction.KeepRecent = keepRecent
if o.MemoryRecallLimit == 0 {
o.MemoryRecallLimit = 5
}
}
}
// MemorySummarizer sets the deterministic summarization hook used by the
// default compacting memory. It is optional; without it, compacted memory uses
// a provider-neutral text summary. The hook receives the older messages being
// removed from active context and returns the replacement summary message.
func MemorySummarizer(fn MemorySummaryFunc) Option {
return func(o *Options) { o.MemoryCompaction.Summarize = fn }
}
// MemoryRecallLimit sets how many archived turns a memory backend may inject
// into a model request for the current Ask. Use 0 to disable retrieval.
func MemoryRecallLimit(n int) Option {
+134 -50
View File
@@ -22,22 +22,29 @@ const (
spanNameModelCall = "agent.model.call"
spanNameToolCall = "agent.tool.call"
AttrRunID = "agent.run.id"
AttrParentRunID = "agent.run.parent_id"
AttrAgentName = "agent.name"
AttrProvider = "agent.model.provider"
AttrModel = "agent.model.name"
AttrLatencyMS = "agent.latency_ms"
AttrInputTokens = "agent.tokens.input"
AttrOutputTokens = "agent.tokens.output"
AttrTotalTokens = "agent.tokens.total"
AttrAttempt = "agent.model.attempt"
AttrMaxAttempts = "agent.model.max_attempts"
AttrToolName = "agent.tool.name"
AttrDelegate = "agent.delegate"
AttrGuardrailBlock = "agent.guardrail.block"
AttrRefusal = "agent.refusal"
AttrInputChars = "agent.input.chars"
AttrRunID = "agent.run.id"
AttrParentRunID = "agent.run.parent_id"
AttrAgentName = "agent.name"
AttrProvider = "agent.model.provider"
AttrModel = "agent.model.name"
AttrLatencyMS = "agent.latency_ms"
AttrInputTokens = "agent.tokens.input"
AttrOutputTokens = "agent.tokens.output"
AttrTotalTokens = "agent.tokens.total"
AttrAttempt = "agent.model.attempt"
AttrMaxAttempts = "agent.model.max_attempts"
AttrToolName = "agent.tool.name"
AttrDelegate = "agent.delegate"
AttrGuardrailBlock = "agent.guardrail.block"
AttrRefusal = "agent.refusal"
AttrInputChars = "agent.input.chars"
AttrErrorKind = "agent.error.kind"
AttrCheckpointStatus = "agent.checkpoint.status"
AttrCheckpointStage = "agent.checkpoint.stage"
AttrFlowName = "agent.flow.name"
AttrFlowStep = "agent.flow.step"
AttrDispatch = "agent.dispatch"
AttrTrigger = "agent.trigger"
)
type RunEvent struct {
@@ -56,7 +63,9 @@ type RunEvent struct {
LatencyMS int64 `json:"latency_ms,omitempty"`
Tokens Usage `json:"tokens,omitempty"`
Refused string `json:"refused,omitempty"`
Status string `json:"status,omitempty"`
Error string `json:"error,omitempty"`
ErrorKind string `json:"error_kind,omitempty"`
InputChars int `json:"input_chars,omitempty"`
}
@@ -66,7 +75,8 @@ type Usage = ai.Usage
// Zero values preserve the full deterministic run list.
type RunListOptions struct {
// Status, when set, keeps only runs with the matching status
// (for example "running", "done", "error", or "refused").
// (for example "running", "done", "canceled", "timeout",
// "rate_limited", "error", or "refused").
Status string
// TraceID, when set, keeps only runs correlated with this trace id.
// A prefix is accepted so operators can paste the shortened trace id
@@ -79,18 +89,19 @@ type RunListOptions struct {
// RunSummary is a compact index entry for a recorded agent run.
type RunSummary struct {
RunID string `json:"run_id"`
Agent string `json:"agent"`
ParentID string `json:"parent_id,omitempty"`
TraceID string `json:"trace_id,omitempty"`
SpanID string `json:"span_id,omitempty"`
StartedAt time.Time `json:"started_at"`
UpdatedAt time.Time `json:"updated_at"`
DurationMS int64 `json:"duration_ms,omitempty"`
Events int `json:"events"`
Status string `json:"status,omitempty"`
LastKind string `json:"last_kind,omitempty"`
LastError string `json:"last_error,omitempty"`
RunID string `json:"run_id"`
Agent string `json:"agent"`
ParentID string `json:"parent_id,omitempty"`
TraceID string `json:"trace_id,omitempty"`
SpanID string `json:"span_id,omitempty"`
StartedAt time.Time `json:"started_at"`
UpdatedAt time.Time `json:"updated_at"`
DurationMS int64 `json:"duration_ms,omitempty"`
Events int `json:"events"`
Status string `json:"status,omitempty"`
LastKind string `json:"last_kind,omitempty"`
LastError string `json:"last_error,omitempty"`
LastErrorKind string `json:"last_error_kind,omitempty"`
}
func (a *agentImpl) tracer() trace.Tracer {
@@ -110,23 +121,28 @@ func (a *agentImpl) startRun(ctx context.Context, message string) (context.Conte
return ctx, func(err error) {
latency := time.Since(start).Milliseconds()
if err != nil {
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "error", LatencyMS: latency, Error: err.Error()})
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "error", LatencyMS: latency, Error: err.Error(), ErrorKind: string(ai.ClassifyError(err))})
return
}
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "done", LatencyMS: latency})
}
}
ctx, span := a.tracer().Start(ctx, spanNameRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(
attribute.String(AttrRunID, info.RunID), attribute.String(AttrParentRunID, info.ParentID), attribute.String(AttrAgentName, info.Agent)))
attrs := appendRunInfoAttributes([]attribute.KeyValue{
attribute.String(AttrRunID, info.RunID),
attribute.String(AttrParentRunID, info.ParentID),
attribute.String(AttrAgentName, info.Agent),
}, info)
ctx, span := a.tracer().Start(ctx, spanNameRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(attrs...))
a.recordSpanEvent(span, runEvent)
return ctx, func(err error) {
latency := time.Since(start).Milliseconds()
span.SetAttributes(attribute.Int64(AttrLatencyMS, latency))
if err != nil {
span.SetAttributes(attribute.String(AttrErrorKind, string(ai.ClassifyError(err))))
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "error", LatencyMS: latency, Error: err.Error()})
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "error", LatencyMS: latency, Error: err.Error(), ErrorKind: string(ai.ClassifyError(err))})
} else {
span.SetStatus(codes.Ok, "")
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "done", LatencyMS: latency})
@@ -157,21 +173,23 @@ func (m *tracedModel) Generate(ctx context.Context, req *ai.Request, opts ...ai.
e := RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "model", Provider: provider, Model: model, Attempt: info.Attempt, MaxAttempts: info.MaxAttempts, LatencyMS: dur, Tokens: usage}
if err != nil {
e.Error = err.Error()
e.ErrorKind = string(ai.ClassifyError(err))
}
m.a.recordRunEvent(e)
return resp, err
}
ctx, span := m.a.tracer().Start(ctx, spanNameModelCall, trace.WithAttributes(
attrs := appendRunInfoAttributes([]attribute.KeyValue{
attribute.String(AttrRunID, info.RunID),
attribute.String(AttrParentRunID, info.ParentID),
attribute.String(AttrAgentName, info.Agent),
attribute.String(AttrProvider, provider),
attribute.String(AttrModel, model),
))
}, info)
ctx, span := m.a.tracer().Start(ctx, spanNameModelCall, trace.WithAttributes(attrs...))
resp, err := m.Model.Generate(ctx, req, opts...)
dur := time.Since(start).Milliseconds()
attrs := []attribute.KeyValue{attribute.Int64(AttrLatencyMS, dur)}
attrs = []attribute.KeyValue{attribute.Int64(AttrLatencyMS, dur)}
if info.Attempt > 0 {
attrs = append(attrs, attribute.Int(AttrAttempt, info.Attempt))
}
@@ -185,6 +203,7 @@ func (m *tracedModel) Generate(ctx context.Context, req *ai.Request, opts ...ai.
}
span.SetAttributes(attrs...)
if err != nil {
span.SetAttributes(attribute.String(AttrErrorKind, string(ai.ClassifyError(err))))
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
} else {
@@ -194,6 +213,7 @@ func (m *tracedModel) Generate(ctx context.Context, req *ai.Request, opts ...ai.
e := RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "model", Provider: provider, Model: model, Attempt: info.Attempt, MaxAttempts: info.MaxAttempts, LatencyMS: dur, Tokens: usage}
if err != nil {
e.Error = err.Error()
e.ErrorKind = string(ai.ClassifyError(err))
}
m.a.recordSpanEvent(span, e)
return resp, err
@@ -220,7 +240,8 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
if a.opts.TraceProvider == nil {
res := next(ctx, call)
dur := time.Since(start).Milliseconds()
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resultError(res)})
resErr := resultError(res)
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
return res
}
@@ -237,8 +258,11 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
if res.Refused != "" {
attrs = append(attrs, attribute.Bool(AttrGuardrailBlock, true), attribute.String(AttrRefusal, res.Refused))
}
span.SetAttributes(attrs...)
resErr := resultError(res)
if kind := classifyToolError(resErr); kind != "" {
attrs = append(attrs, attribute.String(AttrErrorKind, kind))
}
span.SetAttributes(attrs...)
if res.Refused != "" {
span.SetStatus(codes.Error, res.Refused)
} else if resErr != "" {
@@ -247,7 +271,7 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
span.SetStatus(codes.Ok, "")
}
span.End()
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resErr})
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
return res
}
}
@@ -264,6 +288,28 @@ func resultError(res ai.ToolResult) string {
return ""
}
func classifyToolError(err string) string {
switch {
case err == "":
return ""
case strings.Contains(strings.ToLower(err), "context canceled"):
return string(ai.ErrorKindCanceled)
case strings.Contains(strings.ToLower(err), "deadline exceeded"):
return string(ai.ErrorKindTimeout)
default:
return string(ai.ErrorKindProvider)
}
}
func (a *agentImpl) recordTimelineEvent(ctx context.Context, e RunEvent) {
span := trace.SpanFromContext(ctx)
if span.SpanContext().IsValid() {
a.recordSpanEvent(span, e)
return
}
a.recordRunEvent(e)
}
func (a *agentImpl) recordSpanEvent(span trace.Span, e RunEvent) {
if sc := span.SpanContext(); sc.IsValid() {
e.TraceID = sc.TraceID().String()
@@ -309,6 +355,33 @@ func runEventAttributes(e RunEvent) []attribute.KeyValue {
if e.Error != "" {
attrs = append(attrs, attribute.String("agent.error", e.Error))
}
if e.ErrorKind != "" {
attrs = append(attrs, attribute.String(AttrErrorKind, e.ErrorKind))
}
if e.Kind == "checkpoint" {
if e.Status != "" {
attrs = append(attrs, attribute.String(AttrCheckpointStatus, e.Status))
}
if e.Name != "" {
attrs = append(attrs, attribute.String(AttrCheckpointStage, e.Name))
}
}
return attrs
}
func appendRunInfoAttributes(attrs []attribute.KeyValue, info ai.RunInfo) []attribute.KeyValue {
if info.Flow != "" {
attrs = append(attrs, attribute.String(AttrFlowName, info.Flow))
}
if info.Step != "" {
attrs = append(attrs, attribute.String(AttrFlowStep, info.Step))
}
if info.Dispatch != "" {
attrs = append(attrs, attribute.String(AttrDispatch, info.Dispatch))
}
if info.Trigger != "" {
attrs = append(attrs, attribute.String(AttrTrigger, info.Trigger))
}
return attrs
}
@@ -388,6 +461,9 @@ func ListRunSummariesWithOptions(s store.Store, agentName string, opts RunListOp
if e.Error != "" {
summary.LastError = e.Error
}
if e.ErrorKind != "" {
summary.LastErrorKind = e.ErrorKind
}
}
if opts.Status != "" && summary.Status != opts.Status {
continue
@@ -414,24 +490,32 @@ func runStatus(events []RunEvent) string {
}
status := "running"
for _, e := range events {
if e.Error != "" {
status = "error"
}
if e.Refused != "" && status != "error" {
if e.Refused != "" && status == "running" {
status = "refused"
}
switch e.Kind {
case "error":
status = "error"
case "done":
if status == "running" {
status = "done"
}
if e.Error != "" || e.Kind == "error" {
status = runErrorStatus(e.ErrorKind)
}
if e.Kind == "done" && status == "running" {
status = "done"
}
}
return status
}
func runErrorStatus(kind string) string {
switch ai.ErrorKind(kind) {
case ai.ErrorKindCanceled:
return "canceled"
case ai.ErrorKindTimeout:
return "timeout"
case ai.ErrorKindRateLimited:
return "rate_limited"
default:
return "error"
}
}
func LoadRunEvents(s store.Store, agentName, runID string) ([]RunEvent, error) {
st := store.Scope(s, "agent", agentName)
keys, err := st.List(store.ListPrefix("runs/" + runID + "/"))
+100 -7
View File
@@ -10,6 +10,7 @@ import (
"time"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/flow"
"go-micro.dev/v6/store"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/codes"
@@ -241,7 +242,7 @@ func TestAgentOpenTelemetrySpansModelFailure(t *testing.T) {
sawRunError = true
}
case spanNameModelCall:
if attrs[AttrAgentName] == "failing-runner" && attrs[AttrAttempt] == "1" && s.Status().Code == codesError {
if attrs[AttrAgentName] == "failing-runner" && attrs[AttrAttempt] == "1" && attrs[AttrErrorKind] == string(ai.ErrorKindUnknown) && s.Status().Code == codesError {
sawModelError = true
}
}
@@ -263,7 +264,7 @@ func TestAgentOpenTelemetrySpansModelFailure(t *testing.T) {
}
var sawModelEvent bool
for _, event := range events {
if event.Kind == "model" && event.Attempt == 1 && event.MaxAttempts == 1 && event.Error != "" {
if event.Kind == "model" && event.Attempt == 1 && event.MaxAttempts == 1 && event.Error != "" && event.ErrorKind == string(ai.ErrorKindUnknown) {
sawModelEvent = true
}
}
@@ -388,6 +389,74 @@ func TestAgentRunTimelineRecordsModelAndToolWithoutTraceProvider(t *testing.T) {
}
}
func TestAgentCheckpointAndResumeTimelineEvents(t *testing.T) {
exp := tracetest.NewInMemoryExporter()
tp := trace.NewTracerProvider(trace.WithSyncer(exp))
st := store.NewMemoryStore()
cp := flow.StoreCheckpoint(st, "resume-otel-agent")
first := true
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
if first {
first = false
return nil, errors.New("temporary provider failure")
}
return &ai.Response{Reply: "resumed"}, nil
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("resume-otel-agent"), WithStore(st), WithCheckpoint(cp), TraceProvider(tp))
_, err := a.Ask(context.Background(), "resume me")
if err == nil {
t.Fatal("Ask succeeded, want simulated failure")
}
runs, err := cp.List(context.Background())
if err != nil {
t.Fatal(err)
}
if len(runs) != 1 {
t.Fatalf("checkpointed runs = %d, want 1", len(runs))
}
resp, err := Resume(context.Background(), a, runs[0].ID)
if err != nil {
t.Fatalf("Resume: %v", err)
}
if resp.Reply != "resumed" {
t.Fatalf("reply = %q, want resumed", resp.Reply)
}
events, err := LoadRunEvents(st, "resume-otel-agent", runs[0].ID)
if err != nil {
t.Fatal(err)
}
seen := map[string]bool{"checkpoint": false, "resume": false}
for _, e := range events {
if _, ok := seen[e.Kind]; ok {
seen[e.Kind] = true
}
}
for kind, ok := range seen {
if !ok {
t.Fatalf("missing %s event in timeline: %#v", kind, events)
}
}
var resumeSpanEvent bool
for _, s := range exp.GetSpans().Snapshots() {
if s.Name() != spanNameRun {
continue
}
for _, e := range s.Events() {
if e.Name == "agent.resume" {
resumeSpanEvent = true
}
}
}
if !resumeSpanEvent {
t.Fatal("run span missing agent.resume event")
}
}
func TestLoadRunEventsSortsTimelineKeys(t *testing.T) {
st := store.NewMemoryStore()
scoped := store.Scope(st, "agent", "runner")
@@ -429,7 +498,7 @@ func TestListRunSummaries(t *testing.T) {
{Time: time.Unix(0, 1), RunID: "run-a", Agent: "runner", TraceID: "trace-a", SpanID: "span-a", Kind: "run", Name: "first"},
{Time: time.Unix(0, 2), RunID: "run-a", Agent: "runner", Kind: "tool", Name: "probe"},
{Time: time.Unix(0, 3), RunID: "run-b", Agent: "runner", ParentID: "parent", Kind: "run", Name: "second"},
{Time: time.Unix(0, 4), RunID: "run-b", Agent: "runner", ParentID: "parent", Kind: "error", Error: "boom"},
{Time: time.Unix(0, 4), RunID: "run-b", Agent: "runner", ParentID: "parent", Kind: "error", Error: "context deadline exceeded", ErrorKind: string(ai.ErrorKindTimeout)},
}
for _, e := range events {
b, err := json.Marshal(e)
@@ -452,11 +521,35 @@ func TestListRunSummaries(t *testing.T) {
if got[0].RunID != "run-a" || got[0].TraceID != "trace-a" || got[0].SpanID != "span-a" || got[0].Events != 2 || got[0].Status != "running" || got[0].DurationMS != 0 || got[0].LastKind != "tool" || !got[0].UpdatedAt.Equal(time.Unix(0, 2)) {
t.Fatalf("unexpected run-a summary: %#v", got[0])
}
if got[1].RunID != "run-b" || got[1].ParentID != "parent" || got[1].Events != 2 || got[1].Status != "error" || got[1].DurationMS != 0 || got[1].LastKind != "error" || got[1].LastError != "boom" {
if got[1].RunID != "run-b" || got[1].ParentID != "parent" || got[1].Events != 2 || got[1].Status != "timeout" || got[1].DurationMS != 0 || got[1].LastKind != "error" || got[1].LastError != "context deadline exceeded" || got[1].LastErrorKind != string(ai.ErrorKindTimeout) {
t.Fatalf("unexpected run-b summary: %#v", got[1])
}
}
func TestRunStatusClassifiesOperationalErrorKinds(t *testing.T) {
tests := []struct {
name string
kind ai.ErrorKind
want string
}{
{name: "canceled", kind: ai.ErrorKindCanceled, want: "canceled"},
{name: "timeout", kind: ai.ErrorKindTimeout, want: "timeout"},
{name: "rate limited", kind: ai.ErrorKindRateLimited, want: "rate_limited"},
{name: "provider", kind: ai.ErrorKindProvider, want: "error"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got := runStatus([]RunEvent{
{Kind: "run"},
{Kind: "error", Error: "failed", ErrorKind: string(tt.kind)},
})
if got != tt.want {
t.Fatalf("runStatus() = %q, want %q", got, tt.want)
}
})
}
}
func TestListRunSummariesWithOptionsFiltersAndLimits(t *testing.T) {
st := store.NewMemoryStore()
scoped := store.Scope(st, "agent", "runner")
@@ -464,7 +557,7 @@ func TestListRunSummariesWithOptionsFiltersAndLimits(t *testing.T) {
{Time: time.Unix(0, 1), RunID: "run-old", Agent: "runner", Kind: "run"},
{Time: time.Unix(0, 2), RunID: "run-old", Agent: "runner", Kind: "done"},
{Time: time.Unix(0, 3), RunID: "run-new", Agent: "runner", TraceID: "abcdef1234567890", Kind: "run"},
{Time: time.Unix(0, 4), RunID: "run-new", Agent: "runner", Kind: "error", Error: "boom"},
{Time: time.Unix(0, 4), RunID: "run-new", Agent: "runner", Kind: "error", Error: "rate limit exceeded", ErrorKind: string(ai.ErrorKindRateLimited)},
}
for _, e := range events {
b, err := json.Marshal(e)
@@ -476,11 +569,11 @@ func TestListRunSummariesWithOptionsFiltersAndLimits(t *testing.T) {
}
}
got, err := ListRunSummariesWithOptions(st, "runner", RunListOptions{Status: "error", TraceID: "abcdef", Limit: 1})
got, err := ListRunSummariesWithOptions(st, "runner", RunListOptions{Status: "rate_limited", TraceID: "abcdef", Limit: 1})
if err != nil {
t.Fatal(err)
}
if len(got) != 1 || got[0].RunID != "run-new" || got[0].Status != "error" {
if len(got) != 1 || got[0].RunID != "run-new" || got[0].Status != "rate_limited" {
t.Fatalf("filtered summaries = %#v", got)
}
}
+100
View File
@@ -8,6 +8,8 @@ import (
"time"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/flow"
"go-micro.dev/v6/store"
)
func TestAskCancellationAbortsPromptly(t *testing.T) {
@@ -129,3 +131,101 @@ func TestToolCallTimeoutPropagatesDeadlineToCustomTool(t *testing.T) {
t.Fatalf("tool call took %s, want bounded timeout", elapsed)
}
}
func TestAskCheckpointRecordsTerminalOperationalFailureStatus(t *testing.T) {
tests := []struct {
name string
err error
want string
}{
{name: "canceled", err: context.Canceled, want: "canceled"},
{name: "timeout", err: context.DeadlineExceeded, want: "timeout"},
{name: "rate limited", err: testStatusError{code: 429}, want: "rate_limited"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "terminal-"+strings.ReplaceAll(tt.name, " ", "-"))
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
return nil, tt.err
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("terminal-"+strings.ReplaceAll(tt.name, " ", "-")), WithCheckpoint(cp))
_, err := a.Ask(context.Background(), "fail safely")
if err == nil {
t.Fatal("Ask succeeded, want failure")
}
runs, err := cp.List(context.Background())
if err != nil {
t.Fatalf("List: %v", err)
}
if len(runs) != 1 {
t.Fatalf("checkpointed runs = %d, want 1", len(runs))
}
if runs[0].Status != tt.want {
t.Fatalf("run status = %q, want %q", runs[0].Status, tt.want)
}
if len(runs[0].Steps) == 0 || runs[0].Steps[0].Status != tt.want {
t.Fatalf("step status = %#v, want %q", runs[0].Steps, tt.want)
}
if pending, err := Pending(context.Background(), a); err != nil || len(pending) != 0 {
t.Fatalf("Pending = %#v, %v; want no terminal run", pending, err)
}
})
}
}
type testStatusError struct {
code int
}
func (e testStatusError) Error() string { return "provider status error" }
func (e testStatusError) StatusCode() int { return e.code }
func TestToolRetryRetriesTransientToolErrorsThenSucceeds(t *testing.T) {
attempts := 0
a := newTestAgent(
Name("tool-retry-success"),
ToolRetry(3, time.Millisecond),
WithTool("flaky", "flaky tool", nil, func(context.Context, map[string]any) (string, error) {
attempts++
if attempts < 3 {
return "", context.DeadlineExceeded
}
return "ok", nil
}),
)
content := toolContent(a.toolHandler(), "flaky", nil)
if content != "ok" {
t.Fatalf("tool result = %q, want ok", content)
}
if attempts != 3 {
t.Fatalf("attempts = %d, want 3", attempts)
}
}
func TestToolRetryDoesNotRetryGuardrailRefusals(t *testing.T) {
attempts := 0
a := newTestAgent(
Name("tool-retry-refusal"),
MaxSteps(1),
ToolRetry(3, time.Millisecond),
WithTool("counted", "counted tool", nil, func(context.Context, map[string]any) (string, error) {
attempts++
return "ok", nil
}),
)
h := a.toolHandler()
_ = toolContent(h, "counted", nil)
content := toolContent(h, "counted", nil)
if !strings.Contains(content, "step limit reached") {
t.Fatalf("tool result = %q, want step-limit refusal", content)
}
if attempts != 1 {
t.Fatalf("attempts = %d, want only the allowed tool call to execute", attempts)
}
}
+37
View File
@@ -181,6 +181,43 @@ func (a *agentImpl) resumeWithStreamEvents(ctx context.Context, runID string, ev
return a.askLocked(ctx, run.ID, string(run.State.Data), run.ParentID, &run, false)
}
type agentStreamAdapter struct {
stream AgentStream
}
func (s *agentStreamAdapter) Recv() (*ai.Response, error) {
for {
event, err := s.stream.Recv()
if err != nil {
return nil, err
}
if event == nil {
continue
}
switch event.Type {
case StreamEventToken:
if event.Token == "" {
continue
}
return &ai.Response{Reply: event.Token}, nil
case StreamEventDone:
return nil, io.EOF
}
}
}
func (s *agentStreamAdapter) Close() error {
return s.stream.Close()
}
func (a *agentImpl) streamAskAI(ctx context.Context, message string) (ai.Stream, error) {
stream, err := a.StreamAsk(ctx, message)
if err != nil {
return nil, err
}
return &agentStreamAdapter{stream: stream}, nil
}
type agentStream struct {
events <-chan *StreamEvent
done <-chan struct{}
+14 -10
View File
@@ -93,9 +93,10 @@ func (c ToolCall) Scan(v any) error {
// ToolResult represents the result of a tool execution
type ToolResult struct {
ID string // Tool call ID (for correlation)
Value any // Structured result (optional)
Content string // Tool execution result (JSON string), shown to the model
ID string // Tool call ID (for correlation)
Value any // Structured result (optional)
Content string // Tool execution result (JSON string), shown to the model
Attempts int `json:"attempts,omitempty"` // Tool execution attempts, set when retried.
// Refused names the reason a guardrail blocked the call before it ran
// ("max_steps", "loop", "approval"); empty when the call executed. A
// tool wrapper can switch on it to build reliability tooling — react to
@@ -121,13 +122,16 @@ const (
// tell which provider attempt produced the call and whether it is part of a
// retry budget. They are zero when no model-attempt context is known.
type RunInfo struct {
RunID string // correlation id for this agent or flow run
ParentID string // the run that delegated to this one, if any
Agent string // the agent's name
Flow string // the flow's name, when the call is part of a workflow
Step string // the flow step currently executing, when known
Attempt int // current model Generate attempt, starting at 1 when known
MaxAttempts int // configured model Generate attempt budget when known
RunID string // correlation id for this agent or flow run
ParentID string // the run that delegated to this one, if any
Agent string // the agent's name
Flow string // the flow's name, when the call is part of a workflow
Step string // the flow step currently executing, when known
Attempt int // current model Generate attempt, starting at 1 when known
MaxAttempts int // configured model Generate attempt budget when known
VerificationFeedback string // feedback from the previous failed verifier attempt, when retrying a flow step
Dispatch string // how the run was dispatched (direct, broker, schedule, resume) when known
Trigger string // external trigger or schedule label that started the run, when known
}
type runInfoKey struct{}
+307
View File
@@ -0,0 +1,307 @@
package ai_test
import (
"context"
"encoding/json"
"errors"
"io"
"net/http"
"net/http/httptest"
"os"
"reflect"
"strings"
"testing"
"time"
"go-micro.dev/v6/ai"
_ "go-micro.dev/v6/ai/anthropic"
_ "go-micro.dev/v6/ai/atlascloud"
_ "go-micro.dev/v6/ai/gemini"
_ "go-micro.dev/v6/ai/groq"
_ "go-micro.dev/v6/ai/mistral"
_ "go-micro.dev/v6/ai/openai"
_ "go-micro.dev/v6/ai/together"
)
func TestStreamProvidersConformToOpenAICompatibleSSE(t *testing.T) {
providers := conformingStreamProviders(t)
for _, provider := range providers {
provider := provider
t.Run(provider, func(t *testing.T) {
var sawRequest bool
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
sawRequest = true
if r.URL.Path != "/v1/chat/completions" {
t.Fatalf("path = %s, want /v1/chat/completions", r.URL.Path)
}
if got := r.Header.Get("Accept"); got != "text/event-stream" {
t.Fatalf("Accept = %q, want text/event-stream", got)
}
if got := r.Header.Get("Authorization"); got != "Bearer test-key" {
t.Fatalf("Authorization = %q, want bearer API key", got)
}
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
t.Fatalf("decode request: %v", err)
}
if body["model"] == "" {
t.Fatal("request omitted model")
}
if body["stream"] != true {
t.Fatalf("stream = %#v, want true", body["stream"])
}
streamOptions, ok := body["stream_options"].(map[string]any)
if !ok || streamOptions["include_usage"] != true {
t.Fatalf("stream_options = %#v, want include_usage=true", body["stream_options"])
}
messages, ok := body["messages"].([]any)
if !ok || len(messages) != 4 {
t.Fatalf("messages = %#v, want system + history + prompt", body["messages"])
}
wantRoles := []string{"system", "user", "assistant", "user"}
for i, wantRole := range wantRoles {
message, ok := messages[i].(map[string]any)
if !ok || message["role"] != wantRole {
t.Fatalf("message[%d] = %#v, want role %q", i, messages[i], wantRole)
}
}
w.Header().Set("Content-Type", "text/event-stream")
_, _ = w.Write([]byte(": keepalive\n\n"))
_, _ = w.Write([]byte("event: ignored\n\n"))
_, _ = w.Write([]byte("data: {\"choices\":[{\"delta\":{\"content\":\"hel\"}}]}\n\n"))
_, _ = w.Write([]byte("data: {\"choices\":[{\"delta\":{\"content\":\"lo\"}}]}\n\n"))
_, _ = w.Write([]byte("data: {\"choices\":[],\"usage\":{\"prompt_tokens\":3,\"completion_tokens\":2,\"total_tokens\":5}}\n\n"))
_, _ = w.Write([]byte("data: [DONE]\n\n"))
}))
defer ts.Close()
model := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL))
if model == nil {
t.Fatalf("ai.New(%q) returned nil", provider)
}
stream, err := model.Stream(context.Background(), &ai.Request{
SystemPrompt: "system",
Messages: []ai.Message{
{Role: "user", Content: "previous question"},
{Role: "assistant", Content: "previous answer"},
},
Prompt: "current question",
})
if err != nil {
t.Fatalf("Stream returned error: %v", err)
}
defer stream.Close()
if !sawRequest {
t.Fatal("server did not receive stream request")
}
assertStreamReply(t, stream, "hel")
assertStreamReply(t, stream, "lo")
usage, err := stream.Recv()
if err != nil {
t.Fatalf("usage chunk error: %v", err)
}
if usage.Reply != "" || usage.Usage != (ai.Usage{InputTokens: 3, OutputTokens: 2, TotalTokens: 5}) {
t.Fatalf("usage chunk = %#v", usage)
}
if _, err := stream.Recv(); !errors.Is(err, io.EOF) {
t.Fatalf("final error = %v, want EOF", err)
}
})
}
}
func TestStreamProvidersCloseCancelsInFlightRequest(t *testing.T) {
for _, provider := range conformingStreamProviders(t) {
provider := provider
t.Run(provider, func(t *testing.T) {
released := make(chan struct{})
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "text/event-stream")
_, _ = w.Write([]byte("data: {\"choices\":[{\"delta\":{\"content\":\"hel\"}}]}\n\n"))
if f, ok := w.(http.Flusher); ok {
f.Flush()
}
<-r.Context().Done()
close(released)
}))
defer ts.Close()
stream, err := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL)).Stream(context.Background(), &ai.Request{Prompt: "Hello"})
if err != nil {
t.Fatalf("Stream returned error: %v", err)
}
assertStreamReply(t, stream, "hel")
if err := stream.Close(); err != nil {
t.Fatalf("Close returned error: %v", err)
}
if err := stream.Close(); err != nil {
t.Fatalf("second Close returned error: %v", err)
}
select {
case <-released:
case <-time.After(time.Second):
t.Fatal("server did not observe canceled stream request")
}
})
}
}
func TestStreamProvidersPropagateProviderErrors(t *testing.T) {
for _, provider := range conformingStreamProviders(t) {
provider := provider
t.Run(provider, func(t *testing.T) {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "upstream quota exhausted", http.StatusTooManyRequests)
}))
defer ts.Close()
stream, err := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL)).Stream(context.Background(), &ai.Request{Prompt: "Hello"})
if err == nil {
_ = stream.Close()
t.Fatal("Stream returned nil error for provider failure")
}
if !strings.Contains(err.Error(), "429") || !strings.Contains(err.Error(), "upstream quota exhausted") {
t.Fatalf("Stream error = %v, want provider status and body", err)
}
if strings.Contains(err.Error(), "test-key") {
t.Fatal("provider error leaked API key")
}
})
}
}
func TestStreamProvidersHonorCanceledContextBeforeRequest(t *testing.T) {
for _, provider := range conformingStreamProviders(t) {
provider := provider
t.Run(provider, func(t *testing.T) {
var sawRequest bool
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
sawRequest = true
http.Error(w, "unexpected request", http.StatusInternalServerError)
}))
defer ts.Close()
ctx, cancel := context.WithCancel(context.Background())
cancel()
stream, err := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL)).Stream(ctx, &ai.Request{Prompt: "Hello"})
if err == nil {
_ = stream.Close()
t.Fatal("Stream returned nil error for canceled context")
}
if !errors.Is(err, context.Canceled) {
t.Fatalf("Stream error = %v, want context.Canceled", err)
}
if sawRequest {
t.Fatal("provider sent request after context was already canceled")
}
})
}
}
func TestConfiguredProviderStreamsSkipWithoutCredentials(t *testing.T) {
for _, tc := range []struct {
provider string
keyEnv string
modelEnv string
}{
{provider: "openai", keyEnv: "OPENAI_API_KEY", modelEnv: "OPENAI_MODEL"},
{provider: "groq", keyEnv: "GROQ_API_KEY", modelEnv: "GROQ_MODEL"},
{provider: "mistral", keyEnv: "MISTRAL_API_KEY", modelEnv: "MISTRAL_MODEL"},
{provider: "together", keyEnv: "TOGETHER_API_KEY", modelEnv: "TOGETHER_MODEL"},
{provider: "atlascloud", keyEnv: "ATLASCLOUD_API_KEY", modelEnv: "ATLASCLOUD_MODEL"},
} {
tc := tc
t.Run(tc.provider, func(t *testing.T) {
key := os.Getenv(tc.keyEnv)
if key == "" {
t.Skipf("%s not set; skipping configured provider stream check", tc.keyEnv)
}
opts := []ai.Option{ai.WithAPIKey(key)}
if model := os.Getenv(tc.modelEnv); model != "" {
opts = append(opts, ai.WithModel(model))
}
stream, err := ai.New(tc.provider, opts...).Stream(context.Background(), &ai.Request{Prompt: "Reply with exactly: ok"})
if err != nil {
t.Fatalf("Stream returned error: %v", err)
}
defer stream.Close()
deadline := time.After(30 * time.Second)
for {
select {
case <-deadline:
t.Fatal("timed out waiting for provider stream chunk")
default:
}
chunk, err := stream.Recv()
if err != nil {
if errors.Is(err, io.EOF) {
t.Fatal("provider stream ended without content")
}
t.Fatalf("Recv returned error: %v", err)
}
if chunk.Reply != "" {
return
}
}
})
}
}
func TestUnsupportedProvidersReturnStreamingUnsupportedAndStayUnregistered(t *testing.T) {
for _, provider := range []string{"anthropic", "gemini"} {
provider := provider
t.Run(provider, func(t *testing.T) {
if caps := ai.ProviderCapabilities(provider); caps.Stream {
t.Fatalf("ProviderCapabilities(%q).Stream = true, want false", provider)
}
_, err := ai.New(provider, ai.WithAPIKey("test-key")).Stream(context.Background(), &ai.Request{Prompt: "Hello"})
if !errors.Is(err, ai.ErrStreamingUnsupported) {
t.Fatalf("Stream error = %v, want ErrStreamingUnsupported", err)
}
if err != nil && strings.Contains(err.Error(), "test-key") {
t.Fatal("streaming unsupported error leaked API key")
}
})
}
}
func conformingStreamProviders(t *testing.T) []string {
t.Helper()
providers := ai.RegisteredProviders("stream")
allowed := map[string]struct{}{
"atlascloud": {},
"groq": {},
"mistral": {},
"openai": {},
"together": {},
}
var out []string
for _, provider := range providers {
if _, ok := allowed[provider]; ok {
out = append(out, provider)
}
}
want := []string{"atlascloud", "groq", "mistral", "openai", "together"}
if !reflect.DeepEqual(out, want) {
t.Fatalf("conforming stream providers = %#v, want %#v (registered stream providers: %#v)", out, want, providers)
}
return out
}
func assertStreamReply(t *testing.T, stream ai.Stream, want string) {
t.Helper()
chunk, err := stream.Recv()
if err != nil {
t.Fatalf("Recv error = %v, want reply %q", err, want)
}
if chunk.Reply != want {
t.Fatalf("Reply = %q, want %q", chunk.Reply, want)
}
}
+5 -2
View File
@@ -498,10 +498,13 @@ func printBanner(services []*serviceProcess, gw *server.Gateway, watching bool,
fmt.Printf(" Dashboard \033[36mhttp://localhost%s\033[0m\n", gw.Addr())
fmt.Printf(" API \033[36mhttp://localhost%s/api/{service}/{method}\033[0m\n", gw.Addr())
fmt.Printf(" Agent \033[36mhttp://localhost%s/agent\033[0m\n", gw.Addr())
// MCP tools are served on the gateway by default — every endpoint is an
// AI-callable tool, so surface it rather than hiding it behind a flag.
fmt.Printf(" MCP Tools \033[36mhttp://localhost%s/mcp/tools\033[0m\n", gw.Addr())
fmt.Printf(" Health \033[36mhttp://localhost%s/health\033[0m\n", gw.Addr())
if mcpAddr != "" {
fmt.Printf(" MCP \033[36mhttp://localhost%s\033[0m\n", mcpAddr)
fmt.Printf(" MCP Tools \033[36mhttp://localhost%s/mcp/tools\033[0m\n", mcpAddr)
// Optional standalone MCP protocol server (e.g. for MCP clients).
fmt.Printf(" MCP Server \033[36mhttp://localhost%s\033[0m (full MCP protocol)\n", mcpAddr)
fmt.Printf(" WebSocket \033[36mws://localhost%s/mcp/ws\033[0m\n", mcpAddr)
}
}
+45
View File
@@ -0,0 +1,45 @@
# Durable agent run resume
This example shows the agent-side counterpart to `examples/flow-durable`: an
agent run is checkpointed with the same `Checkpoint` interface used by flows,
then resumed after an interruption without repeating a completed side effect.
The sample uses an in-memory store to keep repeated local runs deterministic;
use your service store for process-restart recovery.
Run it with:
```sh
go run ./examples/agent-durable
```
The demo model calls `inventory.reserve`, then fails to mimic a process dying
after the tool call was checkpointed. `micro.AgentPending` finds the unfinished
run and `micro.AgentResume` continues it from the saved checkpoint. The final
`tool executions: 1` line is the important bit: the reservation tool was not
called a second time during resume.
## When to use this instead of a durable flow
Use a durable flow when the path is known ahead of time: ordered service calls,
retries, timers, compensation, and a precise resume stage such as `reserve` or
`charge`. Use a checkpointed agent run when the path is open-ended and the model
may choose tools dynamically, but completed tool side effects still must not be
replayed after a crash or provider failure.
They compose: keep deterministic business process in `flow-durable`, then hand
off the judgment-heavy step to a checkpointed agent when the workflow needs
model-directed tool use. Both use the same `Checkpoint` backend, so inspection
and recovery can share one run-history store.
In a service, use the same pattern at startup:
```go
pending, _ := micro.AgentPending(ctx, agent)
for _, run := range pending {
_, _ = micro.AgentResume(ctx, agent, run.ID)
}
```
`context.Context` cancellation and deadlines are still honored by checkpoint
loads/saves, model calls, and tool calls. Runs with terminal statuses such as
`done`, `canceled`, and `expired` are not returned by `AgentPending`.
+88
View File
@@ -0,0 +1,88 @@
// Package main demonstrates durable agent runs: a checkpointed agent can
// resume after a crash without re-executing completed tool calls.
package main
import (
"context"
"errors"
"fmt"
"sync/atomic"
micro "go-micro.dev/v6"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/store"
)
func main() {
ctx := context.Background()
checkpoint := micro.StoreCheckpoint(store.NewMemoryStore(), "durable-agent-demo")
model := &demoModel{failFirst: true}
ai.Register("durable-demo", func(opts ...ai.Option) ai.Model {
_ = model.Init(opts...)
return model
})
var reservations atomic.Int32
ag := micro.NewAgent("durable-agent-demo",
micro.AgentWithCheckpoint(checkpoint),
micro.AgentProvider("durable-demo"),
micro.AgentTool("inventory.reserve", "reserve inventory exactly once", map[string]any{
"sku": map[string]any{"type": "string"},
}, func(ctx context.Context, input map[string]any) (string, error) {
count := reservations.Add(1)
return fmt.Sprintf("reserved %s (execution %d)", input["sku"], count), nil
}),
)
_, err := ag.Ask(ctx, "reserve sku-123 and confirm")
fmt.Println("initial run:", err)
pending, err := micro.AgentPending(ctx, ag)
if err != nil {
panic(err)
}
if len(pending) == 0 {
panic("expected a checkpointed run to resume")
}
resp, err := micro.AgentResume(ctx, ag, pending[0].ID)
if err != nil {
panic(err)
}
fmt.Println("resumed reply:", resp.Reply)
fmt.Println("tool executions:", reservations.Load())
}
type demoModel struct {
failFirst bool
opts ai.Options
}
func (m *demoModel) Init(opts ...ai.Option) error {
m.opts = ai.NewOptions(opts...)
return nil
}
func (m *demoModel) Options() ai.Options { return m.opts }
func (m *demoModel) String() string { return "durable-demo" }
func (m *demoModel) Generate(ctx context.Context, req *ai.Request, opts ...ai.GenerateOption) (*ai.Response, error) {
if m.opts.ToolHandler != nil {
res := m.opts.ToolHandler(ctx, ai.ToolCall{
ID: "reserve-1",
Name: "inventory.reserve",
Input: map[string]any{"sku": "sku-123"},
})
if res.Content == "" {
return nil, errors.New("reservation tool returned no content")
}
}
if m.failFirst {
m.failFirst = false
return nil, errors.New("simulated process interruption after checkpointed tool call")
}
return &ai.Response{Reply: "sku-123 is reserved; no duplicate reservation was made"}, nil
}
func (m *demoModel) Stream(context.Context, *ai.Request, ...ai.GenerateOption) (ai.Stream, error) {
return nil, ai.ErrStreamingUnsupported
}
+48
View File
@@ -0,0 +1,48 @@
package main
import (
"bytes"
"io"
"os"
"strings"
"testing"
)
func TestDurableAgentExampleResumesWithoutReplayingTool(t *testing.T) {
out := captureStdout(t, main)
if !strings.Contains(out, "simulated process interruption after checkpointed tool call") {
t.Fatalf("example output %q did not show the initial interrupted run", out)
}
if !strings.Contains(out, "resumed reply: sku-123 is reserved; no duplicate reservation was made") {
t.Fatalf("example output %q did not show the resumed response", out)
}
if !strings.Contains(out, "tool executions: 1") {
t.Fatalf("example output %q did not prove the tool was not replayed", out)
}
}
func captureStdout(t *testing.T, fn func()) string {
t.Helper()
old := os.Stdout
r, w, err := os.Pipe()
if err != nil {
t.Fatalf("pipe stdout: %v", err)
}
os.Stdout = w
var buf bytes.Buffer
done := make(chan struct{})
go func() {
_, _ = io.Copy(&buf, r)
close(done)
}()
fn()
_ = w.Close()
os.Stdout = old
<-done
_ = r.Close()
return buf.String()
}
+212
View File
@@ -0,0 +1,212 @@
package flow
import (
"context"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"go-micro.dev/v6/ai"
)
// AnalyzeOptions configures Analyze.
type AnalyzeOptions struct {
// MaxFeedbackSamples bounds the number of representative grader feedback
// strings retained per candidate. Values <= 0 use a small default.
MaxFeedbackSamples int
}
// AnalyzeOption configures Analyze.
type AnalyzeOption func(*AnalyzeOptions)
// AnalyzeMaxFeedbackSamples sets how many grader feedback examples are kept for
// each candidate in the report.
func AnalyzeMaxFeedbackSamples(n int) AnalyzeOption {
return func(o *AnalyzeOptions) { o.MaxFeedbackSamples = n }
}
// Report is the machine-readable output of Analyze. Candidates are ordered from
// worst to best so an agent, CLI, or human can pick the first improvement to try.
type Report struct {
Candidates []Candidate `json:"candidates"`
}
// Candidate identifies one underperforming flow step and the trace evidence that
// made it worth improving.
type Candidate struct {
Step string `json:"step"`
Metric string `json:"metric"`
Score float64 `json:"score"`
Runs int `json:"runs"`
Failures int `json:"failures"`
PassRate float64 `json:"pass_rate"`
ErrorRate float64 `json:"error_rate"`
AverageRetries float64 `json:"average_retries"`
P50Latency time.Duration `json:"p50_latency"`
P95Latency time.Duration `json:"p95_latency"`
SampleFeedback []string `json:"sample_feedback,omitempty"`
RunIDs []string `json:"run_ids,omitempty"`
}
// Analyze aggregates a bounded window of persisted flow runs and returns ranked
// hill-climbing candidates. It uses the same Run records read by Checkpoint.List:
// failed verification fields in step results drive pass-rate and feedback, step
// status drives error rate, and retry attempts contribute retry pressure. An
// empty window returns an empty report.
func Analyze(runs []Run, opts ...AnalyzeOption) Report {
o := AnalyzeOptions{MaxFeedbackSamples: 3}
for _, opt := range opts {
opt(&o)
}
if o.MaxFeedbackSamples <= 0 {
o.MaxFeedbackSamples = 3
}
stats := map[string]*stepStats{}
for _, run := range runs {
for _, step := range run.Steps {
if step.Name == "" {
continue
}
s := stats[step.Name]
if s == nil {
s = &stepStats{}
stats[step.Name] = s
}
s.runs++
s.runIDs = appendUnique(s.runIDs, run.ID)
if step.Attempts > 1 {
s.retries += step.Attempts - 1
}
if step.Status == "failed" || step.Error != "" {
s.errors++
}
if len(run.Steps) > 0 && !run.Started.IsZero() && !run.Updated.IsZero() {
s.latencies = append(s.latencies, run.Updated.Sub(run.Started)/time.Duration(len(run.Steps)))
}
passed, feedback, ok := verificationFields(step.Result)
if ok {
s.graded++
if !passed {
s.gradeFailures++
if feedback != "" && len(s.feedback) < o.MaxFeedbackSamples {
s.feedback = append(s.feedback, feedback)
}
}
}
}
}
report := Report{}
for step, s := range stats {
if s.runs == 0 {
continue
}
failures := s.errors + s.gradeFailures
passRate := 1.0
if s.graded > 0 {
passRate = float64(s.graded-s.gradeFailures) / float64(s.graded)
} else if s.errors > 0 {
passRate = float64(s.runs-s.errors) / float64(s.runs)
}
errorRate := float64(s.errors) / float64(s.runs)
avgRetries := float64(s.retries) / float64(s.runs)
score := float64(s.gradeFailures)*3 + float64(s.errors)*2 + avgRetries
metric := "pass_rate"
if s.gradeFailures == 0 && s.errors > 0 {
metric = "error_rate"
} else if s.gradeFailures == 0 && s.errors == 0 && s.retries > 0 {
metric = "retry_count"
}
report.Candidates = append(report.Candidates, Candidate{
Step: step, Metric: metric, Score: score, Runs: s.runs, Failures: failures,
PassRate: passRate, ErrorRate: errorRate, AverageRetries: avgRetries,
P50Latency: percentile(s.latencies, 0.50), P95Latency: percentile(s.latencies, 0.95),
SampleFeedback: append([]string(nil), s.feedback...), RunIDs: append([]string(nil), s.runIDs...),
})
}
sort.SliceStable(report.Candidates, func(i, j int) bool {
a, b := report.Candidates[i], report.Candidates[j]
if a.Score == b.Score {
return a.Step < b.Step
}
return a.Score > b.Score
})
return report
}
type stepStats struct {
runs, graded, gradeFailures, errors, retries int
feedback, runIDs []string
latencies []time.Duration
}
// PromptOptimizer proposes prompt improvements for a candidate without mutating
// the source flow. Applying the returned prompt stays explicitly gated by the caller.
type PromptOptimizer struct{ model ai.Model }
// LLMOptimizer returns an optimizer that asks model to revise prompts for
// Analyze candidates. The model is injected so tests and callers can use mocks.
func LLMOptimizer(model ai.Model) *PromptOptimizer { return &PromptOptimizer{model: model} }
// OptimizePrompt asks the model for a revised prompt for candidate using the
// current prompt and trace feedback. It returns only the proposal; it never
// modifies a Flow, Step, or Checkpoint.
func (o *PromptOptimizer) OptimizePrompt(ctx context.Context, candidate Candidate, currentPrompt string) (string, error) {
if o == nil || o.model == nil {
return "", fmt.Errorf("flow: LLMOptimizer requires a model")
}
prompt := fmt.Sprintf("Revise this workflow step prompt to improve the failing step.\nStep: %s\nMetric: %s\nScore: %.2f\nFeedback:\n- %s\n\nCurrent prompt:\n%s\n\nReturn only the revised prompt.", candidate.Step, candidate.Metric, candidate.Score, strings.Join(candidate.SampleFeedback, "\n- "), currentPrompt)
resp, err := o.model.Generate(ctx, &ai.Request{Prompt: prompt})
if err != nil {
return "", err
}
proposal := strings.TrimSpace(resp.Answer)
if proposal == "" {
proposal = strings.TrimSpace(resp.Reply)
}
if proposal == "" {
return "", fmt.Errorf("flow: LLMOptimizer returned an empty prompt")
}
return proposal, nil
}
func verificationFields(result string) (bool, string, bool) {
if result == "" {
return false, "", false
}
var obj map[string]any
if err := json.Unmarshal([]byte(result), &obj); err != nil {
return false, "", false
}
v, ok := obj["verification_passed"].(bool)
if !ok {
return false, "", false
}
fb, _ := obj["verification_feedback"].(string)
return v, fb, true
}
func appendUnique(values []string, value string) []string {
if value == "" {
return values
}
for _, v := range values {
if v == value {
return values
}
}
return append(values, value)
}
func percentile(values []time.Duration, p float64) time.Duration {
if len(values) == 0 {
return 0
}
sorted := append([]time.Duration(nil), values...)
sort.Slice(sorted, func(i, j int) bool { return sorted[i] < sorted[j] })
idx := int(float64(len(sorted)-1) * p)
return sorted[idx]
}
+89
View File
@@ -0,0 +1,89 @@
package flow
import (
"context"
"strings"
"testing"
"time"
"go-micro.dev/v6/ai"
)
func TestAnalyzeRanksFailedGraderStepAbovePassingStep(t *testing.T) {
now := time.Now()
runs := []Run{
{ID: "run-1", Started: now, Updated: now.Add(time.Second), Steps: []StepRecord{
{Name: "draft", Status: "done", Attempts: 2, Result: `{"verification_passed":false,"verification_feedback":"cite sources"}`},
{Name: "publish", Status: "done", Attempts: 1, Result: `{"verification_passed":true,"verification_feedback":"ok"}`},
}},
{ID: "run-2", Started: now, Updated: now.Add(2 * time.Second), Steps: []StepRecord{
{Name: "draft", Status: "done", Attempts: 1, Result: `{"verification_passed":false,"verification_feedback":"too vague"}`},
{Name: "publish", Status: "done", Attempts: 1, Result: `{"verification_passed":true,"verification_feedback":"ok"}`},
}},
}
report := Analyze(runs)
if len(report.Candidates) != 2 {
t.Fatalf("Analyze returned %d candidates, want 2", len(report.Candidates))
}
if got := report.Candidates[0].Step; got != "draft" {
t.Fatalf("top candidate = %q, want draft", got)
}
if report.Candidates[0].PassRate != 0 {
t.Fatalf("draft pass rate = %v, want 0", report.Candidates[0].PassRate)
}
}
func TestAnalyzeCarriesFeedbackSamplesAndRunIDs(t *testing.T) {
report := Analyze([]Run{{ID: "run-9", Steps: []StepRecord{{
Name: "grade", Status: "done", Attempts: 3,
Result: `{"verification_passed":false,"verification_feedback":"include totals"}`,
}}}})
if len(report.Candidates) != 1 {
t.Fatalf("candidates = %d, want 1", len(report.Candidates))
}
c := report.Candidates[0]
if len(c.SampleFeedback) != 1 || c.SampleFeedback[0] != "include totals" {
t.Fatalf("feedback = %#v, want include totals", c.SampleFeedback)
}
if len(c.RunIDs) != 1 || c.RunIDs[0] != "run-9" {
t.Fatalf("run ids = %#v, want run-9", c.RunIDs)
}
if c.AverageRetries != 2 {
t.Fatalf("average retries = %v, want 2", c.AverageRetries)
}
}
func TestAnalyzeEmptyWindowReturnsEmptyReport(t *testing.T) {
if got := Analyze(nil); len(got.Candidates) != 0 {
t.Fatalf("empty Analyze candidates = %d, want 0", len(got.Candidates))
}
}
func TestLLMOptimizerReturnsProposalWithoutMutatingFlow(t *testing.T) {
f := New("optimize", Prompt("original prompt"))
before := f.opts.Prompt
optimizer := LLMOptimizer(&optimizerModel{reply: "revised prompt"})
proposal, err := optimizer.OptimizePrompt(context.Background(), Candidate{Step: "draft", Metric: "pass_rate", SampleFeedback: []string{"cite sources"}}, before)
if err != nil {
t.Fatalf("OptimizePrompt returned error: %v", err)
}
if !strings.Contains(proposal, "revised") {
t.Fatalf("proposal = %q, want revised prompt", proposal)
}
if f.opts.Prompt != before {
t.Fatalf("flow prompt mutated to %q, want %q", f.opts.Prompt, before)
}
}
type optimizerModel struct{ reply string }
func (m *optimizerModel) Init(...ai.Option) error { return nil }
func (m *optimizerModel) Options() ai.Options { return ai.Options{} }
func (m *optimizerModel) Generate(context.Context, *ai.Request, ...ai.GenerateOption) (*ai.Response, error) {
return &ai.Response{Reply: m.reply}, nil
}
func (m *optimizerModel) Stream(context.Context, *ai.Request, ...ai.GenerateOption) (ai.Stream, error) {
return nil, ai.ErrStreamingUnsupported
}
func (m *optimizerModel) String() string { return "optimizer" }
+63
View File
@@ -7,6 +7,7 @@ import (
"go-micro.dev/v6/client"
codecbytes "go-micro.dev/v6/codec/bytes"
"go-micro.dev/v6/store"
)
// fakeClient embeds the default client (so NewRequest works) and
@@ -65,3 +66,65 @@ func TestExecuteDispatchesToAgent(t *testing.T) {
t.Errorf("rendered prompt = %q, want %q", results[0].Prompt, "welcome bob")
}
}
// A caller-owned schedule can trigger an agent workflow without a human chat
// prompt and still leave the normal flow run metadata behind for inspection.
func TestScheduledAgentRunHarnessContract(t *testing.T) {
ctx := context.Background()
cp := StoreCheckpoint(store.NewMemoryStore(), "scheduled-contract")
f := New("scheduled-contract",
Trigger("schedule.daily"),
WithCheckpoint(cp),
Steps(Step{Name: "summarize", Run: Dispatch("ops-agent")}),
)
var parentID string
f.client = &fakeClient{
Client: client.DefaultClient,
callFn: func(req client.Request, rsp interface{}) error {
if req.Service() != "ops-agent" || req.Endpoint() != "Agent.Chat" {
t.Fatalf("dispatched to %s.%s, want ops-agent.Agent.Chat", req.Service(), req.Endpoint())
}
reqFrame := req.Body().(*codecbytes.Frame)
var body map[string]string
if err := json.Unmarshal(reqFrame.Data, &body); err != nil {
t.Fatalf("request body: %v", err)
}
parentID = body["parent_id"]
if body["message"] != "run unattended daily ops review" {
t.Fatalf("message = %q, want scheduled payload", body["message"])
}
frame := rsp.(*codecbytes.Frame)
frame.Data = []byte(`{"reply":"review queued","agent":"ops-agent","parent_id":"` + parentID + `"}`)
return nil
},
}
if err := Scheduled(f, "run unattended daily ops review").Tick(ctx); err != nil {
t.Fatalf("scheduled tick: %v", err)
}
if parentID == "" {
t.Fatal("dispatch did not receive the scheduled flow run id as parent_id")
}
runs, err := cp.List(ctx)
if err != nil {
t.Fatalf("list scheduled runs: %v", err)
}
if len(runs) != 1 {
t.Fatalf("got %d runs, want 1", len(runs))
}
run := runs[0]
if run.ID != parentID {
t.Fatalf("run ID = %q, parent_id = %q", run.ID, parentID)
}
if run.Flow != "scheduled-contract" || run.Status != "done" {
t.Fatalf("run = %+v, want scheduled-contract done", run)
}
if got := run.State.String(); got != "review queued" {
t.Fatalf("run result = %q, want agent reply", got)
}
if len(run.Steps) != 1 || run.Steps[0].Name != "summarize" || run.Steps[0].Status != "done" {
t.Fatalf("steps = %+v, want summarize done", run.Steps)
}
}
+45
View File
@@ -0,0 +1,45 @@
package flow_test
import (
"context"
"fmt"
"strings"
"go-micro.dev/v6/flow"
)
func ExampleVerify() {
generate := func(_ context.Context, in flow.State) (flow.State, error) {
if strings.Contains(in.String(), "feedback") {
in.Data = []byte(`{"answer":"include a source"}`)
return in, nil
}
in.Data = []byte(`{"answer":"draft"}`)
return in, nil
}
grader := func(_ context.Context, out flow.State) (bool, string, error) {
return strings.Contains(out.String(), "source"), "add a source", nil
}
out, _ := flow.Verify(generate, grader, flow.VerifyMaxAttempts(2))(context.Background(), flow.State{})
fmt.Println(strings.Contains(out.String(), `"verification_passed":true`))
// Output: true
}
func ExampleAnalyze() {
runs := []flow.Run{{
ID: "run-1",
Steps: []flow.StepRecord{{
Name: "draft",
Status: "done",
Result: `{"verification_passed":false,"verification_feedback":"add a source"}`,
}},
}}
report := flow.Analyze(runs)
fmt.Println(report.Candidates[0].Step)
fmt.Println(report.Candidates[0].SampleFeedback[0])
// Output:
// draft
// add a source
}
+10 -2
View File
@@ -76,6 +76,7 @@ type Result struct {
Answer string `json:"answer,omitempty"`
ToolCalls []string `json:"tool_calls,omitempty"`
Error string `json:"error,omitempty"`
ErrorKind string `json:"error_kind,omitempty"`
Timestamp time.Time `json:"timestamp"`
Duration float64 `json:"duration_seconds"`
}
@@ -141,7 +142,8 @@ func (f *Flow) Register(reg registry.Registry, br broker.Broker, cl client.Clien
if f.opts.TriggerTopic != "" {
sub, err := br.Subscribe(f.opts.TriggerTopic, func(p broker.Event) error {
data := string(p.Message().Body)
if err := f.Execute(context.Background(), data); err != nil {
ctx := ai.WithRunInfo(context.Background(), ai.RunInfo{Dispatch: "broker", Trigger: f.opts.TriggerTopic})
if err := f.Execute(ctx, data); err != nil {
f.log.Logf(logger.ErrorLevel, "Flow %s failed: %v", f.name, err)
}
return nil
@@ -222,7 +224,10 @@ func (f *Flow) Execute(ctx context.Context, data string) error {
}
runID := uuid.New().String()
ctx = ai.WithRunInfo(ctx, ai.RunInfo{RunID: runID, Flow: f.name})
info, _ := ai.RunInfoFrom(ctx)
info.RunID = runID
info.Flow = f.name
ctx = ai.WithRunInfo(ctx, info)
start := time.Now()
@@ -246,6 +251,7 @@ func (f *Flow) Execute(ctx context.Context, data string) error {
result.Duration = time.Since(start).Seconds()
if err != nil {
result.Error = err.Error()
result.ErrorKind = string(ai.ClassifyError(err))
f.record(result)
return err
}
@@ -261,6 +267,7 @@ func (f *Flow) Execute(ctx context.Context, data string) error {
if err != nil {
result.Duration = time.Since(start).Seconds()
result.Error = err.Error()
result.ErrorKind = string(ai.ClassifyError(err))
f.record(result)
return fmt.Errorf("discover tools: %w", err)
}
@@ -274,6 +281,7 @@ func (f *Flow) Execute(ctx context.Context, data string) error {
if err != nil {
result.Error = err.Error()
result.ErrorKind = string(ai.ClassifyError(err))
f.record(result)
return err
}
+45 -14
View File
@@ -16,13 +16,18 @@ const (
spanNameFlowRun = "flow.run"
spanNameFlowStep = "flow.step"
AttrFlowRunID = "flow.run.id"
AttrFlowParentID = "flow.run.parent_id"
AttrFlowName = "flow.name"
AttrFlowStepName = "flow.step.name"
AttrFlowStatus = "flow.status"
AttrFlowAttempts = "flow.step.attempts"
AttrFlowLatencyMS = "flow.latency_ms"
AttrFlowRunID = "flow.run.id"
AttrFlowParentID = "flow.run.parent_id"
AttrFlowName = "flow.name"
AttrFlowStepName = "flow.step.name"
AttrFlowStatus = "flow.status"
AttrFlowAttempts = "flow.step.attempts"
AttrFlowLatencyMS = "flow.latency_ms"
AttrFlowErrorKind = "flow.error.kind"
AttrFlowVerificationStatus = "flow.verification.status"
AttrFlowVerificationNote = "flow.verification.note"
AttrFlowDispatch = "flow.dispatch"
AttrFlowTrigger = "flow.trigger"
)
func (f *Flow) tracer() trace.Tracer {
@@ -33,12 +38,15 @@ func (f *Flow) startRunSpan(ctx context.Context, run Run) (context.Context, func
if f.opts.TraceProvider == nil {
return ctx, func(Run, error) {}
}
ctx, span := f.tracer().Start(ctx, spanNameFlowRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(
info, _ := ai.RunInfoFrom(ctx)
attrs := []attribute.KeyValue{
attribute.String(AttrFlowRunID, run.ID),
attribute.String(AttrFlowParentID, run.ParentID),
attribute.String(AttrFlowName, f.name),
attribute.String(AttrFlowStatus, run.Status),
))
}
attrs = appendRunInfoDispatch(attrs, info)
ctx, span := f.tracer().Start(ctx, spanNameFlowRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(attrs...))
start := time.Now()
return ctx, func(done Run, err error) {
span.SetAttributes(
@@ -47,6 +55,7 @@ func (f *Flow) startRunSpan(ctx context.Context, run Run) (context.Context, func
)
if err != nil {
span.RecordError(err)
span.SetAttributes(attribute.String(AttrFlowErrorKind, string(ai.ClassifyError(err))))
span.SetStatus(codes.Error, err.Error())
} else {
span.SetStatus(codes.Ok, "")
@@ -55,29 +64,51 @@ func (f *Flow) startRunSpan(ctx context.Context, run Run) (context.Context, func
}
}
func (f *Flow) runStepSpan(ctx context.Context, step Step, in State) (State, int, error) {
func (f *Flow) runStepSpan(ctx context.Context, step Step, in State) (State, int, Verification, error) {
if f.opts.TraceProvider == nil {
return f.runStep(ctx, step, in)
}
info, _ := ai.RunInfoFrom(ctx)
ctx, span := f.tracer().Start(ctx, spanNameFlowStep, trace.WithAttributes(
attrs := []attribute.KeyValue{
attribute.String(AttrFlowRunID, info.RunID),
attribute.String(AttrFlowParentID, info.ParentID),
attribute.String(AttrFlowName, f.name),
attribute.String(AttrFlowStepName, step.Name),
))
}
attrs = appendRunInfoDispatch(attrs, info)
ctx, span := f.tracer().Start(ctx, spanNameFlowStep, trace.WithAttributes(attrs...))
start := time.Now()
out, attempts, err := f.runStep(ctx, step, in)
out, attempts, verification, err := f.runStep(ctx, step, in)
span.SetAttributes(
attribute.Int(AttrFlowAttempts, attempts),
attribute.Int64(AttrFlowLatencyMS, time.Since(start).Milliseconds()),
)
if verification.Passed {
span.SetAttributes(attribute.String(AttrFlowVerificationStatus, "passed"))
}
if verification.Feedback != "" {
span.SetAttributes(attribute.String(AttrFlowVerificationNote, verification.Feedback))
if !verification.Passed {
span.SetAttributes(attribute.String(AttrFlowVerificationStatus, "failed"))
}
}
if err != nil {
span.RecordError(err)
span.SetAttributes(attribute.String(AttrFlowErrorKind, string(ai.ClassifyError(err))))
span.SetStatus(codes.Error, err.Error())
} else {
span.SetStatus(codes.Ok, "")
}
span.End()
return out, attempts, err
return out, attempts, verification, err
}
func appendRunInfoDispatch(attrs []attribute.KeyValue, info ai.RunInfo) []attribute.KeyValue {
if info.Dispatch != "" {
attrs = append(attrs, attribute.String(AttrFlowDispatch, info.Dispatch))
}
if info.Trigger != "" {
attrs = append(attrs, attribute.String(AttrFlowTrigger, info.Trigger))
}
return attrs
}
+26
View File
@@ -74,3 +74,29 @@ func flowSpanAttributes(attrs []attribute.KeyValue) map[string]string {
func withTestRunInfo(ctx context.Context, runID string) context.Context {
return ai.WithRunInfo(ctx, ai.RunInfo{RunID: runID, Agent: "planner"})
}
func TestScheduledFlowOpenTelemetryDispatchAttributes(t *testing.T) {
exp := tracetest.NewInMemoryExporter()
tp := trace.NewTracerProvider(trace.WithSyncer(exp))
step := Step{Name: "summarize", Run: func(ctx context.Context, in State) (State, error) {
in.Data = []byte("queued")
return in, nil
}}
f := New("scheduled-observed", Trigger("schedule.daily"), WithCheckpoint(StoreCheckpoint(store.NewMemoryStore(), "scheduled-observed")), TraceProvider(tp), Steps(step))
if err := Scheduled(f, "daily ops review").Tick(context.Background()); err != nil {
t.Fatal(err)
}
for _, span := range exp.GetSpans().Snapshots() {
if span.Name() != spanNameFlowRun {
continue
}
attrs := flowSpanAttributes(span.Attributes())
if attrs[AttrFlowDispatch] != "schedule" || attrs[AttrFlowTrigger] != "schedule.daily" {
t.Fatalf("scheduled run span dispatch attributes = %#v", attrs)
}
return
}
t.Fatal("flow run span not emitted")
}
+60
View File
@@ -0,0 +1,60 @@
package flow
import (
"context"
"time"
"go-micro.dev/v6/ai"
)
// Schedule binds a flow to a recurring work item without introducing a
// scheduler service. It is a small harness contract: callers own the clock,
// Go Micro owns turning each tick into the same inspectable flow run used for
// broker events and direct Execute calls.
type Schedule struct {
flow *Flow
data string
}
// Scheduled returns a deterministic scheduled-run harness for this flow.
// Tests and event loops can call Tick directly; production processes can wire
// the same contract to time.Ticker through RunEvery. Each tick calls Execute, so
// checkpointed run history, parent/run metadata, cancellation, and inspection
// stay on the normal flow surfaces.
func Scheduled(f *Flow, data string) Schedule {
return Schedule{flow: f, data: data}
}
// Tick starts one scheduled run immediately and returns when that run finishes.
func (s Schedule) Tick(ctx context.Context) error {
if ctx == nil {
ctx = context.Background()
}
info, _ := ai.RunInfoFrom(ctx)
info.Dispatch = "schedule"
if info.Trigger == "" {
info.Trigger = s.flow.opts.TriggerTopic
}
if info.Trigger == "" {
info.Trigger = "schedule"
}
return s.flow.Execute(ai.WithRunInfo(ctx, info), s.data)
}
// RunEvery drives scheduled runs from a ticker until ctx is canceled. It does
// not persist schedule definitions or host a scheduler; it only adapts a caller
// owned cadence to Tick.
func (s Schedule) RunEvery(ctx context.Context, interval time.Duration) error {
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return ctx.Err()
case <-ticker.C:
if err := s.Tick(ctx); err != nil {
return err
}
}
}
}
+85 -21
View File
@@ -52,22 +52,53 @@ func (s State) String() string { return string(s.Data) }
// returns the next state.
type StepFunc func(ctx context.Context, in State) (State, error)
// Step is one unit of a flow — a named action with an optional retry
// override. There is one Step kind; the action is the Run func, and the
// Call/LLM/Agent helpers produce the common ones.
// Verifier grades a step output before the flow advances. Returning
// Passed=false converts the grade into a retryable VerificationError, so
// the existing step retry/supervision path can feed Feedback into the next
// attempt through ai.RunInfo.VerificationFeedback.
type Verifier func(ctx context.Context, out State) (Verification, error)
// Verification is the verifier's deterministic grade for one step attempt.
type Verification struct {
Passed bool
Feedback string
}
// VerificationError reports a failed grade. It is returned from runStep so
// existing retry, checkpoint, and trace paths handle verifier failures the
// same way they handle step execution failures.
type VerificationError struct {
Step string
Feedback string
}
func (e *VerificationError) Error() string {
if e.Feedback == "" {
return fmt.Sprintf("flow: verification failed for step %q", e.Step)
}
return fmt.Sprintf("flow: verification failed for step %q: %s", e.Step, e.Feedback)
}
// Step is one unit of a flow — a named action with optional retry and
// verification hooks. There is one Step kind; the action is the Run func,
// and the Call/LLM/Agent helpers produce the common ones.
type Step struct {
Name string
Run StepFunc
Retry int // per-step override of the flow's retry (0 = use the flow default)
Name string
Run StepFunc
Retry int // per-step override of the flow's retry (0 = use the flow default)
Verify Verifier // optional grade; failed grades retry the step with feedback in RunInfo
}
// StepRecord is the recorded outcome of one step within a run.
type StepRecord struct {
Name string `json:"name"`
Status string `json:"status"` // pending | in_progress | done | failed
Attempts int `json:"attempts"`
Result string `json:"result,omitempty"`
Error string `json:"error,omitempty"`
Name string `json:"name"`
Status string `json:"status"` // pending | in_progress | done | failed
Attempts int `json:"attempts"`
Result string `json:"result,omitempty"`
Error string `json:"error,omitempty"`
ErrorKind string `json:"error_kind,omitempty"`
VerificationStatus string `json:"verification_status,omitempty"` // passed | failed
VerificationNote string `json:"verification_note,omitempty"`
}
// Run is the persisted record of one flow execution — what a Checkpoint
@@ -402,7 +433,12 @@ func (f *Flow) Pending(ctx context.Context) ([]Run, error) {
func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
steps := f.opts.Steps
ctx = withDeps(ctx, &runDeps{client: f.client, model: f.model, tools: f.toolSet})
ctx = ai.WithRunInfo(ctx, ai.RunInfo{RunID: run.ID, ParentID: run.ParentID, Agent: f.name, Flow: f.name})
info, _ := ai.RunInfoFrom(ctx)
info.RunID = run.ID
info.ParentID = run.ParentID
info.Agent = f.name
info.Flow = f.name
ctx = ai.WithRunInfo(ctx, info)
ctx, finishSpan := f.startRunSpan(ctx, run)
var spanErr error
defer func() { finishSpan(run, spanErr) }()
@@ -425,12 +461,14 @@ func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
return run, err
}
out, attempts, err := f.runStepSpan(ctx, step, run.State)
out, attempts, verification, err := f.runStepSpan(ctx, step, run.State)
run.Steps[i].Attempts = attempts
applyVerificationRecord(&run.Steps[i], verification)
if err != nil {
spanErr = err
run.Steps[i].Status = "failed"
run.Steps[i].Error = err.Error()
run.Steps[i].ErrorKind = string(ai.ClassifyError(err))
run.Status = "failed"
if saveErr := f.save(ctx, run); saveErr != nil {
spanErr = saveErr
@@ -474,40 +512,65 @@ func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
// runStep runs one step, retrying on error up to the resolved retry count.
// A step with no Run function is a configuration error, and a canceled run
// stops retrying immediately rather than burning the rest of its budget.
func (f *Flow) runStep(ctx context.Context, step Step, in State) (State, int, error) {
func (f *Flow) runStep(ctx context.Context, step Step, in State) (State, int, Verification, error) {
if step.Run == nil {
return in, 0, fmt.Errorf("flow: step %q has no Run function", step.Name)
return in, 0, Verification{}, fmt.Errorf("flow: step %q has no Run function", step.Name)
}
retries := f.opts.Retry
if step.Retry > 0 {
retries = step.Retry
}
var lastErr error
var lastVerification Verification
var feedback string
for attempt := 1; attempt <= retries+1; attempt++ {
// Stop the moment the run's context is canceled or its deadline
// passes — a canceled run shouldn't keep retrying, and the context
// error is surfaced so callers can detect cancellation upstream.
if err := ctx.Err(); err != nil {
return in, attempt - 1, err
return in, attempt - 1, lastVerification, err
}
attemptCtx := ctx
if info, ok := ai.RunInfoFrom(ctx); ok {
info.Step = step.Name
ctx = ai.WithRunInfo(ctx, info)
info.VerificationFeedback = feedback
attemptCtx = ai.WithRunInfo(ctx, info)
}
out, err := step.Run(attemptCtx, in)
if err == nil && step.Verify != nil {
lastVerification, err = step.Verify(attemptCtx, out)
if err == nil && !lastVerification.Passed {
err = &VerificationError{Step: step.Name, Feedback: lastVerification.Feedback}
}
}
out, err := step.Run(ctx, in)
if err == nil {
return out, attempt, nil
return out, attempt, lastVerification, nil
}
lastErr = err
if verr, ok := err.(*VerificationError); ok {
feedback = verr.Feedback
}
if attempt <= retries && f.opts.RetryBackoff > 0 {
select {
case <-time.After(f.opts.RetryBackoff):
case <-ctx.Done():
return in, attempt, ctx.Err()
return in, attempt, lastVerification, ctx.Err()
}
}
}
return in, retries + 1, lastErr
return in, retries + 1, lastVerification, lastErr
}
func applyVerificationRecord(record *StepRecord, verification Verification) {
if verification.Passed {
record.VerificationStatus = "passed"
}
if verification.Feedback != "" {
record.VerificationNote = truncate(verification.Feedback, 200)
if !verification.Passed {
record.VerificationStatus = "failed"
}
}
}
func (f *Flow) save(ctx context.Context, run Run) error {
@@ -555,6 +618,7 @@ func resultFromRun(trigger string, run Run) Result {
r.ToolCalls = append(r.ToolCalls, s.Name+":"+s.Status)
if s.Error != "" {
r.Error = s.Error
r.ErrorKind = s.ErrorKind
}
}
if run.Status == "done" {
+34
View File
@@ -548,3 +548,37 @@ func TestStateSetScan(t *testing.T) {
t.Errorf("round-trip failed: %+v", got)
}
}
func TestFlowFailureRecordsErrorKind(t *testing.T) {
cp := StoreCheckpoint(store.NewMemoryStore(), "failure-kind")
f := New("failure-kind",
WithCheckpoint(cp),
Steps(Step{Name: "limited", Run: func(_ context.Context, in State) (State, error) {
return in, errors.New("rate limit exceeded")
}}),
)
err := f.Execute(context.Background(), "payload")
if err == nil {
t.Fatal("Execute error = nil, want failure")
}
runs, listErr := cp.List(context.Background())
if listErr != nil {
t.Fatalf("List: %v", listErr)
}
if len(runs) != 1 {
t.Fatalf("runs = %d, want 1", len(runs))
}
if got := runs[0].Steps[0].ErrorKind; got != string(ai.ErrorKindRateLimited) {
t.Fatalf("step error kind = %q, want %q", got, ai.ErrorKindRateLimited)
}
results := f.Results()
if len(results) != 1 {
t.Fatalf("results = %d, want 1", len(results))
}
if got := results[0].ErrorKind; got != string(ai.ErrorKindRateLimited) {
t.Fatalf("result error kind = %q, want %q", got, ai.ErrorKindRateLimited)
}
}
+100
View File
@@ -0,0 +1,100 @@
package flow
import (
"context"
"errors"
"testing"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/store"
)
func TestFlowStepVerificationRetriesWithFeedback(t *testing.T) {
var attempts int
var feedback []string
step := Step{
Name: "draft",
Retry: 1,
Run: func(ctx context.Context, in State) (State, error) {
attempts++
info, ok := ai.RunInfoFrom(ctx)
if !ok {
t.Fatal("RunInfo missing from verified step")
}
feedback = append(feedback, info.VerificationFeedback)
if info.VerificationFeedback == "add evidence" {
in.Data = []byte("answer with evidence")
} else {
in.Data = []byte("answer")
}
return in, nil
},
Verify: func(ctx context.Context, out State) (Verification, error) {
if out.String() == "answer with evidence" {
return Verification{Passed: true, Feedback: "meets rubric"}, nil
}
return Verification{Feedback: "add evidence"}, nil
},
}
cp := StoreCheckpoint(store.NewMemoryStore(), "verified")
f := New("verified", WithCheckpoint(cp), Steps(step))
if err := f.Execute(context.Background(), "question"); err != nil {
t.Fatal(err)
}
if attempts != 2 {
t.Fatalf("attempts = %d, want 2", attempts)
}
if len(feedback) != 2 || feedback[0] != "" || feedback[1] != "add evidence" {
t.Fatalf("feedback = %#v, want empty then verifier feedback", feedback)
}
runs, err := cp.List(context.Background())
if err != nil {
t.Fatal(err)
}
if len(runs) != 1 {
t.Fatalf("runs = %d, want 1", len(runs))
}
stepRecord := runs[0].Steps[0]
if stepRecord.Status != "done" || stepRecord.Attempts != 2 || stepRecord.VerificationStatus != "passed" || stepRecord.VerificationNote != "meets rubric" {
t.Fatalf("step record = %#v", stepRecord)
}
}
func TestFlowStepVerificationFailureIsCheckpointed(t *testing.T) {
step := Step{
Name: "grade",
Run: func(ctx context.Context, in State) (State, error) {
in.Data = []byte("bad")
return in, nil
},
Verify: func(ctx context.Context, out State) (Verification, error) {
return Verification{Feedback: "missing citation"}, nil
},
}
cp := StoreCheckpoint(store.NewMemoryStore(), "verified-fail")
f := New("verified-fail", WithCheckpoint(cp), Steps(step))
err := f.Execute(context.Background(), "question")
if err == nil {
t.Fatal("Execute succeeded, want verification failure")
}
var verr *VerificationError
if !errors.As(err, &verr) {
t.Fatalf("error = %T %v, want VerificationError", err, err)
}
if verr.Feedback != "missing citation" {
t.Fatalf("feedback = %q, want missing citation", verr.Feedback)
}
runs, listErr := cp.List(context.Background())
if listErr != nil {
t.Fatal(listErr)
}
if len(runs) != 1 {
t.Fatalf("runs = %d, want 1", len(runs))
}
stepRecord := runs[0].Steps[0]
if runs[0].Status != "failed" || stepRecord.VerificationStatus != "failed" || stepRecord.VerificationNote != "missing citation" {
t.Fatalf("run = %#v step = %#v", runs[0], stepRecord)
}
}
+191
View File
@@ -0,0 +1,191 @@
package flow
import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
"go-micro.dev/v6/ai"
)
// Grader checks a step output against a rubric. It returns pass=true when the
// output is acceptable; otherwise feedback should explain what the next attempt
// should fix.
type Grader func(ctx context.Context, out State) (pass bool, feedback string, err error)
// VerifyOptions configure Verify.
type VerifyOptions struct {
// MaxAttempts bounds how many times the body can run. Default 2.
MaxAttempts int
// Backoff waits between failed grades. Zero means retry immediately.
Backoff time.Duration
// FeedbackField is the JSON field used to thread grader feedback into the
// next attempt's input. Default "feedback".
FeedbackField string
}
// VerifyOption configures Verify.
type VerifyOption func(*VerifyOptions)
// VerifyMaxAttempts sets the total attempt budget for Verify. Values <= 0 use
// the default of 2.
func VerifyMaxAttempts(n int) VerifyOption { return func(o *VerifyOptions) { o.MaxAttempts = n } }
// VerifyBackoff sets the delay between failed verification attempts.
func VerifyBackoff(d time.Duration) VerifyOption { return func(o *VerifyOptions) { o.Backoff = d } }
// VerifyFeedbackField sets the JSON field used to pass grader feedback to the
// next body attempt. Empty values use "feedback".
func VerifyFeedbackField(field string) VerifyOption {
return func(o *VerifyOptions) { o.FeedbackField = field }
}
// Verify runs body, grades its output, and retries with grader feedback threaded
// into the next input until the grader passes or MaxAttempts is exhausted. It is
// a StepFunc, so it composes directly as Step.Run with Loop, LLM, Call, Agent, or
// any code-defined step.
//
// On a failed grade, Verify adds the feedback to the next attempt's input as a
// JSON field named "feedback" (or VerifyFeedbackField). When all attempts fail,
// it returns the last output without error, annotated with verification fields so
// the run can keep the bounded failure outcome in its state:
// "verification_passed": false, "verification_feedback", and
// "verification_attempts".
func Verify(body StepFunc, grader Grader, opts ...VerifyOption) StepFunc {
o := VerifyOptions{MaxAttempts: 2, FeedbackField: "feedback"}
for _, op := range opts {
op(&o)
}
if o.MaxAttempts <= 0 {
o.MaxAttempts = 2
}
if o.FeedbackField == "" {
o.FeedbackField = "feedback"
}
return func(ctx context.Context, in State) (State, error) {
if body == nil {
return in, fmt.Errorf("flow: Verify requires a body step")
}
if grader == nil {
return in, fmt.Errorf("flow: Verify requires a grader")
}
cur := in
last := in
feedback := ""
for attempt := 1; attempt <= o.MaxAttempts; attempt++ {
if err := ctx.Err(); err != nil {
return last, err
}
if feedback != "" {
var err error
cur, err = stateWithField(cur, o.FeedbackField, feedback)
if err != nil {
return last, err
}
}
out, err := body(ctx, cur)
if err != nil {
return last, fmt.Errorf("verify attempt %d: %w", attempt, err)
}
last = out
pass, fb, err := grader(ctx, out)
if err != nil {
return last, fmt.Errorf("verify grade attempt %d: %w", attempt, err)
}
if pass {
return stateWithVerification(out, true, fb, attempt)
}
feedback = fb
cur = in
if attempt < o.MaxAttempts && o.Backoff > 0 {
select {
case <-time.After(o.Backoff):
case <-ctx.Done():
return last, ctx.Err()
}
}
}
return stateWithVerification(last, false, feedback, o.MaxAttempts)
}
}
// LLMGrader returns a grader that asks the flow model to judge the latest output
// against rubric. The model should answer with pass/fail plus short feedback.
// It reuses the flow's configured model, so it must run inside a flow.
func LLMGrader(rubric string) Grader {
return func(ctx context.Context, out State) (bool, string, error) {
d := depsFrom(ctx)
if d == nil || d.model == nil {
return false, "", fmt.Errorf("flow: LLMGrader requires a flow model (set Provider/APIKey)")
}
prompt := fmt.Sprintf("Grade the latest result against this rubric:\n%s\n\nLatest result:\n%s\n\nAnswer with PASS or FAIL on the first line, followed by one short feedback sentence.", rubric, out.String())
resp, err := d.model.Generate(ctx, &ai.Request{Prompt: prompt})
if err != nil {
return false, "", err
}
reply := resp.Answer
if reply == "" {
reply = resp.Reply
}
return parseGrade(reply)
}
}
func parseGrade(reply string) (bool, string, error) {
text := strings.TrimSpace(reply)
if text == "" {
return false, "", fmt.Errorf("flow: LLMGrader returned an empty grade")
}
lines := strings.SplitN(text, "\n", 2)
first := strings.ToLower(strings.TrimSpace(lines[0]))
feedback := ""
if len(lines) > 1 {
feedback = strings.TrimSpace(lines[1])
}
pass := strings.HasPrefix(first, "pass") || isAffirmative(first)
if !pass && feedback == "" {
feedback = text
}
return pass, feedback, nil
}
func stateWithField(s State, field, value string) (State, error) {
var obj map[string]any
if len(s.Data) > 0 && json.Unmarshal(s.Data, &obj) == nil && obj != nil {
obj[field] = value
return stateWithObject(s, obj)
}
obj = map[string]any{field: value}
if len(s.Data) > 0 {
obj["data"] = s.String()
}
return stateWithObject(s, obj)
}
func stateWithVerification(s State, passed bool, feedback string, attempts int) (State, error) {
var obj map[string]any
if len(s.Data) > 0 && json.Unmarshal(s.Data, &obj) == nil && obj != nil {
obj["verification_passed"] = passed
obj["verification_feedback"] = feedback
obj["verification_attempts"] = attempts
return stateWithObject(s, obj)
}
obj = map[string]any{
"data": s.String(),
"verification_passed": passed,
"verification_feedback": feedback,
"verification_attempts": attempts,
}
return stateWithObject(s, obj)
}
func stateWithObject(s State, obj map[string]any) (State, error) {
b, err := json.Marshal(obj)
if err != nil {
return s, err
}
s.Data = b
return s, nil
}
+96
View File
@@ -0,0 +1,96 @@
package flow
import (
"context"
"strings"
"testing"
)
func TestVerifyPassesFirstTry(t *testing.T) {
attempts := 0
step := Verify(func(_ context.Context, in State) (State, error) {
attempts++
in.Data = []byte(`{"answer":"ok"}`)
return in, nil
}, func(context.Context, State) (bool, string, error) {
return true, "looks good", nil
}, VerifyMaxAttempts(3))
out, err := step(context.Background(), State{})
if err != nil {
t.Fatalf("Verify returned error: %v", err)
}
if attempts != 1 {
t.Fatalf("body attempts = %d, want 1", attempts)
}
var got map[string]any
if err := out.Scan(&got); err != nil {
t.Fatalf("scan output: %v", err)
}
if got["verification_passed"] != true {
t.Fatalf("verification_passed = %v, want true", got["verification_passed"])
}
}
func TestVerifyRetriesWithFeedback(t *testing.T) {
attempts := 0
var secondInput map[string]string
step := Verify(func(_ context.Context, in State) (State, error) {
attempts++
if attempts == 2 {
if err := in.Scan(&secondInput); err != nil {
t.Fatalf("scan second input: %v", err)
}
}
in.Data = []byte(`{"answer":"draft"}`)
return in, nil
}, func(_ context.Context, _ State) (bool, string, error) {
return attempts >= 2, "include citations", nil
}, VerifyMaxAttempts(3))
out, err := step(context.Background(), State{Data: []byte(`{"topic":"agents"}`)})
if err != nil {
t.Fatalf("Verify returned error: %v", err)
}
if attempts != 2 {
t.Fatalf("body attempts = %d, want 2", attempts)
}
if secondInput["feedback"] != "include citations" {
t.Fatalf("feedback = %q, want include citations", secondInput["feedback"])
}
if !strings.Contains(out.String(), `"verification_passed":true`) {
t.Fatalf("output missing successful verification annotation: %s", out.String())
}
}
func TestVerifyExhaustsAttemptsReturnsLastOutput(t *testing.T) {
attempts := 0
step := Verify(func(_ context.Context, in State) (State, error) {
attempts++
in.Data = []byte(`{"answer":"still wrong"}`)
return in, nil
}, func(context.Context, State) (bool, string, error) {
return false, "try again", nil
}, VerifyMaxAttempts(2))
out, err := step(context.Background(), State{})
if err != nil {
t.Fatalf("Verify returned error: %v", err)
}
if attempts != 2 {
t.Fatalf("body attempts = %d, want 2", attempts)
}
var got map[string]any
if err := out.Scan(&got); err != nil {
t.Fatalf("scan output: %v", err)
}
if got["verification_passed"] != false {
t.Fatalf("verification_passed = %v, want false", got["verification_passed"])
}
if got["verification_feedback"] != "try again" {
t.Fatalf("verification_feedback = %v, want try again", got["verification_feedback"])
}
if got["verification_attempts"] != float64(2) {
t.Fatalf("verification_attempts = %v, want 2", got["verification_attempts"])
}
}
+3 -1
View File
@@ -180,6 +180,8 @@ type Provider struct {
type Capabilities struct {
Streaming bool `json:"streaming"`
PushNotifications bool `json:"pushNotifications"`
TaskResubscribe bool `json:"taskResubscribe"`
InputRequired bool `json:"inputRequired"`
}
// Skill is a capability advertised on the Agent Card.
@@ -349,7 +351,7 @@ func Card(name, url, description string, services []string) AgentCard {
URL: url,
Version: "1.0.0",
ProtocolVersion: protocolVersion,
Capabilities: Capabilities{Streaming: true, PushNotifications: true},
Capabilities: Capabilities{Streaming: true, PushNotifications: true, TaskResubscribe: true, InputRequired: true},
DefaultInputModes: []string{"text/plain"},
DefaultOutputModes: []string{"text/plain"},
Skills: skills,
+3
View File
@@ -91,6 +91,9 @@ func TestAgentCardFromRegistry(t *testing.T) {
if card.ProtocolVersion == "" {
t.Errorf("card missing protocolVersion: %+v", card)
}
if !card.Capabilities.TaskResubscribe || !card.Capabilities.InputRequired {
t.Errorf("card capabilities = %+v, want task resubscribe and input-required advertised", card.Capabilities)
}
if got := skillIDs(card.Skills); strings.Join(got, ",") != "task,project" {
t.Errorf("skill IDs = %v, want [task project]", got)
}
+74 -1
View File
@@ -1,11 +1,13 @@
package a2a
import (
"bufio"
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"strings"
"time"
"github.com/google/uuid"
@@ -135,6 +137,77 @@ func (c *Client) SendMessage(ctx context.Context, message Message) (*Task, error
return &task, nil
}
// Resubscribe reconnects to a retained or active task stream and returns task
// snapshots as the remote agent emits updates. The returned channel is closed
// when the task reaches a terminal state or ctx is canceled.
func (c *Client) Resubscribe(ctx context.Context, taskID string) (<-chan Task, <-chan error) {
tasks := make(chan Task, 8)
errs := make(chan error, 1)
go func() {
defer close(tasks)
defer close(errs)
body, _ := json.Marshal(map[string]any{
"jsonrpc": "2.0",
"id": uuid.New().String(),
"method": "tasks/resubscribe",
"params": getParams{ID: taskID},
})
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.url, bytes.NewReader(body))
if err != nil {
errs <- err
return
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "text/event-stream")
resp, err := c.http.Do(req)
if err != nil {
errs <- err
return
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
errs <- fmt.Errorf("tasks/resubscribe: status %d", resp.StatusCode)
return
}
scanner := bufio.NewScanner(resp.Body)
for scanner.Scan() {
line := scanner.Text()
if !strings.HasPrefix(line, "data:") {
continue
}
payload := strings.TrimSpace(strings.TrimPrefix(line, "data:"))
if payload == "" {
continue
}
var out struct {
Result Task `json:"result"`
Error *rpcError `json:"error"`
}
if err := json.Unmarshal([]byte(payload), &out); err != nil {
errs <- err
return
}
if out.Error != nil {
errs <- fmt.Errorf("a2a tasks/resubscribe: %s (%d)", out.Error.Message, out.Error.Code)
return
}
select {
case <-ctx.Done():
errs <- ctx.Err()
return
case tasks <- out.Result:
}
if terminal(out.Result.Status.State) {
return
}
}
if err := scanner.Err(); err != nil {
errs <- err
}
}()
return tasks, errs
}
// SetPushNotificationConfig asks the remote agent to POST updates for taskID to cfg.URL.
func (c *Client) SetPushNotificationConfig(ctx context.Context, taskID string, cfg PushNotificationConfig) error {
_, err := c.call(ctx, "tasks/pushNotificationConfig/set", pushConfigParams{
@@ -193,7 +266,7 @@ func (c *Client) call(ctx context.Context, method string, params any) (json.RawM
func terminal(state string) bool {
switch state {
case "completed", "failed", "canceled", "rejected":
case "completed", "failed", "canceled", "rejected", "input-required":
return true
}
return false
+56
View File
@@ -3,8 +3,10 @@ package a2a
import (
"context"
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
)
@@ -110,3 +112,57 @@ func TestClientContinuesTaskAndConfiguresPush(t *testing.T) {
t.Fatal("timed out waiting for push update")
}
}
func TestClientResubscribeStreamsRetainedAndLiveTask(t *testing.T) {
d := newDispatcher()
initial := &Task{ID: "task-1", ContextID: "ctx-1", Kind: "task", Status: TaskStatus{State: stateWorking, Timestamp: time.Now().UTC().Format(time.RFC3339)}}
d.store(initial)
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
d.serve(w, r, func(context.Context, string) (string, error) { return "", nil })
}))
defer ts.Close()
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
tasks, errs := NewClient(ts.URL).Resubscribe(ctx, initial.ID)
first := <-tasks
if first.ID != initial.ID || first.Status.State != stateWorking {
t.Fatalf("first resubscribe task = %+v, want retained working task", first)
}
final := &Task{ID: initial.ID, ContextID: initial.ContextID, Kind: "task", Status: TaskStatus{State: stateCompleted, Timestamp: time.Now().UTC().Format(time.RFC3339)}, Artifacts: []Artifact{textArtifact("done")}}
d.store(final)
second := <-tasks
if second.ID != final.ID || second.Status.State != stateCompleted || textOf(second.Artifacts[0].Parts) != "done" {
t.Fatalf("second resubscribe task = %+v, want live completed task", second)
}
if _, ok := <-tasks; ok {
t.Fatal("resubscribe task channel stayed open after terminal update")
}
select {
case err := <-errs:
if err != nil {
t.Fatalf("resubscribe error = %v", err)
}
default:
}
}
func TestClientSendMessageReturnsInputRequiredTask(t *testing.T) {
card := Card("solo", "http://localhost:4000", "", []string{"task"})
h := NewAgentHandler(card, func(context.Context, string) (string, error) {
return "", errors.New("input-required: provide approval code")
})
ts := httptest.NewServer(h)
defer ts.Close()
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
task, err := NewClient(ts.URL).SendMessage(ctx, Message{Parts: []Part{{Kind: "text", Text: "approve?"}}})
if err != nil {
t.Fatalf("SendMessage: %v", err)
}
if task.Status.State != stateInputRequired || !strings.Contains(textOf(task.Artifacts[0].Parts), "provide approval code") {
t.Fatalf("task = %+v, want input-required handoff", task)
}
}
+12
View File
@@ -29,6 +29,18 @@ consistent across providers.
The companion `TestAgentProviderConformanceFakeError` keeps provider error
propagation covered locally without relying on external credentials.
## Local no-secret conformance
Use `make provider-conformance-mock` to run the same provider conformance harness
through the deterministic mock provider. That target requires no API keys and is
what `make harness` delegates to after the 0→1 and 0→hero scenarios, so every PR
continues to exercise the provider-facing agent/tool contract without spending
live model credits.
Use `make provider-conformance` when you want the live-provider sweep: providers
without keys are skipped, and configured providers must satisfy the same harness
contract.
## Scheduled CI
The daily/manual `Harness (E2E)` workflow runs the same matrix with
+2 -2
View File
@@ -21,8 +21,8 @@ changes, architectural rewrites. Those go to the human.
## Work queue (ranked)
1. **Broaden end-to-end agent streaming coverage** ([#3402](https://github.com/micro/go-micro/issues/3402)) — with durable agent checkpoint/resume guarded by #3406, streaming is now the highest-value Next-phase developer-visible seam: provider tokens need to move consistently through `ai.Stream`, `micro chat`, `Agent.Chat`, and A2A streaming. CI-safe local coverage plus provider-gated conformance should lock down chunk ordering, terminal/error events, and fallback behavior.
2. **Emit OpenTelemetry spans for agent runs** ([#3403](https://github.com/micro/go-micro/issues/3403)) — once long runs can resume and stream, operators need the same run story in traces that developers see via `micro runs`: lifecycle boundaries, model/tool calls, approvals, retries, failures, and cancellation correlated by run ID without leaking sensitive payloads.
1. **CI-verify cross-provider agent conformance** ([#3537](https://github.com/micro/go-micro/issues/3537)) — retry/backoff resilience and durable checkpoint/resume have shipped, so the highest-value remaining Now-phase gap is proving the same services → agents → workflows scenario works across every supported model provider. This keeps the harness honest where provider drift most directly breaks developer trust: model calls, tool calls, skip/pass/fail reporting, and CI-safe execution when credentials are absent.
2. **Emit OpenTelemetry spans for agent run timelines** ([#3525](https://github.com/micro/go-micro/issues/3525)) — recent work made runs inspectable, correlated trace metadata through scheduled dispatch, verified restart resume, and added opt-in tool retries; the next step is to turn that RunInfo foundation into standard OTel spans for agent runs, model calls, tool calls, checkpoint/resume, and failures. This keeps `micro runs` useful while making the harness observable in the systems developers already run.
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
architecture-review pass._
@@ -0,0 +1,184 @@
// A2A stream fallback harness.
//
// It exercises the gateway boundary that fronts an agent over A2A. The agent is
// configured with tools and memory, but its model streaming path deliberately
// reports ai.ErrStreamingUnsupported; the A2A gateway must fall back to the
// normal Ask path and still complete the same tool-calling run.
package main
import (
"bufio"
"bytes"
"context"
"encoding/json"
"errors"
"flag"
"fmt"
"io"
"net/http"
"net/http/httptest"
"os"
"strings"
"time"
"go-micro.dev/v6/agent"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/gateway/a2a"
"go-micro.dev/v6/registry"
"go-micro.dev/v6/store"
)
type mockModel struct{ opts ai.Options }
func newMock(opts ...ai.Option) ai.Model {
m := &mockModel{}
_ = m.Init(opts...)
return m
}
func (m *mockModel) Init(opts ...ai.Option) error {
for _, o := range opts {
o(&m.opts)
}
return nil
}
func (m *mockModel) Options() ai.Options { return m.opts }
func (m *mockModel) String() string { return "mock" }
func (m *mockModel) Stream(context.Context, *ai.Request, ...ai.GenerateOption) (ai.Stream, error) {
return nil, ai.ErrStreamingUnsupported
}
func (m *mockModel) Generate(ctx context.Context, req *ai.Request, _ ...ai.GenerateOption) (*ai.Response, error) {
if req.Prompt == "" {
return nil, errors.New("missing prompt")
}
if len(req.Messages) == 0 || req.Messages[len(req.Messages)-1].Role != "user" {
return nil, fmt.Errorf("missing user history: %+v", req.Messages)
}
if len(req.Tools) == 0 || m.opts.ToolHandler == nil {
return nil, errors.New("missing tools or tool handler")
}
res := m.opts.ToolHandler(ctx, ai.ToolCall{ID: "a2a-fallback-call", Name: "fallback_echo", Input: map[string]any{"value": "a2a-fallback"}})
if res.Content == "" {
return nil, errors.New("empty tool result")
}
return &ai.Response{Reply: "fallback completed", Answer: res.Content, ToolCalls: []ai.ToolCall{{ID: "a2a-fallback-call", Name: "fallback_echo", Input: map[string]any{"value": "a2a-fallback"}, Result: res.Content}}}, nil
}
func providerKey(provider string) string {
if v := os.Getenv("MICRO_AI_API_KEY"); v != "" {
return v
}
env := map[string]string{
"anthropic": "ANTHROPIC_API_KEY", "openai": "OPENAI_API_KEY",
"gemini": "GEMINI_API_KEY", "groq": "GROQ_API_KEY", "mistral": "MISTRAL_API_KEY",
"together": "TOGETHER_API_KEY", "atlascloud": "ATLASCLOUD_API_KEY",
}[provider]
return os.Getenv(env)
}
func main() {
provider := flag.String("provider", "mock", "LLM provider: mock (default), anthropic, openai, ...")
flag.Parse()
apiKey := ""
if *provider == "mock" {
ai.Register("mock", newMock)
} else {
apiKey = providerKey(*provider)
if apiKey == "" {
fmt.Printf("no API key for provider %q — set MICRO_AI_API_KEY or the provider's key env\n", *provider)
return
}
}
fmt.Printf("\n\033[1mA2A streaming fallback conformance (provider: %s)\033[0m\n", *provider)
reg := registry.NewMemoryRegistry()
st := store.NewMemoryStore()
var sawTool, sawRunInfo bool
ag := agent.New(
agent.Name("a2a-fallback"),
agent.Provider(*provider),
agent.APIKey(apiKey),
agent.Prompt("Use fallback_echo exactly once with value a2a-fallback, then answer with the tool result."),
agent.WithRegistry(reg),
agent.WithStore(st),
agent.WithMemory(agent.NewInMemory(8)),
agent.ModelCallTimeout(45*time.Second),
agent.WithTool("fallback_echo", "Echo the A2A fallback marker.", map[string]any{
"value": map[string]any{"type": "string", "description": "value to echo"},
}, func(ctx context.Context, input map[string]any) (string, error) {
sawTool = true
info, ok := ai.RunInfoFrom(ctx)
if !ok || info.RunID == "" || info.Agent != "a2a-fallback" {
return "", fmt.Errorf("unexpected run info: %+v", info)
}
sawRunInfo = true
if input["value"] != "a2a-fallback" {
return "", fmt.Errorf("unexpected value %v", input["value"])
}
return `{"marker":"a2a-fallback-ok"}`, nil
}),
)
card := a2a.Card("a2a-fallback", "http://example.invalid/a2a-fallback", "", nil)
handler := a2a.NewAgentStreamHandler(card, func(ctx context.Context, text string) (string, error) {
resp, err := ag.Ask(ctx, text)
if err != nil {
return "", err
}
return resp.Reply, nil
}, ag.Stream)
body := []byte(`{"jsonrpc":"2.0","id":1,"method":"message/stream","params":{"message":{"role":"user","parts":[{"kind":"text","text":"Run the A2A fallback conformance check."}],"kind":"message"}}}`)
req := httptest.NewRequest(http.MethodPost, "/", bytes.NewReader(body))
req.Header.Set("Content-Type", "application/json")
rr := httptest.NewRecorder()
handler.ServeHTTP(rr, req)
res := rr.Result()
defer res.Body.Close()
if res.StatusCode != http.StatusOK {
b, _ := io.ReadAll(res.Body)
fmt.Fprintf(os.Stderr, "unexpected status %d: %s\n", res.StatusCode, b)
os.Exit(1)
}
if ct := res.Header.Get("Content-Type"); !strings.HasPrefix(ct, "text/event-stream") {
fmt.Fprintf(os.Stderr, "content-type = %q, want text/event-stream\n", ct)
os.Exit(1)
}
payload, err := readSSEData(res.Body)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if !strings.Contains(payload, "a2a-fallback-ok") {
fmt.Fprintf(os.Stderr, "stream payload missing marker: %s\n", payload)
os.Exit(1)
}
if !sawTool || !sawRunInfo {
fmt.Fprintf(os.Stderr, "tool=%v runInfo=%v\n", sawTool, sawRunInfo)
os.Exit(1)
}
fmt.Println("\n\033[32m✓ A2A message/stream fell back to Ask and preserved tool/run metadata\033[0m")
}
func readSSEData(r io.Reader) (string, error) {
scanner := bufio.NewScanner(r)
var payload strings.Builder
for scanner.Scan() {
line := scanner.Text()
if data, ok := strings.CutPrefix(line, "data: "); ok {
payload.WriteString(data)
payload.WriteByte('\n')
}
}
if err := scanner.Err(); err != nil {
return "", err
}
if payload.Len() == 0 {
return "", errors.New("no SSE data received")
}
if !json.Valid([]byte(strings.TrimSpace(payload.String()))) {
return "", fmt.Errorf("SSE data is not JSON: %s", payload.String())
}
return payload.String(), nil
}
@@ -15,6 +15,9 @@ agent test and the harnesses in `internal/harness`:
- `agent-flow` — a workflow event that drives an agent to call services.
- `plan-delegate` — plan persistence plus agent-to-agent delegation and service
calls.
- `a2a-stream-fallback` — A2A `message/stream` through the gateway, including
fallback from unsupported provider streaming to the tool-calling `Ask` path while
preserving run metadata.
The command also emits the registered provider capability matrix so the run shows
which providers advertise model, image, video, and streaming support.
@@ -49,7 +52,8 @@ Provider keys are read from `MICRO_AI_API_KEY` or the provider-specific variable
| AtlasCloud | `ATLASCLOUD_API_KEY` |
Use `-require-configured` when you want a selected provider without a key to fail
instead of skip:
instead of skip. This is useful for manually checking that a required provider
secret is actually wired into CI before relying on that provider as covered:
```sh
go run ./internal/harness/provider-conformance \
@@ -60,10 +64,14 @@ go run ./internal/harness/provider-conformance \
## Scheduled CI behavior
The `Harness (E2E)` workflow runs on pushes and pull requests with deterministic
mock LLMs, including `provider-conformance -providers mock`. On the daily schedule and manual dispatch it also runs the live
provider conformance job. That job:
mock LLMs, including `provider-conformance -providers mock`. On the daily
schedule and manual dispatch it also runs the live provider conformance job. A
manual dispatch can narrow `providers` or `harnesses`, and can set
`require_configured=true` to fail fast when an expected repository secret is
missing; scheduled runs keep the safe default and report missing keys as skips.
That job:
1. runs the same `agent`, `universe`, `agent-flow`, and `plan-delegate` harness list,
1. runs the same `agent`, `universe`, `agent-flow`, `plan-delegate`, and `a2a-stream-fallback` harness list,
2. reads the provider keys from repository secrets,
3. skips providers whose secrets are absent,
4. fails when any configured provider fails a harness, and
+26 -2
View File
@@ -34,6 +34,8 @@ import (
_ "go-micro.dev/v6/ai/together"
)
const defaultHarnesses = "agent,universe,agent-flow,plan-delegate,a2a-stream-fallback"
var providerEnv = map[string]string{
"anthropic": "ANTHROPIC_API_KEY",
"openai": "OPENAI_API_KEY",
@@ -45,8 +47,8 @@ var providerEnv = map[string]string{
}
func main() {
providersFlag := flag.String("providers", "anthropic,openai,gemini,groq,mistral,together,atlascloud", "comma-separated providers to check; use mock for deterministic local checks")
harnessesFlag := flag.String("harnesses", "agent,universe,agent-flow,plan-delegate", "comma-separated harness names under internal/harness; agent runs the provider tool-call conformance test")
providersFlag := flag.String("providers", defaultProviders(), "comma-separated providers to check; use mock for deterministic local checks")
harnessesFlag := flag.String("harnesses", defaultHarnesses, "comma-separated harness names under internal/harness; agent runs the provider tool-call conformance test")
timeoutFlag := flag.Duration("timeout", 10*time.Minute, "timeout per provider/harness run")
requireConfiguredFlag := flag.Bool("require-configured", false, "fail when a selected live provider is missing an API key")
capabilitiesFlag := flag.Bool("capabilities", true, "print the registered provider capability matrix before running conformance")
@@ -169,6 +171,8 @@ func writeSummaryMarkdown(path string, summary conformanceSummary) error {
var b strings.Builder
b.WriteString("# Provider conformance summary\n\n")
fmt.Fprintf(&b, "Passed: %d. Skipped providers: %d. Failed: %d.\n\n", summary.Passed, summary.Skipped, summary.Failed)
fmt.Fprintf(&b, "Providers: %s.\n\n", markdownList(summary.Providers))
fmt.Fprintf(&b, "Harnesses: %s.\n\n", markdownList(summary.Harnesses))
b.WriteString("## Capability matrix\n\n")
b.WriteString(capabilityMarkdown(summary.Capabilities))
b.WriteString("\n## Harness results\n\n")
@@ -194,6 +198,17 @@ func capabilityMarkdown(rows []ai.CapabilityRow) string {
return b.String()
}
func markdownList(values []string) string {
if len(values) == 0 {
return "—"
}
escaped := make([]string, 0, len(values))
for _, value := range values {
escaped = append(escaped, "`"+strings.ReplaceAll(value, "`", "\\`")+"`")
}
return strings.Join(escaped, ", ")
}
func markdownCell(s string) string {
if s == "" {
return "—"
@@ -279,6 +294,15 @@ func repoRoot() string {
}
}
func defaultProviders() string {
providers := make([]string, 0, len(providerEnv))
for provider := range providerEnv {
providers = append(providers, provider)
}
slices.Sort(providers)
return strings.Join(providers, ",")
}
func knownProviders() string {
providers := make([]string, 0, len(providerEnv)+1)
providers = append(providers, "mock")
@@ -36,6 +36,18 @@ func TestValidateSelectionRejectsUnsafeHarnessName(t *testing.T) {
}
}
func TestDefaultProvidersTracksLiveProviderSet(t *testing.T) {
got := defaultProviders()
for _, want := range []string{"anthropic", "openai", "gemini", "groq", "mistral", "together", "atlascloud"} {
if !strings.Contains(got, want) {
t.Fatalf("defaultProviders() = %q, want %q", got, want)
}
}
if strings.Contains(got, "mock") {
t.Fatalf("defaultProviders() = %q, should not include mock in live scheduled defaults", got)
}
}
func TestCapabilityMatrixHasRegisteredProviders(t *testing.T) {
rows := ai.CapabilityRows()
if len(rows) == 0 {
@@ -105,6 +117,8 @@ func TestWriteSummaryMarkdown(t *testing.T) {
for _, want := range []string{
"# Provider conformance summary",
"Passed: 1. Skipped providers: 1. Failed: 0.",
"Providers: —.",
"Harnesses: —.",
"| mock | ✅ | — | — | — |",
"| mock | agent-flow | passed | — |",
"| live | — | skipped | missing \\| key |",
@@ -115,6 +129,14 @@ func TestWriteSummaryMarkdown(t *testing.T) {
}
}
func TestMarkdownListEscapesBackticks(t *testing.T) {
got := markdownList([]string{"agent", "bad`name"})
want := "`agent`, `bad\\`name`"
if got != want {
t.Fatalf("markdownList() = %q, want %q", got, want)
}
}
func TestWriteSummaryJSON(t *testing.T) {
path := filepath.Join(t.TempDir(), "summary.json")
summary := conformanceSummary{
@@ -0,0 +1,48 @@
package zerotoheroci
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestZeroToHeroReferenceDocs(t *testing.T) {
root := filepath.Clean(filepath.Join("..", "..", ".."))
guide := readFile(t, filepath.Join(root, "internal", "website", "docs", "guides", "zero-to-hero.md"))
for _, want := range []string{
"make harness",
"go test ./cmd/micro/cli/new -run TestZeroToOne -count=1",
"go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1",
"go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1",
"./internal/harness/zero-to-hero-ci/run.sh",
"go run ./internal/harness/agent-flow",
"make provider-conformance-mock",
"internal/harness/plan-delegate",
"internal/harness/universe",
} {
if !strings.Contains(guide, want) {
t.Fatalf("0→hero guide missing %q", want)
}
}
readme := readFile(t, filepath.Join(root, "README.md"))
if !strings.Contains(readme, "internal/website/docs/guides/zero-to-hero.md") {
t.Fatal("README does not point to the canonical 0→hero guide")
}
nav := readFile(t, filepath.Join(root, "internal", "website", "_data", "navigation.yml"))
if !strings.Contains(nav, "0→hero Reference") || !strings.Contains(nav, "/docs/guides/zero-to-hero.html") {
t.Fatal("website navigation does not expose the canonical 0→hero guide")
}
}
func readFile(t *testing.T, name string) string {
t.Helper()
data, err := os.ReadFile(name)
if err != nil {
t.Fatalf("read %s: %v", name, err)
}
return string(data)
}
+3
View File
@@ -32,6 +32,8 @@ examples:
- title: Real-World Examples
url: /docs/examples/realworld/
guides:
- title: 0→hero Reference
url: /docs/guides/zero-to-hero.html
- title: The Agent Harness
url: /docs/guides/agent-harness.html
- title: Agents and Workflows
@@ -70,6 +72,7 @@ project:
- title: Server (optional)
url: /docs/server.html
search_order:
- /docs/guides/zero-to-hero.html
- /docs/getting-started.html
- /docs/mcp.html
- /docs/architecture.html
+4 -4
View File
@@ -135,8 +135,8 @@ The store backend determines durability — file-backed by default, Postgres or
| **What** | Capability | Intelligence | Event orchestration |
| **Does** | Handles requests | Manages services | Reacts to events |
| **Knows** | Its endpoints | Its services' endpoints | Its trigger topic |
| **State** | Store | Store (memory) | Stateless per event |
| **Create** | `micro.NewService()` | `micro.NewAgent()` | `micro.NewFlow()` |
| **State** | Store | Store-backed memory | Checkpointed run history |
| **Create** | `micro.NewService("name")` | `micro.NewAgent("name")` | `micro.NewFlow("name")` |
| **Package** | `service/` | `agent/` | `flow/` |
They compose:
@@ -164,7 +164,7 @@ Or build an agent in Go:
```go
package main
import "go-micro.dev/v5"
import "go-micro.dev/v6"
func main() {
agent := micro.NewAgent("task-mgr",
@@ -176,7 +176,7 @@ func main() {
}
```
The agent package is at `go-micro.dev/v5/agent`. The full interface design is documented in [AGENT_DESIGN.md](https://github.com/micro/go-micro/blob/master/internal/docs/AGENT_DESIGN.md).
The agent implementation lives under `go-micro.dev/v6/agent`; most users create agents through the top-level `go-micro.dev/v6` API. The full interface design is documented in [AGENT_DESIGN.md](https://github.com/micro/go-micro/blob/master/internal/docs/AGENT_DESIGN.md).
---
+89
View File
@@ -0,0 +1,89 @@
---
layout: blog
title: "An Agent Is a Service: Where Agent Frameworks Are Going"
permalink: /blog/32
description: "A field guide to the agent-framework landscape — from LangChain and the first wave, through the two layers of a harness and the rise of loop engineering, to where the frameworks diverge. And why Go Micro's answer is that an agent is a service."
---
# An Agent Is a Service: Where Agent Frameworks Are Going
*June 30, 2026 • By the Go Micro Team*
There are now a lot of ways to build an agent. LangChain and LangGraph, LlamaIndex, CrewAI, Microsoft's AutoGen, Google's ADK, the model labs' own SDKs, and — most recently in our own backyard — [tRPC-Agent-Go](https://github.com/trpc-group/trpc-agent-go) from Tencent. They are not all solving the same problem, and the places they differ tell you a lot about where this is heading.
This is a field guide to that landscape, and an honest account of where Go Micro sits in it.
## The first wave: a model in a loop
The first wave of agent frameworks solved one thing: get a model to call tools in a loop until a task is done. LangChain, more than any other project, defined that category in 2022 — chains, then agents, then graphs. LlamaIndex came at it from the data and retrieval side. CrewAI and AutoGen leaned into multi-agent orchestration — crews and conversations of role-played agents. The model labs shipped their own agent SDKs so you could stay close to the metal.
That first problem — model, tools, a loop — is now largely commoditized. Every SDK does it, and they mostly do it well. Which means the interesting question has moved. It is no longer "how do I get a model to use a tool." It is everything that happens *around* the loop once the agent has to do real work: connect to real systems, hold state across restarts, recover from failure, be observed, be scheduled, and be reached by other agents. That is the part that decides whether an agent makes it out of a demo.
LangChain itself is the clearest evidence. The framework was the distribution; the value moved to *operating* agents — which is why their commercial product is LangSmith (observability, evaluation, monitoring), not the framework. The lesson the pioneer taught is that the framework gets you to a running agent, and the hard, durable, valuable problems are in operating it.
## "Agent = Model + Harness" — but a harness has two layers
LangChain has a good framing for this: an agent is a model plus a *harness* — the runtime around the model that makes it useful. The framing is right. What is usually left implicit is that "harness" has two distinct layers, and almost all the frameworks live in the first one.
**The intra-agent harness** is the runtime around a *single model*: the system prompt, the tool definitions, context management and compaction, the sandbox, self-verification, and the continuation loop that keeps the model going until it is done. LangChain and LangGraph, deepagents, Claude Code, and the model labs' SDKs are excellent at this. It is real, hard work, and it is most of what people mean when they say "agent framework."
**The operational harness** is the distributed substrate an agent *operates inside*: services exposed as typed tools, discovery and RPC, durable and resumable runs, observability, scheduling, and the protocols agents use to reach each other. This is the layer where a single agent stops being a script and becomes part of a system — where many agents, many services, and many workflows have to compose without falling over.
The first layer produces an agent. The second is where that agent has to live. Most frameworks build the first and leave the second to you — you bring your own services, your own discovery, your own durability, your own deployment. That is the gap that matters now, because the moment you have more than one agent or one service, the operational harness *is* the product.
## The loop is the new frontier
If the first wave was "a model in a loop," the direction now is what LangChain has started calling [loop engineering](https://www.langchain.com/blog/the-art-of-loop-engineering): stacking loops around the agent. It is a useful map. There is the **agent loop** (model calls tools until done), the **verification loop** (a grader checks the output against a rubric and sends failures back with feedback), the **event-driven loop** (the agent is triggered by webhooks, schedules, or messages instead of a human typing), and the **hill-climbing loop** (production traces feed back to improve the prompts, tools, and graders over time).
Notice that only the first of those four is the intra-agent harness. The other three — verification, event-driven triggers, learning from traces — are the operational harness. The frontier is moving from "answer a prompt" to **scheduled, looping, work-performing agents**: agents that run on a cadence, do real work, check their own output, and get better. That is exactly the layer that is underbuilt, and it is the layer that decides whether agents are dependable.
## Where the frameworks are going
Survey the field and a shape emerges. LangChain and LangGraph pair graph-based orchestration with LangSmith for operations, funded to build the team that operates the platform. CrewAI and AutoGen are converging on multi-agent orchestration patterns. Google's ADK is a strong code-first framework with first-class evaluation, tuned for Gemini and Google Cloud. tRPC-Agent-Go brings a production-grade Go agent SDK — LLM, Chain, Parallel, Cycle, and Graph agents; tools; MCP and A2A; memory and RAG; evaluation; agent self-evolution; OpenTelemetry — maintained by Tencent's tRPC group and validated inside Tencent.
They differ in the details, but most share two structural choices. They are an **agent SDK you run alongside your services** — the agents are a layer, and your service tier lives somewhere else and is called into. And they are **graph-centric** — you compose agents and tools into graphs and conditional workflows. That is a coherent, well-trodden approach, and for a lot of teams it is exactly right.
Go Micro starts somewhere else.
## Where Go Micro fits: an agent is a service
Go Micro's position is a single claim: **an agent is a service.** Not a layer bolted onto a service tier — the same runtime.
The reasoning is straightforward. The moment an agent has to discover services, call them, hold state, and recover from failure, it *is* a distributed system. That is precisely the problem a service framework already solves. So instead of building an agent SDK that sits next to your services, Go Micro makes agents and services the same primitives:
- **Every service endpoint is automatically an AI-callable tool**, derived from registry metadata. You do not wire tools into a graph; you write a service and it is already a tool, reachable over MCP.
- **An agent is a service.** It registers, is discovered, load-balances, exposes an `Agent.Chat` RPC, keeps store-backed memory, and is reachable over A2A — the same lifecycle as anything else you run.
- **Workflows are durable code paths, not a graph DSL.** Use a `flow` of checkpointed steps where the path is known; dispatch to an agent where it is not. The deterministic parts are plain, resumable Go; the dynamic parts are agents.
The premise is that the line between "your services" and "your agents" is accidental complexity. Remove it, and there is less to wire, less to keep in sync, and a much shorter path from a service to an agent that uses it. The operational harness — discovery, RPC, pub/sub, durable runs, observability, deployment — is not something you assemble around the framework. It *is* the framework.
This is also why Go Micro is deliberately not a graph DSL. Graphs are expressive, and for some teams that visual, declarative model is the draw. But a graph is one more thing to learn and maintain next to your services. "It is just services and durable flows" is a smaller surface to hold in your head, and it composes with everything a service already does.
## A concrete contrast: tRPC-Agent-Go
Because it is the closest neighbour — a serious, production Go framework — tRPC-Agent-Go makes the fork concrete. It is an agent SDK that runs alongside your tRPC services, organised around graph, chain, parallel, and cycle agents. Go Micro is one runtime where the agent *is* the service and orchestration is durable flows.
We will be honest about where they are ahead: tRPC-Agent-Go ships a first-class evaluation framework, agent self-evolution, AG-UI streaming, and RAG today. Go Micro has the trace foundation (OpenTelemetry run timelines, `micro runs`) and has the verification/grader loop and richer memory on the roadmap — but if you need those right now, they are further along there, with a large team behind them. Pretending the checklists match would help no one.
What Go Micro offers in return is the thing an SDK-alongside-your-services cannot: services that become tools with zero glue, agents that are first-class services, and one set of primitives — service, agent, flow — instead of a service stack plus an agent layer plus a graph runtime.
## The direction we're building
If scheduled, looping, work-performing agents are where this goes, then the operational harness is the thing to get right, and loops are the organising idea. Go Micro already has the agent loop, durable event-driven flows, and the trace foundation for learning. The verification loop — grade a step's output against a rubric and route failures back with feedback — is the next primitive, building on the supervised loop and retry machinery already there. Durable agent runs, streaming end to end, and richer observability are on the same line. The aim is not to win a feature checklist; it is to be the runtime where an operating agent is dependable.
There is one more piece of evidence we find hard to argue with: Go Micro is increasingly built by its own loop — an autonomous improvement loop running in CI, opening and merging its own changes against a thesis. An agent harness, operated by agents, building itself. If it is good enough to do that, it is good enough to operate yours.
## Open protocols, different homes
None of this is winner-take-all, and it should not be. Every serious framework here speaks **MCP** for tools and **A2A** for agents. A Go Micro agent and a tRPC-Agent-Go agent can call each other; either can consume the other's tools; an ADK or LangGraph agent can plug into a Go Micro runtime over A2A, and the reverse. The protocols are the commons.
So the real question is not which framework wins. It is where your agents should *live*. The answer that Go Micro is built around is that when an agent has to operate inside a real system, it is a distributed system — and the simplest place to build it is the runtime where your services already live.
---
*Go Micro is an open source agent harness and service framework for Go. [Star us on GitHub](https://github.com/micro/go-micro).*
<div class="post-nav">
<div><a href="/blog/31">&larr; How Go Micro Builds Itself</a></div>
<div><a href="/blog/">All Posts</a></div>
</div>
+3 -3
View File
@@ -11,7 +11,7 @@ description: "Unified service creation, cleaner handler registration, and modula
*March 4, 2026 — By the Go Micro Team*
Go Micro has always prioritized getting out of your way. But over time, the API accumulated multiple ways to do the same thing — `micro.NewService()`, `micro.NewService()`, `service.New()`, three different handler registration patterns. If you're building something for AI agents or running a modular monolith, you shouldn't have to choose between equivalent APIs.
Go Micro has always prioritized getting out of your way. But over time, the API accumulated multiple ways to do the same thing — `micro.New(name)`, `micro.NewService(micro.Name(...))`, `service.New()`, three different handler registration patterns. If you're building something for AI agents or running a modular monolith, you shouldn't have to choose between equivalent APIs.
We've cleaned it up. Here's what changed and why.
@@ -33,7 +33,7 @@ service := micro.NewService("greeter")
service := micro.NewService("greeter", micro.Address(":8080"))
```
Name is always the first argument. Options follow. `NewService` still works (it's deprecated, not removed), but every example, doc, and guide now uses `micro.NewService()`.
Name is always the first argument. Options follow. `micro.New` still works as a deprecated alias, but every example, doc, and guide now uses `micro.NewService("name")`.
## Clean Handler Registration
@@ -107,7 +107,7 @@ Your Go comments become tool descriptions. Your struct tags become parameter sch
If you're building new services, use `micro.NewService("name", opts...)` and `service.Handle()`. That's it.
If you have existing code using `micro.NewService()` or `service.Server().Handle()`, everything still works — we didn't break anything. But the docs, examples, and guides all point to the new patterns now.
If you have existing code using `micro.New("name")` or `service.Server().Handle()`, it still works. For v6, update name-less `micro.NewService(opts...)` calls to `micro.NewService("name", opts...)`. The docs, examples, and guides all point to the new patterns now.
The goal is simple: when someone asks "how do I create a service?", there should be exactly one answer.
+7
View File
@@ -11,6 +11,13 @@ permalink: /blog/
<div class="posts">
<article style="margin-bottom: 2rem; padding-bottom: 1.5rem; border-bottom: 1px solid #e5e5e5;">
<h2 style="margin: 0 0 0.5rem;"><a href="/blog/32">An Agent Is a Service: Where Agent Frameworks Are Going</a></h2>
<p class="meta" style="color: #666; font-size: 0.85rem;">June 30, 2026</p>
<p>A field guide to the agent-framework landscape — LangChain and the first wave, the two layers of a harness, the rise of loop engineering, and where the frameworks (LangGraph, ADK, CrewAI, AutoGen, tRPC-Agent-Go) diverge. And why Go Micro's answer is that an agent is a service.</p>
<a href="/blog/32">Read more &rarr;</a>
</article>
<article style="margin-bottom: 2rem; padding-bottom: 1.5rem; border-bottom: 1px solid #e5e5e5;">
<h2 style="margin: 0 0 0.5rem;"><a href="/blog/31">How Go Micro Builds Itself</a></h2>
<p class="meta" style="color: #666; font-size: 0.85rem;">June 25, 2026</p>
+1 -1
View File
@@ -16,7 +16,7 @@ Go Micro has three core abstractions:
## Prerequisites
- **Go 1.21+** for development. The `curl` install below gives you the `micro` binary without Go, but `micro run` compiles your services, so you'll want Go installed to build them.
- **Go 1.24+** for development. The `curl` install below gives you the `micro` binary without Go, but `micro run` compiles your services, so you'll want Go installed to build them.
- An **LLM provider key** (Anthropic, OpenAI, Gemini, …) *only* for the AI features — `micro run --prompt`, `micro chat`, and agents. Plain services need no key. Set it before running, e.g. `export ANTHROPIC_API_KEY=sk-ant-...`.
## Install
@@ -31,7 +31,7 @@ your stack — the harness *is* the stack.
| Tools | Every service endpoint is an MCP-callable tool from registry metadata — no extra code | Shipped |
| Memory | Store-backed agent memory (`AgentMemory`), durable across restarts | Shipped |
| Guardrails | `MaxSteps`, `LoopLimit`, `ApproveTool`, tool wrappers — enforced at the call site | Shipped |
| Workflows | Durable flows; `flow.Loop` for run-until-done | Shipped |
| Workflows | Durable flows; `micro.FlowLoop` for run-until-done | Shipped |
| Planning / delegation | Built-in `plan` and `delegate` tools on every agent | Shipped |
| Discovery & RPC | Registry + client; agents and services find and call each other | Shipped |
| Interop | MCP (tools), A2A (agents), x402 (paid tools) | Shipped |
@@ -69,6 +69,14 @@ if err != nil {
_ = resp
```
Choose the boundary deliberately: use a durable flow when the steps are known
(`reserve`, `charge`, `confirm`) and each step has deterministic retry/resume
semantics. Use a checkpointed agent run when the model is deciding which tools to
call or how many turns it needs, but the side effects of completed tool calls
still need crash-safe resume. Flows and agents share the same `Checkpoint`
interface, so a flow can safely dispatch to a checkpointed agent for the
open-ended part.
For human-in-the-loop runs that pause through the built-in `request_input` tool,
resume with the operator's response:
+2 -2
View File
@@ -18,11 +18,11 @@ you're done") has no natural ceiling. So a usable loop needs two things:
1. a **stop condition** — how it decides it's done, and
2. a **hard cap** — a guardrail that guarantees it always terminates.
Go Micro gives you both as a flow step: `flow.Loop`.
Go Micro gives you both as a flow step: `micro.FlowLoop`.
## The shape
`flow.Loop` is a `StepFunc`, so it drops into a flow's ordered, checkpointed
`micro.FlowLoop` is a `StepFunc`, so it drops into a flow's ordered, checkpointed
step list like any other step. It runs a **body** step repeatedly, carrying the
flow `State` from one pass to the next, until a stop condition fires or the
iteration cap is hit — whichever comes first.
@@ -89,14 +89,32 @@ Agents use store-backed conversation memory by default, scoped under the agent's
name. That makes short restarts boring: the next `Ask` reloads the retained
history from the same store backend you already use for services and flows.
Long-running agents can also keep model context bounded without losing useful
prior context:
prior context. If you want retrieval without summaries, enable bounded active
context plus a durable archive of every turn:
```go
a := micro.NewAgent("conductor",
micro.AgentServices("task"),
micro.AgentProvider("anthropic"),
micro.AgentRetrievalMemory(40), // active messages kept in prompt context
micro.AgentMemoryRecallLimit(5), // archived turns recalled per Ask
)
```
`AgentRetrievalMemory(activeLimit)` switches the default memory to a store-backed
retriever. The active conversation is capped at `activeLimit`, every turn is
archived in the same scoped store used by the agent, and future asks inject
matching archived turns ahead of active context. The built-in ranking is
deterministic and credential-free for CI.
When you also want a rolling summary in active context, use compacting memory:
```go
a := micro.NewAgent("conductor",
micro.AgentServices("task"),
micro.AgentProvider("anthropic"),
micro.AgentCompactMemory(40, 12), // max active messages, recent messages kept verbatim
micro.AgentMemoryRecallLimit(5), // archived turns recalled per Ask
micro.AgentMemoryRecallLimit(5), // compacted turns recalled per Ask
)
```
@@ -105,8 +123,7 @@ deterministic compactor. Once active history grows past `maxMessages`, older
turns move into the durable archive, a provider-neutral summary is injected into
active context, and the newest `keepRecent` messages stay verbatim. On future
asks, archived turns whose text matches the current request are recalled ahead of
the active context. The built-in retrieval is intentionally simple and
credential-free for CI; teams that need embeddings or a vector database can still
the active context. Teams that need embeddings or a vector database can still
provide their own `AgentMemory` implementation.
This is harness memory, not prompt-layer orchestration: services remain the
@@ -227,6 +227,53 @@ and an ADK agent (in any language) can call each other over A2A, and either can
consume the other's MCP tools. A common pattern is to run Go Micro as the service
mesh / runtime and let ADK (or any A2A agent) plug into it.
## vs tRPC-Agent-Go
[tRPC-Agent-Go](https://github.com/trpc-group/trpc-agent-go) (maintained by tRPC-Group,
validated inside Tencent) is a production-grade Go framework for agent systems:
LLM / Chain / Parallel / Cycle / Graph agents, function tools, MCP, A2A, AG-UI, Redis
memory and RAG, evaluation, agent self-evolution, and OpenTelemetry. It's a serious,
well-resourced project.
They overlap heavily on agents but take a different approach. tRPC-Agent-Go is an **agent
SDK you run alongside your services** — you compose agents and tools into graphs and
conditional workflows, and your microservices (tRPC) live separately and are called
into. Go Micro starts from the premise that **an agent is a service** — one runtime
where every endpoint is automatically a tool, an agent registers and is discovered and
load-balanced like anything else, and workflows are durable code paths rather than a
graph DSL. The premise is that the line between "your services" and "your agents" is
accidental complexity; remove it and there's less to wire and keep in sync.
| | Go Micro | tRPC-Agent-Go |
|---|----------|---------------|
| **Primary unit** | A harnessed service (an agent is a service with an LLM inside) | An agent |
| **Orchestration** | Durable `flow` steps + `Loop` — plain code paths | Graph / Chain / Parallel / Cycle agents (graph DSL) |
| **Services as tools** | Every endpoint is automatically an MCP tool | Function tools + MCP, wired explicitly |
| **Service runtime** | Built in — agents *are* services (registry, RPC, load balancing, pub/sub) | Runs alongside your existing service stack (tRPC) |
| **MCP / A2A** | Both, generated from the registry | Both |
| **Evaluation / self-evolution** | Verification loop on the roadmap; not yet first-class | First-class today |
| **Memory / RAG** | Store-backed memory (Postgres, NATS KV, file); RAG on the roadmap | In-memory / Redis memory; RAG today |
| **Observability** | OpenTelemetry run timelines, `micro runs` | OpenTelemetry, Langfuse examples |
| **Backing** | Independent, community | tRPC-Group / Tencent |
### When to choose tRPC-Agent-Go
- You want a graph/workflow DSL for composing agents and tools
- You're on tRPC, or want to add agents alongside an existing service stack
- You want first-class evaluation and self-evolution today, with a large team behind it
### When to choose Go Micro
- You want one runtime where services, agents, and flows are the same primitives —
registered, discoverable, and deployed the same way
- You want your existing services to become agent tools with zero extra code
- You prefer durable flows and plain code paths over a graph DSL, in a small,
independent framework you can hold in your head
### They interoperate
Both speak **MCP** and **A2A**, so a Go Micro agent and a tRPC-Agent-Go agent can call
each other over A2A, and either can consume the other's MCP tools. You can run Go Micro
as the service-and-agent runtime and still reach an agent built on tRPC-Agent-Go.
## Feature Deep Dive
### Service Discovery
@@ -13,7 +13,7 @@ They are exposed to the model as ordinary tools. There is no separate graph runt
## Prerequisites
- Go 1.21+
- Go 1.24+
- An API key for any supported provider (Anthropic, OpenAI, Gemini, Groq, Mistral, Together, Atlas Cloud)
```bash
@@ -0,0 +1,84 @@
---
layout: default
---
# 0→hero reference path
The 0→hero path is the maintained, no-secret reference for the Go Micro
services → agents → workflows lifecycle. It ties the CLI inner loop and the
runtime harness together so a contributor can prove the framework still works as
one system, not as separate demos.
Use it when you want to answer: "Can I scaffold a service, run it locally, talk
to an agent, inspect durable work, and reach the deployment boundary without
cloud credentials?"
## What the contract covers
| Boundary | Contract | CI check |
| --- | --- | --- |
| Scaffold | `micro new` generates a runnable service with and without MCP support. | `go test ./cmd/micro/cli/new -run TestZeroToOne -count=1` |
| Run | `micro run` remains the local development entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Chat | `micro chat` remains the interactive agent entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Inspect | `micro inspect agent`, `micro inspect flow`, and `micro flow runs` remain discoverable for run history. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Deploy | `micro deploy --dry-run` resolves deploy targets without touching remote infrastructure. | `go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1` |
| Runtime | Real services, agents, durable flows, store-backed history, delegation, and A2A run with only the model mocked. | `./internal/harness/zero-to-hero-ci/run.sh` and `make provider-conformance-mock` |
## Run the whole no-secret path
From the repository root:
```sh
make harness
```
That target runs the scaffold contract, the CLI boundary smoke tests, the
0→hero runtime harnesses, the event-driven agent-flow harness, and mock provider
conformance. It is intentionally deterministic: no provider key, cloud account,
SSH access, or remote service is required.
## Run focused checks while iterating
Use the smaller checks when you are working on one seam:
```sh
# Scaffold → run/call contract.
go test ./cmd/micro/cli/new -run TestZeroToOne -count=1
# CLI inner-loop commands: run, chat, inspect, flow runs, deploy --dry-run.
go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1
go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1
# Durable services → agents → workflows reference scenarios.
./internal/harness/zero-to-hero-ci/run.sh
# Event-as-prompt agent flow.
go run ./internal/harness/agent-flow
# Cross-provider semantics with the deterministic mock provider.
make provider-conformance-mock
```
## Reference scenarios
- [`internal/harness/plan-delegate`](https://github.com/micro/go-micro/tree/master/internal/harness/plan-delegate)
is the compact 0→hero scenario: real task and notify services, a conductor
agent, a comms agent, plan persistence, delegation, and a workflow handoff.
- [`internal/harness/universe`](https://github.com/micro/go-micro/tree/master/internal/harness/universe)
boots a larger mini-world: inventory, payment, order confirmation, a concierge
agent, durable checkpoint/resume, agent run history, flow run history, and A2A
reachability.
- [`internal/harness/agent-flow`](https://github.com/micro/go-micro/tree/master/internal/harness/agent-flow)
shows the event-driven path where a `user.created` event prompts an agent to
call services and complete onboarding.
Together these scenarios keep the North Star executable: services expose typed
capabilities, agents use those capabilities with memory and guardrails, and
workflows compose the work over time.
## Keeping the guide honest
If you change the CLI inner loop, durable flow APIs, agent run history, or the
provider/tool semantics, update this guide and the harness in the same PR. The
point of 0→hero is not a polished sample app that drifts from reality; it is a
CI-verifiable contract that the documented lifecycle still works.
+28
View File
@@ -116,6 +116,24 @@ func AgentLoopLimit(n int) AgentOption { return agent.LoopLimit(n) }
// each action the agent takes.
func AgentApproveTool(fn ApproveFunc) AgentOption { return agent.ApproveTool(fn) }
// AgentModelCallTimeout sets the timeout for each provider Generate call.
func AgentModelCallTimeout(d time.Duration) AgentOption { return agent.ModelCallTimeout(d) }
// AgentModelRetry sets the provider retry budget and backoff for transient failures.
func AgentModelRetry(maxAttempts int, backoff time.Duration) AgentOption {
return agent.ModelRetry(maxAttempts, backoff)
}
// AgentToolCallTimeout sets the timeout for each agent tool execution.
func AgentToolCallTimeout(d time.Duration) AgentOption { return agent.ToolCallTimeout(d) }
// AgentToolRetry sets the tool retry budget and backoff for transient failures.
// Attempts include the first call. Retries are opt-in because tools can have
// side effects; keep handlers idempotent before enabling this.
func AgentToolRetry(maxAttempts int, backoff time.Duration) AgentOption {
return agent.ToolRetry(maxAttempts, backoff)
}
// Memory is an agent's pluggable conversation memory.
type Memory = agent.Memory
@@ -128,6 +146,12 @@ type ToolFunc = agent.ToolFunc
// NewMemory returns the default store-backed agent memory.
func NewMemory(s store.Store, key string, limit int) Memory { return agent.NewMemory(s, key, limit) }
// NewRetrievalMemory returns store-backed memory with bounded active context
// and durable retrieval over every prior turn.
func NewRetrievalMemory(s store.Store, key string, activeLimit int) Memory {
return agent.NewRetrievalMemory(s, key, activeLimit)
}
// NewCompactingMemory returns store-backed memory with deterministic
// summarization and retrieval controls.
func NewCompactingMemory(s store.Store, key string, maxMessages, keepRecent int) Memory {
@@ -140,6 +164,10 @@ func NewInMemory(limit int) Memory { return agent.NewInMemory(limit) }
// AgentMemory sets the agent's conversation memory (default: store-backed).
func AgentMemory(m Memory) AgentOption { return agent.WithMemory(m) }
// AgentRetrievalMemory enables deterministic default-memory retrieval without
// compaction; activeLimit bounds active context while every turn is archived.
func AgentRetrievalMemory(activeLimit int) AgentOption { return agent.RetrievalMemory(activeLimit) }
// AgentCompactMemory enables deterministic default-memory compaction and
// retrieval for long-running agents.
func AgentCompactMemory(maxMessages, keepRecent int) AgentOption {