Compare commits

..

20 Commits

Author SHA1 Message Date
Codex 82ba4323b9 docs: align public AI harness facts
Harness (E2E) / Harnesses (mock LLM) (push) Waiting to run
Harness (E2E) / Provider harnesses (live LLM conformance) (push) Waiting to run
Lint / golangci-lint (push) Waiting to run
Run Tests / Unit Tests (push) Waiting to run
Run Tests / Etcd Integration Tests (push) Waiting to run
2026-07-01 08:39:58 +00:00
Asim Aslam 04e8759d41 test agent checkpoint resume after restart (#3529)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 09:06:46 +01:00
Asim Aslam 78725135aa docs(priorities): refresh architect queue (#3527)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 08:33:01 +01:00
Asim Aslam 2ff64ff0b2 Document canonical 0-to-hero reference path (#3522)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 07:12:17 +01:00
Asim Aslam 7a8e7cd9ae docs(priorities): advance queue to 0-to-hero reference (#3520)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 06:37:15 +01:00
Asim Aslam b58eed1698 Add retrieval-backed agent memory (#3518)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 04:59:16 +01:00
Asim Aslam c92c9cc244 docs(priorities): refresh architect queue (#3516)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 04:06:59 +01:00
Asim Aslam a2bf43e9ef test: broaden stream provider conformance (#3512)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 03:37:26 +01:00
Asim Aslam 49bac7e4a8 trace scheduled flow dispatch metadata (#3510)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 02:42:31 +01:00
Asim Aslam 5aac7e3ca0 docs(priorities): refresh architect queue after scheduling (#3507)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 01:51:21 +01:00
Asim Aslam d86585bf5c Add scheduled flow agent harness (#3505)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 01:02:29 +01:00
Asim Aslam 88e2b58711 docs(priorities): refresh architect queue (#3503)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 00:38:51 +01:00
Asim Aslam d259383645 Harden agent terminal failure statuses (#3499)
Co-authored-by: Codex <codex@openai.com>
2026-07-01 00:03:21 +01:00
Asim Aslam 2d7ee300a4 docs(priorities): refresh architect queue after conformance (#3497)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 23:40:19 +01:00
Asim Aslam 57fa4e3b7a Add mock provider conformance target (#3495)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 23:09:19 +01:00
Asim Aslam 0f1917f26b docs(priorities): refresh architect queue after verification (#3493)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 22:40:05 +01:00
Asim Aslam 6e9c5e87e9 Add flow step verification loop (#3489)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 22:18:01 +01:00
Asim Aslam d8bb892425 docs(priorities): refresh architect queue after a2a continuity (#3487)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 21:37:56 +01:00
Asim Aslam 064a112c6b a2a: expose resubscribe and input-required support (#3484)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 21:23:43 +01:00
Asim Aslam 24d103c658 docs(priorities): refresh architect queue after memory (#3482)
Co-authored-by: Codex <codex@openai.com>
2026-06-30 20:47:39 +01:00
33 changed files with 1091 additions and 81 deletions
+8 -1
View File
@@ -8,7 +8,7 @@ LDFLAGS = -X $(GIT_IMPORT).BuildDate=$(BUILD_DATE) -X $(GIT_IMPORT).GitCommit=$(
# GORELEASER_DOCKER_IMAGE = ghcr.io/goreleaser/goreleaser-cross:v1.25.7
GORELEASER_DOCKER_IMAGE = ghcr.io/goreleaser/goreleaser:latest
.PHONY: test test-race test-coverage harness provider-conformance lint fmt install-tools proto clean help gorelease-dry-run gorelease-dry-run-docker
.PHONY: test test-race test-coverage harness provider-conformance-mock provider-conformance lint fmt install-tools proto clean help gorelease-dry-run gorelease-dry-run-docker
# Default target
help:
@@ -19,6 +19,7 @@ help:
@echo " make test-coverage - Run tests with coverage"
@echo " make lint - Run linter"
@echo " make harness - Run deterministic getting-started and end-to-end harnesses"
@echo " make provider-conformance-mock - Run cross-provider harness with deterministic mock provider"
@echo " make provider-conformance - Run harnesses against configured live providers"
@echo " make fmt - Format code"
@echo " make install-tools - Install development tools"
@@ -50,6 +51,12 @@ harness:
go test ./cmd/micro/cli/new -run TestZeroToOne -count=1
./internal/harness/zero-to-hero-ci/run.sh
go run ./internal/harness/agent-flow
$(MAKE) provider-conformance-mock
# Run the shared provider conformance contract with the deterministic mock
# provider. This is the no-secret path used by CI and local dogfooding to keep
# provider-facing agent/tool semantics covered on every machine.
provider-conformance-mock:
go run ./internal/harness/provider-conformance -providers mock
# Run the same harnesses against every configured live provider. Providers
+4 -3
View File
@@ -65,7 +65,7 @@ curl -X POST http://localhost:8080/api/helloworld/Helloworld.Call \
```
This scaffold → run → call path is covered by the no-secret CI harness. To run
the same local contract (including the 0→hero services → agents → workflows path,
the same local contract (including the [0→hero services → agents → workflows path](internal/website/docs/guides/zero-to-hero.md),
chat/inspect CLI boundaries, and deploy dry-run), use:
```bash
@@ -401,8 +401,8 @@ Swap providers with a single import — same interface everywhere:
| Google Gemini | `gemini-2.5-flash` |
| Groq | `llama-3.3-70b-versatile` |
| Mistral | `mistral-large-latest` |
| Together AI | `Llama-3.3-70B-Instruct-Turbo` |
| Atlas Cloud | `llama-3.3-70b` |
| Together AI | `meta-llama/Llama-3.3-70B-Instruct-Turbo` |
| Atlas Cloud | `deepseek-ai/DeepSeek-V3-0324` |
```go
m := ai.New("anthropic", ai.WithAPIKey(key))
@@ -423,6 +423,7 @@ See [all examples](examples/README.md).
- [Getting Started](internal/website/docs/getting-started.md)
- [AI Integration](internal/website/docs/ai-integration.md)
- [0→hero Reference](internal/website/docs/guides/zero-to-hero.md)
- [Agents and Workflows](internal/website/docs/guides/agents-and-workflows.md)
- [Agent Design](internal/docs/AGENT_DESIGN.md)
- [Plan & Delegate](internal/website/docs/guides/plan-delegate.md)
+8 -3
View File
@@ -179,6 +179,8 @@ func (a *agentImpl) setupWithToolHandler(handler ai.ToolHandler) {
a.mem = NewInMemory(a.opts.HistoryLimit)
case a.opts.MemoryCompaction.MaxMessages > 0:
a.mem = NewCompactingMemoryWithOptions(a.stateStore(), "history", a.opts.MemoryCompaction)
case a.opts.MemoryRetrievalLimit > 0:
a.mem = NewRetrievalMemory(a.stateStore(), "history", a.opts.MemoryRetrievalLimit)
default:
a.mem = NewMemory(a.stateStore(), "history", a.opts.HistoryLimit)
}
@@ -298,12 +300,15 @@ func (a *agentImpl) askLocked(ctx context.Context, runID, message, parentRunID s
Backoff: a.opts.ModelRetryBackoff,
})
if err != nil {
run.Status = "failed"
run.Steps[0].Status = "failed"
run.Steps[0].Error = err.Error()
run.Status = agentRunFailureStatus(err)
if a.currentRun != nil {
run.Steps = a.currentRun.Steps
}
if len(run.Steps) == 0 {
run.Steps = []flow.StepRecord{{Name: agentAskStep}}
}
run.Steps[0].Status = run.Status
run.Steps[0].Error = err.Error()
_ = a.saveRun(ctx, run)
return nil, err
}
+14 -1
View File
@@ -171,13 +171,26 @@ func (a *agentImpl) pending(ctx context.Context) ([]flow.Run, error) {
func terminalAgentRunStatus(status string) bool {
switch status {
case "done", "canceled", "expired":
case "done", "canceled", "timeout", "rate_limited", "expired":
return true
default:
return false
}
}
func agentRunFailureStatus(err error) string {
switch ai.ClassifyError(err) {
case ai.ErrorKindCanceled:
return "canceled"
case ai.ErrorKindTimeout:
return "timeout"
case ai.ErrorKindRateLimited:
return "rate_limited"
default:
return "failed"
}
}
func (a *agentImpl) checkpointToolWrap(next ai.ToolHandler) ai.ToolHandler {
return func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
if a.opts.Checkpoint == nil || a.currentRun == nil {
+69
View File
@@ -102,6 +102,75 @@ func TestResumeFailedCheckpointDoesNotReplayCompletedTool(t *testing.T) {
}
}
func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "restart-resume-agent")
toolRuns := 0
modelCalls := 0
failFirst := true
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
modelCalls++
if opts.ToolHandler != nil {
res := opts.ToolHandler(ctx, ai.ToolCall{ID: "call-1", Name: "external.provision", Input: map[string]any{"service": "api"}})
if res.Content != "provisioned" {
t.Fatalf("tool result = %q, want provisioned", res.Content)
}
}
if failFirst {
failFirst = false
return nil, errors.New("process stopped after tool checkpoint")
}
return &ai.Response{Reply: "resumed after restart"}, nil
}
defer func() { fakeGen = nil }()
newAgent := func() *agentImpl {
return newTestAgent(Name("restart-resume-agent"), WithCheckpoint(cp),
WithTool("external.provision", "provision service once", nil, func(context.Context, map[string]any) (string, error) {
toolRuns++
return "provisioned", nil
}))
}
first := newAgent()
_, err := first.Ask(ctx, "provision api")
if err == nil {
t.Fatal("Ask succeeded, want simulated process stop")
}
if toolRuns != 1 {
t.Fatalf("tool executions after failed Ask = %d, want 1", toolRuns)
}
runs, err := Pending(ctx, first)
if err != nil {
t.Fatalf("Pending before restart: %v", err)
}
if len(runs) != 1 {
t.Fatalf("Pending before restart returned %d runs, want 1", len(runs))
}
restarted := newAgent()
resp, err := Resume(ctx, restarted, runs[0].ID)
if err != nil {
t.Fatalf("Resume after restart: %v", err)
}
if resp.Reply != "resumed after restart" || resp.RunID != runs[0].ID {
t.Fatalf("response = %#v, want resumed reply on original run id", resp)
}
if toolRuns != 1 {
t.Fatalf("tool executions after restart resume = %d, want checkpointed tool not replayed", toolRuns)
}
if modelCalls != 2 {
t.Fatalf("model calls = %d, want initial call plus resumed call", modelCalls)
}
loaded, ok, err := cp.Load(ctx, runs[0].ID)
if err != nil || !ok {
t.Fatalf("Load resumed run ok=%v err=%v", ok, err)
}
if loaded.Status != "done" || loaded.ParentID != runs[0].ParentID {
t.Fatalf("loaded run status/parent = %s/%s, want done/%s", loaded.Status, loaded.ParentID, runs[0].ParentID)
}
}
func TestResumeFailedCheckpointDoesNotDuplicateCompactedMemory(t *testing.T) {
ctx := context.Background()
st := store.NewMemoryStore()
+25 -6
View File
@@ -56,6 +56,16 @@ func NewMemory(s store.Store, key string, limit int) Memory {
return m
}
// NewRetrievalMemory returns store-backed memory that keeps a bounded active
// conversation and archives every turn for retrieval. It is useful when callers
// want relevant durable recall without summary compaction in the active context.
// A nil store or empty key keeps only the active in-process buffer.
func NewRetrievalMemory(s store.Store, key string, activeLimit int) Memory {
m := &storeMemory{store: s, key: key, hist: ai.NewHistory(activeLimit), retrieveAll: true}
m.load()
return m
}
// NewCompactingMemory returns store-backed memory with explicit compaction and
// retrieval controls. It keeps all messages in the backing store, compacts older
// turns into a deterministic summary when the conversation exceeds maxMessages,
@@ -100,16 +110,20 @@ func NewInMemory(limit int) Memory {
// storeMemory is the default Memory: an ai.History buffer optionally
// persisted to a store.
type storeMemory struct {
mu sync.Mutex
store store.Store
key string
hist *ai.History
compaction MemoryCompaction
archive []ai.Message
mu sync.Mutex
store store.Store
key string
hist *ai.History
compaction MemoryCompaction
archive []ai.Message
retrieveAll bool
}
func (m *storeMemory) Add(role, content string) {
m.mu.Lock()
if m.retrieveAll {
m.archive = append(m.archive, ai.Message{Role: role, Content: content})
}
m.hist.Add(role, content)
m.mu.Unlock()
m.compact()
@@ -133,6 +147,8 @@ func (m *storeMemory) Clear() {
// Recall returns archived messages whose content contains words from query.
// It is deterministic and provider-neutral: no embeddings or model calls are
// required, but semantic/vector stores can replace Memory for richer retrieval.
// When created with NewRetrievalMemory the archive contains every persisted
// turn; when created with NewCompactingMemory it contains compacted older turns.
func (m *storeMemory) Recall(query string, limit int) []ai.Message {
m.mu.Lock()
defer m.mu.Unlock()
@@ -186,6 +202,9 @@ func (m *storeMemory) load() {
}
m.mu.Lock()
m.archive = state.Archive
if m.retrieveAll && len(m.archive) == 0 {
m.archive = append(m.archive, state.Messages...)
}
for _, msg := range state.Messages {
m.hist.Add(msg.Role, msg.Content)
}
+43
View File
@@ -64,6 +64,49 @@ func TestWithMemoryUsed(t *testing.T) {
}
}
func TestRetrievalMemoryArchivesAllTurnsAndRanksRelevant(t *testing.T) {
st := store.NewMemoryStore()
m := NewRetrievalMemory(st, "agent/retrieval/history", 2)
m.Add("user", "alpha budget is 42")
m.Add("assistant", "noted")
m.Add("user", "beta owner is lee")
m.Add("assistant", "tracked")
m.Add("user", "alpha owner is sam")
if got := len(m.Messages()); got != 2 {
t.Fatalf("active messages = %d, want bounded history of 2", got)
}
recall, ok := m.(MemoryRecall)
if !ok {
t.Fatal("retrieval memory should support recall")
}
recalled := recall.Recall("alpha budget", 2)
if len(recalled) == 0 {
t.Fatal("expected relevant recalled turns")
}
if got := recalled[0].Content.(string); !strings.Contains(got, "alpha budget is 42") {
t.Fatalf("top recall = %q, want archived alpha budget turn", got)
}
}
func TestRetrievalMemoryPersistsArchiveAcrossReload(t *testing.T) {
st := store.NewMemoryStore()
m := NewRetrievalMemory(st, "agent/retrieval/reload", 1)
m.Add("user", "alpha budget is 42")
m.Add("assistant", "noted")
m.Add("user", "beta budget is 7")
reloaded := NewRetrievalMemory(st, "agent/retrieval/reload", 1)
recalled := reloaded.(MemoryRecall).Recall("alpha budget", 1)
if len(recalled) != 1 {
t.Fatalf("recalled %d messages, want 1", len(recalled))
}
if got := recalled[0].Content.(string); !strings.Contains(got, "alpha budget is 42") {
t.Fatalf("reloaded recall = %q, want alpha budget", got)
}
}
func TestCompactingMemoryRecallRanksSpecificMatches(t *testing.T) {
m := NewCompactingMemory(store.NewMemoryStore(), "agent/rank/history", 3, 1).(MemoryRecall)
writer := m.(Memory)
+16
View File
@@ -64,6 +64,10 @@ type Options struct {
// Memory is the agent's conversation memory. Nil = the default
// store-backed memory (durable across restarts).
Memory Memory
// MemoryRetrievalLimit enables retrieval-backed default memory without
// compaction. The active conversation stays bounded to this many messages
// while every turn is archived for deterministic recall.
MemoryRetrievalLimit int
// MemoryCompaction enables deterministic compaction/retrieval on the
// default store-backed memory. Custom Memory implementations can expose
// retrieval by implementing MemoryRecall.
@@ -240,6 +244,18 @@ func WithMemory(m Memory) Option {
return func(o *Options) { o.Memory = m }
}
// RetrievalMemory enables deterministic, store-backed retrieval memory for
// the default agent memory without compaction. Active context is capped at
// activeLimit messages while every turn is archived in the store for Recall.
func RetrievalMemory(activeLimit int) Option {
return func(o *Options) {
o.MemoryRetrievalLimit = activeLimit
if o.MemoryRecallLimit == 0 {
o.MemoryRecallLimit = 5
}
}
}
// CompactMemory enables deterministic, store-backed memory compaction for the
// default agent memory. Older turns are summarized once active context exceeds
// maxMessages, keepRecent newest turns remain verbatim, and recalled archived
+30 -5
View File
@@ -41,6 +41,10 @@ const (
AttrErrorKind = "agent.error.kind"
AttrCheckpointStatus = "agent.checkpoint.status"
AttrCheckpointStage = "agent.checkpoint.stage"
AttrFlowName = "agent.flow.name"
AttrFlowStep = "agent.flow.step"
AttrDispatch = "agent.dispatch"
AttrTrigger = "agent.trigger"
)
type RunEvent struct {
@@ -124,8 +128,12 @@ func (a *agentImpl) startRun(ctx context.Context, message string) (context.Conte
}
}
ctx, span := a.tracer().Start(ctx, spanNameRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(
attribute.String(AttrRunID, info.RunID), attribute.String(AttrParentRunID, info.ParentID), attribute.String(AttrAgentName, info.Agent)))
attrs := appendRunInfoAttributes([]attribute.KeyValue{
attribute.String(AttrRunID, info.RunID),
attribute.String(AttrParentRunID, info.ParentID),
attribute.String(AttrAgentName, info.Agent),
}, info)
ctx, span := a.tracer().Start(ctx, spanNameRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(attrs...))
a.recordSpanEvent(span, runEvent)
return ctx, func(err error) {
latency := time.Since(start).Milliseconds()
@@ -171,16 +179,17 @@ func (m *tracedModel) Generate(ctx context.Context, req *ai.Request, opts ...ai.
return resp, err
}
ctx, span := m.a.tracer().Start(ctx, spanNameModelCall, trace.WithAttributes(
attrs := appendRunInfoAttributes([]attribute.KeyValue{
attribute.String(AttrRunID, info.RunID),
attribute.String(AttrParentRunID, info.ParentID),
attribute.String(AttrAgentName, info.Agent),
attribute.String(AttrProvider, provider),
attribute.String(AttrModel, model),
))
}, info)
ctx, span := m.a.tracer().Start(ctx, spanNameModelCall, trace.WithAttributes(attrs...))
resp, err := m.Model.Generate(ctx, req, opts...)
dur := time.Since(start).Milliseconds()
attrs := []attribute.KeyValue{attribute.Int64(AttrLatencyMS, dur)}
attrs = []attribute.KeyValue{attribute.Int64(AttrLatencyMS, dur)}
if info.Attempt > 0 {
attrs = append(attrs, attribute.Int(AttrAttempt, info.Attempt))
}
@@ -360,6 +369,22 @@ func runEventAttributes(e RunEvent) []attribute.KeyValue {
return attrs
}
func appendRunInfoAttributes(attrs []attribute.KeyValue, info ai.RunInfo) []attribute.KeyValue {
if info.Flow != "" {
attrs = append(attrs, attribute.String(AttrFlowName, info.Flow))
}
if info.Step != "" {
attrs = append(attrs, attribute.String(AttrFlowStep, info.Step))
}
if info.Dispatch != "" {
attrs = append(attrs, attribute.String(AttrDispatch, info.Dispatch))
}
if info.Trigger != "" {
attrs = append(attrs, attribute.String(AttrTrigger, info.Trigger))
}
return attrs
}
func (a *agentImpl) recordRunEvent(e RunEvent) {
if e.RunID == "" {
return
+55
View File
@@ -8,6 +8,8 @@ import (
"time"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/flow"
"go-micro.dev/v6/store"
)
func TestAskCancellationAbortsPromptly(t *testing.T) {
@@ -129,3 +131,56 @@ func TestToolCallTimeoutPropagatesDeadlineToCustomTool(t *testing.T) {
t.Fatalf("tool call took %s, want bounded timeout", elapsed)
}
}
func TestAskCheckpointRecordsTerminalOperationalFailureStatus(t *testing.T) {
tests := []struct {
name string
err error
want string
}{
{name: "canceled", err: context.Canceled, want: "canceled"},
{name: "timeout", err: context.DeadlineExceeded, want: "timeout"},
{name: "rate limited", err: testStatusError{code: 429}, want: "rate_limited"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "terminal-"+strings.ReplaceAll(tt.name, " ", "-"))
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
return nil, tt.err
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("terminal-"+strings.ReplaceAll(tt.name, " ", "-")), WithCheckpoint(cp))
_, err := a.Ask(context.Background(), "fail safely")
if err == nil {
t.Fatal("Ask succeeded, want failure")
}
runs, err := cp.List(context.Background())
if err != nil {
t.Fatalf("List: %v", err)
}
if len(runs) != 1 {
t.Fatalf("checkpointed runs = %d, want 1", len(runs))
}
if runs[0].Status != tt.want {
t.Fatalf("run status = %q, want %q", runs[0].Status, tt.want)
}
if len(runs[0].Steps) == 0 || runs[0].Steps[0].Status != tt.want {
t.Fatalf("step status = %#v, want %q", runs[0].Steps, tt.want)
}
if pending, err := Pending(context.Background(), a); err != nil || len(pending) != 0 {
t.Fatalf("Pending = %#v, %v; want no terminal run", pending, err)
}
})
}
}
type testStatusError struct {
code int
}
func (e testStatusError) Error() string { return "provider status error" }
func (e testStatusError) StatusCode() int { return e.code }
+10 -7
View File
@@ -121,13 +121,16 @@ const (
// tell which provider attempt produced the call and whether it is part of a
// retry budget. They are zero when no model-attempt context is known.
type RunInfo struct {
RunID string // correlation id for this agent or flow run
ParentID string // the run that delegated to this one, if any
Agent string // the agent's name
Flow string // the flow's name, when the call is part of a workflow
Step string // the flow step currently executing, when known
Attempt int // current model Generate attempt, starting at 1 when known
MaxAttempts int // configured model Generate attempt budget when known
RunID string // correlation id for this agent or flow run
ParentID string // the run that delegated to this one, if any
Agent string // the agent's name
Flow string // the flow's name, when the call is part of a workflow
Step string // the flow step currently executing, when known
Attempt int // current model Generate attempt, starting at 1 when known
MaxAttempts int // configured model Generate attempt budget when known
VerificationFeedback string // feedback from the previous failed verifier attempt, when retrying a flow step
Dispatch string // how the run was dispatched (direct, broker, schedule, resume) when known
Trigger string // external trigger or schedule label that started the run, when known
}
type runInfoKey struct{}
+104
View File
@@ -7,6 +7,7 @@ import (
"io"
"net/http"
"net/http/httptest"
"os"
"reflect"
"strings"
"testing"
@@ -150,6 +151,109 @@ func TestStreamProvidersCloseCancelsInFlightRequest(t *testing.T) {
}
}
func TestStreamProvidersPropagateProviderErrors(t *testing.T) {
for _, provider := range conformingStreamProviders(t) {
provider := provider
t.Run(provider, func(t *testing.T) {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "upstream quota exhausted", http.StatusTooManyRequests)
}))
defer ts.Close()
stream, err := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL)).Stream(context.Background(), &ai.Request{Prompt: "Hello"})
if err == nil {
_ = stream.Close()
t.Fatal("Stream returned nil error for provider failure")
}
if !strings.Contains(err.Error(), "429") || !strings.Contains(err.Error(), "upstream quota exhausted") {
t.Fatalf("Stream error = %v, want provider status and body", err)
}
if strings.Contains(err.Error(), "test-key") {
t.Fatal("provider error leaked API key")
}
})
}
}
func TestStreamProvidersHonorCanceledContextBeforeRequest(t *testing.T) {
for _, provider := range conformingStreamProviders(t) {
provider := provider
t.Run(provider, func(t *testing.T) {
var sawRequest bool
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
sawRequest = true
http.Error(w, "unexpected request", http.StatusInternalServerError)
}))
defer ts.Close()
ctx, cancel := context.WithCancel(context.Background())
cancel()
stream, err := ai.New(provider, ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL)).Stream(ctx, &ai.Request{Prompt: "Hello"})
if err == nil {
_ = stream.Close()
t.Fatal("Stream returned nil error for canceled context")
}
if !errors.Is(err, context.Canceled) {
t.Fatalf("Stream error = %v, want context.Canceled", err)
}
if sawRequest {
t.Fatal("provider sent request after context was already canceled")
}
})
}
}
func TestConfiguredProviderStreamsSkipWithoutCredentials(t *testing.T) {
for _, tc := range []struct {
provider string
keyEnv string
modelEnv string
}{
{provider: "openai", keyEnv: "OPENAI_API_KEY", modelEnv: "OPENAI_MODEL"},
{provider: "groq", keyEnv: "GROQ_API_KEY", modelEnv: "GROQ_MODEL"},
{provider: "mistral", keyEnv: "MISTRAL_API_KEY", modelEnv: "MISTRAL_MODEL"},
{provider: "together", keyEnv: "TOGETHER_API_KEY", modelEnv: "TOGETHER_MODEL"},
{provider: "atlascloud", keyEnv: "ATLASCLOUD_API_KEY", modelEnv: "ATLASCLOUD_MODEL"},
} {
tc := tc
t.Run(tc.provider, func(t *testing.T) {
key := os.Getenv(tc.keyEnv)
if key == "" {
t.Skipf("%s not set; skipping configured provider stream check", tc.keyEnv)
}
opts := []ai.Option{ai.WithAPIKey(key)}
if model := os.Getenv(tc.modelEnv); model != "" {
opts = append(opts, ai.WithModel(model))
}
stream, err := ai.New(tc.provider, opts...).Stream(context.Background(), &ai.Request{Prompt: "Reply with exactly: ok"})
if err != nil {
t.Fatalf("Stream returned error: %v", err)
}
defer stream.Close()
deadline := time.After(30 * time.Second)
for {
select {
case <-deadline:
t.Fatal("timed out waiting for provider stream chunk")
default:
}
chunk, err := stream.Recv()
if err != nil {
if errors.Is(err, io.EOF) {
t.Fatal("provider stream ended without content")
}
t.Fatalf("Recv returned error: %v", err)
}
if chunk.Reply != "" {
return
}
}
})
}
}
func TestUnsupportedProvidersReturnStreamingUnsupportedAndStayUnregistered(t *testing.T) {
for _, provider := range []string{"anthropic", "gemini"} {
provider := provider
+63
View File
@@ -7,6 +7,7 @@ import (
"go-micro.dev/v6/client"
codecbytes "go-micro.dev/v6/codec/bytes"
"go-micro.dev/v6/store"
)
// fakeClient embeds the default client (so NewRequest works) and
@@ -65,3 +66,65 @@ func TestExecuteDispatchesToAgent(t *testing.T) {
t.Errorf("rendered prompt = %q, want %q", results[0].Prompt, "welcome bob")
}
}
// A caller-owned schedule can trigger an agent workflow without a human chat
// prompt and still leave the normal flow run metadata behind for inspection.
func TestScheduledAgentRunHarnessContract(t *testing.T) {
ctx := context.Background()
cp := StoreCheckpoint(store.NewMemoryStore(), "scheduled-contract")
f := New("scheduled-contract",
Trigger("schedule.daily"),
WithCheckpoint(cp),
Steps(Step{Name: "summarize", Run: Dispatch("ops-agent")}),
)
var parentID string
f.client = &fakeClient{
Client: client.DefaultClient,
callFn: func(req client.Request, rsp interface{}) error {
if req.Service() != "ops-agent" || req.Endpoint() != "Agent.Chat" {
t.Fatalf("dispatched to %s.%s, want ops-agent.Agent.Chat", req.Service(), req.Endpoint())
}
reqFrame := req.Body().(*codecbytes.Frame)
var body map[string]string
if err := json.Unmarshal(reqFrame.Data, &body); err != nil {
t.Fatalf("request body: %v", err)
}
parentID = body["parent_id"]
if body["message"] != "run unattended daily ops review" {
t.Fatalf("message = %q, want scheduled payload", body["message"])
}
frame := rsp.(*codecbytes.Frame)
frame.Data = []byte(`{"reply":"review queued","agent":"ops-agent","parent_id":"` + parentID + `"}`)
return nil
},
}
if err := Scheduled(f, "run unattended daily ops review").Tick(ctx); err != nil {
t.Fatalf("scheduled tick: %v", err)
}
if parentID == "" {
t.Fatal("dispatch did not receive the scheduled flow run id as parent_id")
}
runs, err := cp.List(ctx)
if err != nil {
t.Fatalf("list scheduled runs: %v", err)
}
if len(runs) != 1 {
t.Fatalf("got %d runs, want 1", len(runs))
}
run := runs[0]
if run.ID != parentID {
t.Fatalf("run ID = %q, parent_id = %q", run.ID, parentID)
}
if run.Flow != "scheduled-contract" || run.Status != "done" {
t.Fatalf("run = %+v, want scheduled-contract done", run)
}
if got := run.State.String(); got != "review queued" {
t.Fatalf("run result = %q, want agent reply", got)
}
if len(run.Steps) != 1 || run.Steps[0].Name != "summarize" || run.Steps[0].Status != "done" {
t.Fatalf("steps = %+v, want summarize done", run.Steps)
}
}
+6 -2
View File
@@ -142,7 +142,8 @@ func (f *Flow) Register(reg registry.Registry, br broker.Broker, cl client.Clien
if f.opts.TriggerTopic != "" {
sub, err := br.Subscribe(f.opts.TriggerTopic, func(p broker.Event) error {
data := string(p.Message().Body)
if err := f.Execute(context.Background(), data); err != nil {
ctx := ai.WithRunInfo(context.Background(), ai.RunInfo{Dispatch: "broker", Trigger: f.opts.TriggerTopic})
if err := f.Execute(ctx, data); err != nil {
f.log.Logf(logger.ErrorLevel, "Flow %s failed: %v", f.name, err)
}
return nil
@@ -223,7 +224,10 @@ func (f *Flow) Execute(ctx context.Context, data string) error {
}
runID := uuid.New().String()
ctx = ai.WithRunInfo(ctx, ai.RunInfo{RunID: runID, Flow: f.name})
info, _ := ai.RunInfoFrom(ctx)
info.RunID = runID
info.Flow = f.name
ctx = ai.WithRunInfo(ctx, info)
start := time.Now()
+43 -15
View File
@@ -16,14 +16,18 @@ const (
spanNameFlowRun = "flow.run"
spanNameFlowStep = "flow.step"
AttrFlowRunID = "flow.run.id"
AttrFlowParentID = "flow.run.parent_id"
AttrFlowName = "flow.name"
AttrFlowStepName = "flow.step.name"
AttrFlowStatus = "flow.status"
AttrFlowAttempts = "flow.step.attempts"
AttrFlowLatencyMS = "flow.latency_ms"
AttrFlowErrorKind = "flow.error.kind"
AttrFlowRunID = "flow.run.id"
AttrFlowParentID = "flow.run.parent_id"
AttrFlowName = "flow.name"
AttrFlowStepName = "flow.step.name"
AttrFlowStatus = "flow.status"
AttrFlowAttempts = "flow.step.attempts"
AttrFlowLatencyMS = "flow.latency_ms"
AttrFlowErrorKind = "flow.error.kind"
AttrFlowVerificationStatus = "flow.verification.status"
AttrFlowVerificationNote = "flow.verification.note"
AttrFlowDispatch = "flow.dispatch"
AttrFlowTrigger = "flow.trigger"
)
func (f *Flow) tracer() trace.Tracer {
@@ -34,12 +38,15 @@ func (f *Flow) startRunSpan(ctx context.Context, run Run) (context.Context, func
if f.opts.TraceProvider == nil {
return ctx, func(Run, error) {}
}
ctx, span := f.tracer().Start(ctx, spanNameFlowRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(
info, _ := ai.RunInfoFrom(ctx)
attrs := []attribute.KeyValue{
attribute.String(AttrFlowRunID, run.ID),
attribute.String(AttrFlowParentID, run.ParentID),
attribute.String(AttrFlowName, f.name),
attribute.String(AttrFlowStatus, run.Status),
))
}
attrs = appendRunInfoDispatch(attrs, info)
ctx, span := f.tracer().Start(ctx, spanNameFlowRun, trace.WithSpanKind(trace.SpanKindInternal), trace.WithAttributes(attrs...))
start := time.Now()
return ctx, func(done Run, err error) {
span.SetAttributes(
@@ -57,23 +64,34 @@ func (f *Flow) startRunSpan(ctx context.Context, run Run) (context.Context, func
}
}
func (f *Flow) runStepSpan(ctx context.Context, step Step, in State) (State, int, error) {
func (f *Flow) runStepSpan(ctx context.Context, step Step, in State) (State, int, Verification, error) {
if f.opts.TraceProvider == nil {
return f.runStep(ctx, step, in)
}
info, _ := ai.RunInfoFrom(ctx)
ctx, span := f.tracer().Start(ctx, spanNameFlowStep, trace.WithAttributes(
attrs := []attribute.KeyValue{
attribute.String(AttrFlowRunID, info.RunID),
attribute.String(AttrFlowParentID, info.ParentID),
attribute.String(AttrFlowName, f.name),
attribute.String(AttrFlowStepName, step.Name),
))
}
attrs = appendRunInfoDispatch(attrs, info)
ctx, span := f.tracer().Start(ctx, spanNameFlowStep, trace.WithAttributes(attrs...))
start := time.Now()
out, attempts, err := f.runStep(ctx, step, in)
out, attempts, verification, err := f.runStep(ctx, step, in)
span.SetAttributes(
attribute.Int(AttrFlowAttempts, attempts),
attribute.Int64(AttrFlowLatencyMS, time.Since(start).Milliseconds()),
)
if verification.Passed {
span.SetAttributes(attribute.String(AttrFlowVerificationStatus, "passed"))
}
if verification.Feedback != "" {
span.SetAttributes(attribute.String(AttrFlowVerificationNote, verification.Feedback))
if !verification.Passed {
span.SetAttributes(attribute.String(AttrFlowVerificationStatus, "failed"))
}
}
if err != nil {
span.RecordError(err)
span.SetAttributes(attribute.String(AttrFlowErrorKind, string(ai.ClassifyError(err))))
@@ -82,5 +100,15 @@ func (f *Flow) runStepSpan(ctx context.Context, step Step, in State) (State, int
span.SetStatus(codes.Ok, "")
}
span.End()
return out, attempts, err
return out, attempts, verification, err
}
func appendRunInfoDispatch(attrs []attribute.KeyValue, info ai.RunInfo) []attribute.KeyValue {
if info.Dispatch != "" {
attrs = append(attrs, attribute.String(AttrFlowDispatch, info.Dispatch))
}
if info.Trigger != "" {
attrs = append(attrs, attribute.String(AttrFlowTrigger, info.Trigger))
}
return attrs
}
+26
View File
@@ -74,3 +74,29 @@ func flowSpanAttributes(attrs []attribute.KeyValue) map[string]string {
func withTestRunInfo(ctx context.Context, runID string) context.Context {
return ai.WithRunInfo(ctx, ai.RunInfo{RunID: runID, Agent: "planner"})
}
func TestScheduledFlowOpenTelemetryDispatchAttributes(t *testing.T) {
exp := tracetest.NewInMemoryExporter()
tp := trace.NewTracerProvider(trace.WithSyncer(exp))
step := Step{Name: "summarize", Run: func(ctx context.Context, in State) (State, error) {
in.Data = []byte("queued")
return in, nil
}}
f := New("scheduled-observed", Trigger("schedule.daily"), WithCheckpoint(StoreCheckpoint(store.NewMemoryStore(), "scheduled-observed")), TraceProvider(tp), Steps(step))
if err := Scheduled(f, "daily ops review").Tick(context.Background()); err != nil {
t.Fatal(err)
}
for _, span := range exp.GetSpans().Snapshots() {
if span.Name() != spanNameFlowRun {
continue
}
attrs := flowSpanAttributes(span.Attributes())
if attrs[AttrFlowDispatch] != "schedule" || attrs[AttrFlowTrigger] != "schedule.daily" {
t.Fatalf("scheduled run span dispatch attributes = %#v", attrs)
}
return
}
t.Fatal("flow run span not emitted")
}
+60
View File
@@ -0,0 +1,60 @@
package flow
import (
"context"
"time"
"go-micro.dev/v6/ai"
)
// Schedule binds a flow to a recurring work item without introducing a
// scheduler service. It is a small harness contract: callers own the clock,
// Go Micro owns turning each tick into the same inspectable flow run used for
// broker events and direct Execute calls.
type Schedule struct {
flow *Flow
data string
}
// Scheduled returns a deterministic scheduled-run harness for this flow.
// Tests and event loops can call Tick directly; production processes can wire
// the same contract to time.Ticker through RunEvery. Each tick calls Execute, so
// checkpointed run history, parent/run metadata, cancellation, and inspection
// stay on the normal flow surfaces.
func Scheduled(f *Flow, data string) Schedule {
return Schedule{flow: f, data: data}
}
// Tick starts one scheduled run immediately and returns when that run finishes.
func (s Schedule) Tick(ctx context.Context) error {
if ctx == nil {
ctx = context.Background()
}
info, _ := ai.RunInfoFrom(ctx)
info.Dispatch = "schedule"
if info.Trigger == "" {
info.Trigger = s.flow.opts.TriggerTopic
}
if info.Trigger == "" {
info.Trigger = "schedule"
}
return s.flow.Execute(ai.WithRunInfo(ctx, info), s.data)
}
// RunEvery drives scheduled runs from a ticker until ctx is canceled. It does
// not persist schedule definitions or host a scheduler; it only adapts a caller
// owned cadence to Tick.
func (s Schedule) RunEvery(ctx context.Context, interval time.Duration) error {
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return ctx.Err()
case <-ticker.C:
if err := s.Tick(ctx); err != nil {
return err
}
}
}
}
+83 -22
View File
@@ -52,23 +52,53 @@ func (s State) String() string { return string(s.Data) }
// returns the next state.
type StepFunc func(ctx context.Context, in State) (State, error)
// Step is one unit of a flow — a named action with an optional retry
// override. There is one Step kind; the action is the Run func, and the
// Call/LLM/Agent helpers produce the common ones.
// Verifier grades a step output before the flow advances. Returning
// Passed=false converts the grade into a retryable VerificationError, so
// the existing step retry/supervision path can feed Feedback into the next
// attempt through ai.RunInfo.VerificationFeedback.
type Verifier func(ctx context.Context, out State) (Verification, error)
// Verification is the verifier's deterministic grade for one step attempt.
type Verification struct {
Passed bool
Feedback string
}
// VerificationError reports a failed grade. It is returned from runStep so
// existing retry, checkpoint, and trace paths handle verifier failures the
// same way they handle step execution failures.
type VerificationError struct {
Step string
Feedback string
}
func (e *VerificationError) Error() string {
if e.Feedback == "" {
return fmt.Sprintf("flow: verification failed for step %q", e.Step)
}
return fmt.Sprintf("flow: verification failed for step %q: %s", e.Step, e.Feedback)
}
// Step is one unit of a flow — a named action with optional retry and
// verification hooks. There is one Step kind; the action is the Run func,
// and the Call/LLM/Agent helpers produce the common ones.
type Step struct {
Name string
Run StepFunc
Retry int // per-step override of the flow's retry (0 = use the flow default)
Name string
Run StepFunc
Retry int // per-step override of the flow's retry (0 = use the flow default)
Verify Verifier // optional grade; failed grades retry the step with feedback in RunInfo
}
// StepRecord is the recorded outcome of one step within a run.
type StepRecord struct {
Name string `json:"name"`
Status string `json:"status"` // pending | in_progress | done | failed
Attempts int `json:"attempts"`
Result string `json:"result,omitempty"`
Error string `json:"error,omitempty"`
ErrorKind string `json:"error_kind,omitempty"`
Name string `json:"name"`
Status string `json:"status"` // pending | in_progress | done | failed
Attempts int `json:"attempts"`
Result string `json:"result,omitempty"`
Error string `json:"error,omitempty"`
ErrorKind string `json:"error_kind,omitempty"`
VerificationStatus string `json:"verification_status,omitempty"` // passed | failed
VerificationNote string `json:"verification_note,omitempty"`
}
// Run is the persisted record of one flow execution — what a Checkpoint
@@ -403,7 +433,12 @@ func (f *Flow) Pending(ctx context.Context) ([]Run, error) {
func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
steps := f.opts.Steps
ctx = withDeps(ctx, &runDeps{client: f.client, model: f.model, tools: f.toolSet})
ctx = ai.WithRunInfo(ctx, ai.RunInfo{RunID: run.ID, ParentID: run.ParentID, Agent: f.name, Flow: f.name})
info, _ := ai.RunInfoFrom(ctx)
info.RunID = run.ID
info.ParentID = run.ParentID
info.Agent = f.name
info.Flow = f.name
ctx = ai.WithRunInfo(ctx, info)
ctx, finishSpan := f.startRunSpan(ctx, run)
var spanErr error
defer func() { finishSpan(run, spanErr) }()
@@ -426,8 +461,9 @@ func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
return run, err
}
out, attempts, err := f.runStepSpan(ctx, step, run.State)
out, attempts, verification, err := f.runStepSpan(ctx, step, run.State)
run.Steps[i].Attempts = attempts
applyVerificationRecord(&run.Steps[i], verification)
if err != nil {
spanErr = err
run.Steps[i].Status = "failed"
@@ -476,40 +512,65 @@ func (f *Flow) runFrom(ctx context.Context, run Run) (Run, error) {
// runStep runs one step, retrying on error up to the resolved retry count.
// A step with no Run function is a configuration error, and a canceled run
// stops retrying immediately rather than burning the rest of its budget.
func (f *Flow) runStep(ctx context.Context, step Step, in State) (State, int, error) {
func (f *Flow) runStep(ctx context.Context, step Step, in State) (State, int, Verification, error) {
if step.Run == nil {
return in, 0, fmt.Errorf("flow: step %q has no Run function", step.Name)
return in, 0, Verification{}, fmt.Errorf("flow: step %q has no Run function", step.Name)
}
retries := f.opts.Retry
if step.Retry > 0 {
retries = step.Retry
}
var lastErr error
var lastVerification Verification
var feedback string
for attempt := 1; attempt <= retries+1; attempt++ {
// Stop the moment the run's context is canceled or its deadline
// passes — a canceled run shouldn't keep retrying, and the context
// error is surfaced so callers can detect cancellation upstream.
if err := ctx.Err(); err != nil {
return in, attempt - 1, err
return in, attempt - 1, lastVerification, err
}
attemptCtx := ctx
if info, ok := ai.RunInfoFrom(ctx); ok {
info.Step = step.Name
ctx = ai.WithRunInfo(ctx, info)
info.VerificationFeedback = feedback
attemptCtx = ai.WithRunInfo(ctx, info)
}
out, err := step.Run(attemptCtx, in)
if err == nil && step.Verify != nil {
lastVerification, err = step.Verify(attemptCtx, out)
if err == nil && !lastVerification.Passed {
err = &VerificationError{Step: step.Name, Feedback: lastVerification.Feedback}
}
}
out, err := step.Run(ctx, in)
if err == nil {
return out, attempt, nil
return out, attempt, lastVerification, nil
}
lastErr = err
if verr, ok := err.(*VerificationError); ok {
feedback = verr.Feedback
}
if attempt <= retries && f.opts.RetryBackoff > 0 {
select {
case <-time.After(f.opts.RetryBackoff):
case <-ctx.Done():
return in, attempt, ctx.Err()
return in, attempt, lastVerification, ctx.Err()
}
}
}
return in, retries + 1, lastErr
return in, retries + 1, lastVerification, lastErr
}
func applyVerificationRecord(record *StepRecord, verification Verification) {
if verification.Passed {
record.VerificationStatus = "passed"
}
if verification.Feedback != "" {
record.VerificationNote = truncate(verification.Feedback, 200)
if !verification.Passed {
record.VerificationStatus = "failed"
}
}
}
func (f *Flow) save(ctx context.Context, run Run) error {
+100
View File
@@ -0,0 +1,100 @@
package flow
import (
"context"
"errors"
"testing"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/store"
)
func TestFlowStepVerificationRetriesWithFeedback(t *testing.T) {
var attempts int
var feedback []string
step := Step{
Name: "draft",
Retry: 1,
Run: func(ctx context.Context, in State) (State, error) {
attempts++
info, ok := ai.RunInfoFrom(ctx)
if !ok {
t.Fatal("RunInfo missing from verified step")
}
feedback = append(feedback, info.VerificationFeedback)
if info.VerificationFeedback == "add evidence" {
in.Data = []byte("answer with evidence")
} else {
in.Data = []byte("answer")
}
return in, nil
},
Verify: func(ctx context.Context, out State) (Verification, error) {
if out.String() == "answer with evidence" {
return Verification{Passed: true, Feedback: "meets rubric"}, nil
}
return Verification{Feedback: "add evidence"}, nil
},
}
cp := StoreCheckpoint(store.NewMemoryStore(), "verified")
f := New("verified", WithCheckpoint(cp), Steps(step))
if err := f.Execute(context.Background(), "question"); err != nil {
t.Fatal(err)
}
if attempts != 2 {
t.Fatalf("attempts = %d, want 2", attempts)
}
if len(feedback) != 2 || feedback[0] != "" || feedback[1] != "add evidence" {
t.Fatalf("feedback = %#v, want empty then verifier feedback", feedback)
}
runs, err := cp.List(context.Background())
if err != nil {
t.Fatal(err)
}
if len(runs) != 1 {
t.Fatalf("runs = %d, want 1", len(runs))
}
stepRecord := runs[0].Steps[0]
if stepRecord.Status != "done" || stepRecord.Attempts != 2 || stepRecord.VerificationStatus != "passed" || stepRecord.VerificationNote != "meets rubric" {
t.Fatalf("step record = %#v", stepRecord)
}
}
func TestFlowStepVerificationFailureIsCheckpointed(t *testing.T) {
step := Step{
Name: "grade",
Run: func(ctx context.Context, in State) (State, error) {
in.Data = []byte("bad")
return in, nil
},
Verify: func(ctx context.Context, out State) (Verification, error) {
return Verification{Feedback: "missing citation"}, nil
},
}
cp := StoreCheckpoint(store.NewMemoryStore(), "verified-fail")
f := New("verified-fail", WithCheckpoint(cp), Steps(step))
err := f.Execute(context.Background(), "question")
if err == nil {
t.Fatal("Execute succeeded, want verification failure")
}
var verr *VerificationError
if !errors.As(err, &verr) {
t.Fatalf("error = %T %v, want VerificationError", err, err)
}
if verr.Feedback != "missing citation" {
t.Fatalf("feedback = %q, want missing citation", verr.Feedback)
}
runs, listErr := cp.List(context.Background())
if listErr != nil {
t.Fatal(listErr)
}
if len(runs) != 1 {
t.Fatalf("runs = %d, want 1", len(runs))
}
stepRecord := runs[0].Steps[0]
if runs[0].Status != "failed" || stepRecord.VerificationStatus != "failed" || stepRecord.VerificationNote != "missing citation" {
t.Fatalf("run = %#v step = %#v", runs[0], stepRecord)
}
}
+3 -1
View File
@@ -180,6 +180,8 @@ type Provider struct {
type Capabilities struct {
Streaming bool `json:"streaming"`
PushNotifications bool `json:"pushNotifications"`
TaskResubscribe bool `json:"taskResubscribe"`
InputRequired bool `json:"inputRequired"`
}
// Skill is a capability advertised on the Agent Card.
@@ -349,7 +351,7 @@ func Card(name, url, description string, services []string) AgentCard {
URL: url,
Version: "1.0.0",
ProtocolVersion: protocolVersion,
Capabilities: Capabilities{Streaming: true, PushNotifications: true},
Capabilities: Capabilities{Streaming: true, PushNotifications: true, TaskResubscribe: true, InputRequired: true},
DefaultInputModes: []string{"text/plain"},
DefaultOutputModes: []string{"text/plain"},
Skills: skills,
+3
View File
@@ -91,6 +91,9 @@ func TestAgentCardFromRegistry(t *testing.T) {
if card.ProtocolVersion == "" {
t.Errorf("card missing protocolVersion: %+v", card)
}
if !card.Capabilities.TaskResubscribe || !card.Capabilities.InputRequired {
t.Errorf("card capabilities = %+v, want task resubscribe and input-required advertised", card.Capabilities)
}
if got := skillIDs(card.Skills); strings.Join(got, ",") != "task,project" {
t.Errorf("skill IDs = %v, want [task project]", got)
}
+74 -1
View File
@@ -1,11 +1,13 @@
package a2a
import (
"bufio"
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"strings"
"time"
"github.com/google/uuid"
@@ -135,6 +137,77 @@ func (c *Client) SendMessage(ctx context.Context, message Message) (*Task, error
return &task, nil
}
// Resubscribe reconnects to a retained or active task stream and returns task
// snapshots as the remote agent emits updates. The returned channel is closed
// when the task reaches a terminal state or ctx is canceled.
func (c *Client) Resubscribe(ctx context.Context, taskID string) (<-chan Task, <-chan error) {
tasks := make(chan Task, 8)
errs := make(chan error, 1)
go func() {
defer close(tasks)
defer close(errs)
body, _ := json.Marshal(map[string]any{
"jsonrpc": "2.0",
"id": uuid.New().String(),
"method": "tasks/resubscribe",
"params": getParams{ID: taskID},
})
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.url, bytes.NewReader(body))
if err != nil {
errs <- err
return
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "text/event-stream")
resp, err := c.http.Do(req)
if err != nil {
errs <- err
return
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
errs <- fmt.Errorf("tasks/resubscribe: status %d", resp.StatusCode)
return
}
scanner := bufio.NewScanner(resp.Body)
for scanner.Scan() {
line := scanner.Text()
if !strings.HasPrefix(line, "data:") {
continue
}
payload := strings.TrimSpace(strings.TrimPrefix(line, "data:"))
if payload == "" {
continue
}
var out struct {
Result Task `json:"result"`
Error *rpcError `json:"error"`
}
if err := json.Unmarshal([]byte(payload), &out); err != nil {
errs <- err
return
}
if out.Error != nil {
errs <- fmt.Errorf("a2a tasks/resubscribe: %s (%d)", out.Error.Message, out.Error.Code)
return
}
select {
case <-ctx.Done():
errs <- ctx.Err()
return
case tasks <- out.Result:
}
if terminal(out.Result.Status.State) {
return
}
}
if err := scanner.Err(); err != nil {
errs <- err
}
}()
return tasks, errs
}
// SetPushNotificationConfig asks the remote agent to POST updates for taskID to cfg.URL.
func (c *Client) SetPushNotificationConfig(ctx context.Context, taskID string, cfg PushNotificationConfig) error {
_, err := c.call(ctx, "tasks/pushNotificationConfig/set", pushConfigParams{
@@ -193,7 +266,7 @@ func (c *Client) call(ctx context.Context, method string, params any) (json.RawM
func terminal(state string) bool {
switch state {
case "completed", "failed", "canceled", "rejected":
case "completed", "failed", "canceled", "rejected", "input-required":
return true
}
return false
+56
View File
@@ -3,8 +3,10 @@ package a2a
import (
"context"
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
)
@@ -110,3 +112,57 @@ func TestClientContinuesTaskAndConfiguresPush(t *testing.T) {
t.Fatal("timed out waiting for push update")
}
}
func TestClientResubscribeStreamsRetainedAndLiveTask(t *testing.T) {
d := newDispatcher()
initial := &Task{ID: "task-1", ContextID: "ctx-1", Kind: "task", Status: TaskStatus{State: stateWorking, Timestamp: time.Now().UTC().Format(time.RFC3339)}}
d.store(initial)
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
d.serve(w, r, func(context.Context, string) (string, error) { return "", nil })
}))
defer ts.Close()
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
tasks, errs := NewClient(ts.URL).Resubscribe(ctx, initial.ID)
first := <-tasks
if first.ID != initial.ID || first.Status.State != stateWorking {
t.Fatalf("first resubscribe task = %+v, want retained working task", first)
}
final := &Task{ID: initial.ID, ContextID: initial.ContextID, Kind: "task", Status: TaskStatus{State: stateCompleted, Timestamp: time.Now().UTC().Format(time.RFC3339)}, Artifacts: []Artifact{textArtifact("done")}}
d.store(final)
second := <-tasks
if second.ID != final.ID || second.Status.State != stateCompleted || textOf(second.Artifacts[0].Parts) != "done" {
t.Fatalf("second resubscribe task = %+v, want live completed task", second)
}
if _, ok := <-tasks; ok {
t.Fatal("resubscribe task channel stayed open after terminal update")
}
select {
case err := <-errs:
if err != nil {
t.Fatalf("resubscribe error = %v", err)
}
default:
}
}
func TestClientSendMessageReturnsInputRequiredTask(t *testing.T) {
card := Card("solo", "http://localhost:4000", "", []string{"task"})
h := NewAgentHandler(card, func(context.Context, string) (string, error) {
return "", errors.New("input-required: provide approval code")
})
ts := httptest.NewServer(h)
defer ts.Close()
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
task, err := NewClient(ts.URL).SendMessage(ctx, Message{Parts: []Part{{Kind: "text", Text: "approve?"}}})
if err != nil {
t.Fatalf("SendMessage: %v", err)
}
if task.Status.State != stateInputRequired || !strings.Contains(textOf(task.Artifacts[0].Parts), "provide approval code") {
t.Fatalf("task = %+v, want input-required handoff", task)
}
}
+12
View File
@@ -29,6 +29,18 @@ consistent across providers.
The companion `TestAgentProviderConformanceFakeError` keeps provider error
propagation covered locally without relying on external credentials.
## Local no-secret conformance
Use `make provider-conformance-mock` to run the same provider conformance harness
through the deterministic mock provider. That target requires no API keys and is
what `make harness` delegates to after the 0→1 and 0→hero scenarios, so every PR
continues to exercise the provider-facing agent/tool contract without spending
live model credits.
Use `make provider-conformance` when you want the live-provider sweep: providers
without keys are skipped, and configured providers must satisfy the same harness
contract.
## Scheduled CI
The daily/manual `Harness (E2E)` workflow runs the same matrix with
+3 -3
View File
@@ -21,9 +21,9 @@ changes, architectural rewrites. Those go to the human.
## Work queue (ranked)
1. **Support A2A task resubscribe and input-required handoffs** ([#3474](https://github.com/micro/go-micro/issues/3474)) — the earlier Now/Next hardening seams are closed for this pass: cross-provider dispatch, failure classification, checkpoint/resume trace coverage, RunInfo spans, end-to-end A2A streaming, and memory summarization/retrieval hooks have shipped. The highest-value remaining roadmap gap is live-operation continuity over A2A: remote agents must be able to reconnect to task streams and handle explicit `input-required` pauses so Go Micro agents remain dependable neighbours over open protocols.
2. **Add a CI-verifiable agent verification loop** ([#3481](https://github.com/micro/go-micro/issues/3481)) — blog/32 makes the next strategic frontier explicit: scheduled, looping, work-performing agents need a verification/grader loop around the agent loop. This ranks just behind the open A2A continuity issue because it is an internal operational-harness primitive rather than a named roadmap item, but it is the next cohesive step after RunInfo spans, failure semantics, and memory: grade outputs, route feedback through existing retry/supervision paths, and make dependable agent work observable without drifting into a graph DSL.
1. **Add durable checkpoint/resume for agent runs** ([#3524](https://github.com/micro/go-micro/issues/3524)) — the 0→hero reference app has now shipped, and the next highest-value lifecycle gap is making long-running agent work survive restarts the way flows already do. This is the clearest bridge from services → agents → workflows: flows can already checkpoint deterministic orchestration, but the dynamic agent loop still needs a CI-verifiable resume contract so scheduled, looping agents can be operated rather than merely invoked.
2. **Emit OpenTelemetry spans for agent run timelines** ([#3525](https://github.com/micro/go-micro/issues/3525)) — recent work made runs inspectable and correlated trace metadata through scheduled dispatch; the next step is to turn that RunInfo foundation into standard OTel spans for agent runs, model calls, and tool calls. This keeps `micro runs` useful while making the harness observable in the systems developers already run.
3. **Add retry and timeout resilience to agent tool execution** ([#3526](https://github.com/micro/go-micro/issues/3526)) — flow retry/backoff and cancellation safety have shipped, but the agent loop still needs the same failure semantics around tool/model calls: bounded retries, deadline propagation, cancellation, and visible retry/timeout outcomes. This belongs high because operability is the difference between an agent demo and a dependable service.
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
architecture-review pass._
@@ -0,0 +1,48 @@
package zerotoheroci
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestZeroToHeroReferenceDocs(t *testing.T) {
root := filepath.Clean(filepath.Join("..", "..", ".."))
guide := readFile(t, filepath.Join(root, "internal", "website", "docs", "guides", "zero-to-hero.md"))
for _, want := range []string{
"make harness",
"go test ./cmd/micro/cli/new -run TestZeroToOne -count=1",
"go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1",
"go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1",
"./internal/harness/zero-to-hero-ci/run.sh",
"go run ./internal/harness/agent-flow",
"make provider-conformance-mock",
"internal/harness/plan-delegate",
"internal/harness/universe",
} {
if !strings.Contains(guide, want) {
t.Fatalf("0→hero guide missing %q", want)
}
}
readme := readFile(t, filepath.Join(root, "README.md"))
if !strings.Contains(readme, "internal/website/docs/guides/zero-to-hero.md") {
t.Fatal("README does not point to the canonical 0→hero guide")
}
nav := readFile(t, filepath.Join(root, "internal", "website", "_data", "navigation.yml"))
if !strings.Contains(nav, "0→hero Reference") || !strings.Contains(nav, "/docs/guides/zero-to-hero.html") {
t.Fatal("website navigation does not expose the canonical 0→hero guide")
}
}
func readFile(t *testing.T, name string) string {
t.Helper()
data, err := os.ReadFile(name)
if err != nil {
t.Fatalf("read %s: %v", name, err)
}
return string(data)
}
+3
View File
@@ -32,6 +32,8 @@ examples:
- title: Real-World Examples
url: /docs/examples/realworld/
guides:
- title: 0→hero Reference
url: /docs/guides/zero-to-hero.html
- title: The Agent Harness
url: /docs/guides/agent-harness.html
- title: Agents and Workflows
@@ -70,6 +72,7 @@ project:
- title: Server (optional)
url: /docs/server.html
search_order:
- /docs/guides/zero-to-hero.html
- /docs/getting-started.html
- /docs/mcp.html
- /docs/architecture.html
+4 -4
View File
@@ -135,8 +135,8 @@ The store backend determines durability — file-backed by default, Postgres or
| **What** | Capability | Intelligence | Event orchestration |
| **Does** | Handles requests | Manages services | Reacts to events |
| **Knows** | Its endpoints | Its services' endpoints | Its trigger topic |
| **State** | Store | Store (memory) | Stateless per event |
| **Create** | `micro.NewService()` | `micro.NewAgent()` | `micro.NewFlow()` |
| **State** | Store | Store-backed memory | Checkpointed run history |
| **Create** | `micro.NewService("name")` | `micro.NewAgent("name")` | `micro.NewFlow("name")` |
| **Package** | `service/` | `agent/` | `flow/` |
They compose:
@@ -164,7 +164,7 @@ Or build an agent in Go:
```go
package main
import "go-micro.dev/v5"
import "go-micro.dev/v6"
func main() {
agent := micro.NewAgent("task-mgr",
@@ -176,7 +176,7 @@ func main() {
}
```
The agent package is at `go-micro.dev/v5/agent`. The full interface design is documented in [AGENT_DESIGN.md](https://github.com/micro/go-micro/blob/master/internal/docs/AGENT_DESIGN.md).
The agent implementation lives under `go-micro.dev/v6/agent`; most users create agents through the top-level `go-micro.dev/v6` API. The full interface design is documented in [AGENT_DESIGN.md](https://github.com/micro/go-micro/blob/master/internal/docs/AGENT_DESIGN.md).
---
@@ -31,7 +31,7 @@ your stack — the harness *is* the stack.
| Tools | Every service endpoint is an MCP-callable tool from registry metadata — no extra code | Shipped |
| Memory | Store-backed agent memory (`AgentMemory`), durable across restarts | Shipped |
| Guardrails | `MaxSteps`, `LoopLimit`, `ApproveTool`, tool wrappers — enforced at the call site | Shipped |
| Workflows | Durable flows; `flow.Loop` for run-until-done | Shipped |
| Workflows | Durable flows; `micro.FlowLoop` for run-until-done | Shipped |
| Planning / delegation | Built-in `plan` and `delegate` tools on every agent | Shipped |
| Discovery & RPC | Registry + client; agents and services find and call each other | Shipped |
| Interop | MCP (tools), A2A (agents), x402 (paid tools) | Shipped |
+2 -2
View File
@@ -18,11 +18,11 @@ you're done") has no natural ceiling. So a usable loop needs two things:
1. a **stop condition** — how it decides it's done, and
2. a **hard cap** — a guardrail that guarantees it always terminates.
Go Micro gives you both as a flow step: `flow.Loop`.
Go Micro gives you both as a flow step: `micro.FlowLoop`.
## The shape
`flow.Loop` is a `StepFunc`, so it drops into a flow's ordered, checkpointed
`micro.FlowLoop` is a `StepFunc`, so it drops into a flow's ordered, checkpointed
step list like any other step. It runs a **body** step repeatedly, carrying the
flow `State` from one pass to the next, until a stop condition fires or the
iteration cap is hit — whichever comes first.
@@ -89,14 +89,32 @@ Agents use store-backed conversation memory by default, scoped under the agent's
name. That makes short restarts boring: the next `Ask` reloads the retained
history from the same store backend you already use for services and flows.
Long-running agents can also keep model context bounded without losing useful
prior context:
prior context. If you want retrieval without summaries, enable bounded active
context plus a durable archive of every turn:
```go
a := micro.NewAgent("conductor",
micro.AgentServices("task"),
micro.AgentProvider("anthropic"),
micro.AgentRetrievalMemory(40), // active messages kept in prompt context
micro.AgentMemoryRecallLimit(5), // archived turns recalled per Ask
)
```
`AgentRetrievalMemory(activeLimit)` switches the default memory to a store-backed
retriever. The active conversation is capped at `activeLimit`, every turn is
archived in the same scoped store used by the agent, and future asks inject
matching archived turns ahead of active context. The built-in ranking is
deterministic and credential-free for CI.
When you also want a rolling summary in active context, use compacting memory:
```go
a := micro.NewAgent("conductor",
micro.AgentServices("task"),
micro.AgentProvider("anthropic"),
micro.AgentCompactMemory(40, 12), // max active messages, recent messages kept verbatim
micro.AgentMemoryRecallLimit(5), // archived turns recalled per Ask
micro.AgentMemoryRecallLimit(5), // compacted turns recalled per Ask
)
```
@@ -105,8 +123,7 @@ deterministic compactor. Once active history grows past `maxMessages`, older
turns move into the durable archive, a provider-neutral summary is injected into
active context, and the newest `keepRecent` messages stay verbatim. On future
asks, archived turns whose text matches the current request are recalled ahead of
the active context. The built-in retrieval is intentionally simple and
credential-free for CI; teams that need embeddings or a vector database can still
the active context. Teams that need embeddings or a vector database can still
provide their own `AgentMemory` implementation.
This is harness memory, not prompt-layer orchestration: services remain the
@@ -0,0 +1,84 @@
---
layout: default
---
# 0→hero reference path
The 0→hero path is the maintained, no-secret reference for the Go Micro
services → agents → workflows lifecycle. It ties the CLI inner loop and the
runtime harness together so a contributor can prove the framework still works as
one system, not as separate demos.
Use it when you want to answer: "Can I scaffold a service, run it locally, talk
to an agent, inspect durable work, and reach the deployment boundary without
cloud credentials?"
## What the contract covers
| Boundary | Contract | CI check |
| --- | --- | --- |
| Scaffold | `micro new` generates a runnable service with and without MCP support. | `go test ./cmd/micro/cli/new -run TestZeroToOne -count=1` |
| Run | `micro run` remains the local development entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Chat | `micro chat` remains the interactive agent entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Inspect | `micro inspect agent`, `micro inspect flow`, and `micro flow runs` remain discoverable for run history. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
| Deploy | `micro deploy --dry-run` resolves deploy targets without touching remote infrastructure. | `go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1` |
| Runtime | Real services, agents, durable flows, store-backed history, delegation, and A2A run with only the model mocked. | `./internal/harness/zero-to-hero-ci/run.sh` and `make provider-conformance-mock` |
## Run the whole no-secret path
From the repository root:
```sh
make harness
```
That target runs the scaffold contract, the CLI boundary smoke tests, the
0→hero runtime harnesses, the event-driven agent-flow harness, and mock provider
conformance. It is intentionally deterministic: no provider key, cloud account,
SSH access, or remote service is required.
## Run focused checks while iterating
Use the smaller checks when you are working on one seam:
```sh
# Scaffold → run/call contract.
go test ./cmd/micro/cli/new -run TestZeroToOne -count=1
# CLI inner-loop commands: run, chat, inspect, flow runs, deploy --dry-run.
go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1
go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1
# Durable services → agents → workflows reference scenarios.
./internal/harness/zero-to-hero-ci/run.sh
# Event-as-prompt agent flow.
go run ./internal/harness/agent-flow
# Cross-provider semantics with the deterministic mock provider.
make provider-conformance-mock
```
## Reference scenarios
- [`internal/harness/plan-delegate`](https://github.com/micro/go-micro/tree/master/internal/harness/plan-delegate)
is the compact 0→hero scenario: real task and notify services, a conductor
agent, a comms agent, plan persistence, delegation, and a workflow handoff.
- [`internal/harness/universe`](https://github.com/micro/go-micro/tree/master/internal/harness/universe)
boots a larger mini-world: inventory, payment, order confirmation, a concierge
agent, durable checkpoint/resume, agent run history, flow run history, and A2A
reachability.
- [`internal/harness/agent-flow`](https://github.com/micro/go-micro/tree/master/internal/harness/agent-flow)
shows the event-driven path where a `user.created` event prompts an agent to
call services and complete onboarding.
Together these scenarios keep the North Star executable: services expose typed
capabilities, agents use those capabilities with memory and guardrails, and
workflows compose the work over time.
## Keeping the guide honest
If you change the CLI inner loop, durable flow APIs, agent run history, or the
provider/tool semantics, update this guide and the harness in the same PR. The
point of 0→hero is not a polished sample app that drifts from reality; it is a
CI-verifiable contract that the documented lifecycle still works.
+10
View File
@@ -128,6 +128,12 @@ type ToolFunc = agent.ToolFunc
// NewMemory returns the default store-backed agent memory.
func NewMemory(s store.Store, key string, limit int) Memory { return agent.NewMemory(s, key, limit) }
// NewRetrievalMemory returns store-backed memory with bounded active context
// and durable retrieval over every prior turn.
func NewRetrievalMemory(s store.Store, key string, activeLimit int) Memory {
return agent.NewRetrievalMemory(s, key, activeLimit)
}
// NewCompactingMemory returns store-backed memory with deterministic
// summarization and retrieval controls.
func NewCompactingMemory(s store.Store, key string, maxMessages, keepRecent int) Memory {
@@ -140,6 +146,10 @@ func NewInMemory(limit int) Memory { return agent.NewInMemory(limit) }
// AgentMemory sets the agent's conversation memory (default: store-backed).
func AgentMemory(m Memory) AgentOption { return agent.WithMemory(m) }
// AgentRetrievalMemory enables deterministic default-memory retrieval without
// compaction; activeLimit bounds active context while every turn is archived.
func AgentRetrievalMemory(activeLimit int) AgentOption { return agent.RetrievalMemory(activeLimit) }
// AgentCompactMemory enables deterministic default-memory compaction and
// retrieval for long-running agents.
func AgentCompactMemory(maxMessages, keepRecent int) AgentOption {