Compare commits

..

13 Commits

Author SHA1 Message Date
Codex 749a3ee382 docs: refresh planner priorities for 4248
Harness (E2E) / Harnesses (mock LLM) (push) Waiting to run
Harness (E2E) / Provider harnesses (live LLM conformance) (push) Waiting to run
Lint / golangci-lint (push) Waiting to run
Run Tests / Unit Tests (push) Waiting to run
Run Tests / Etcd Integration Tests (push) Waiting to run
2026-07-07 13:56:07 +00:00
Asim Aslam 5dc2a32233 fix atlascloud minimax tool fallback (#4247)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 14:30:38 +01:00
Asim Aslam d8562edcf6 docs: refresh planner priorities for 4240 (#4242)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 14:02:07 +01:00
Asim Aslam 5a85cba982 docs: add examples wayfinding index (#4239)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 13:34:38 +01:00
Asim Aslam 019420c123 docs: refresh planner priorities for 4233 (#4234)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 12:13:53 +01:00
Asim Aslam 5b32014de1 agent: trace tool retry attempts (#4232)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 10:58:59 +01:00
Asim Aslam 32e522a2a9 docs: refresh planner priorities for 4229 (#4230)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 10:18:29 +01:00
Asim Aslam 811d617b46 docs: refresh coherence changelog (#4228)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 09:15:15 +01:00
Asim Aslam 32d0c46676 agent: cancel stream ask on close (#4226)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 09:05:08 +01:00
Asim Aslam 34e7cf1f5f docs: refresh planner priorities for 4222 (#4224)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 08:26:57 +01:00
Asim Aslam 091fb1d4df agent: add resume pending helper (#4221)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 07:05:27 +01:00
Asim Aslam 1c56d77c86 docs: refresh planner priorities for 4216 (#4219)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 06:32:27 +01:00
Asim Aslam 6f220089d5 test agent retry side-effect dedupe (#4215)
Co-authored-by: Codex <codex@openai.com>
2026-07-07 04:57:01 +01:00
16 changed files with 562 additions and 21 deletions
+3 -2
View File
@@ -21,8 +21,9 @@ changes, architectural rewrites. Those go to the human.
## Work queue (ranked)
1. **Propagate cancellation and retry signals through provider model calls** ([#4175](https://github.com/micro/go-micro/issues/4175)) — the installed first-agent on-ramp shipped in #4207 and AtlasCloud guarded-delegate conformance closed in #4211, so the highest open Now-phase risk is failure handling under real provider conditions: cancellation/deadline propagation and retry/backoff must not duplicate tool side effects. This keeps the services → agents → workflows lifecycle dependable across providers without changing public APIs.
2. **Resume durable agent runs from checkpoints** ([#4202](https://github.com/micro/go-micro/issues/4202)) — once provider failure semantics are guarded, move into the Next-phase durable-agent-loop work: long-running agents should recover like flows, checkpoint progress, and avoid replaying completed tool side effects after interruption.
1. **Prevent duplicate delegated notifications in plan-delegate harness** ([#4245](https://github.com/micro/go-micro/issues/4245)) — Now-phase safety is the current red edge after the AtlasCloud 400 fallback shipped: the services → agents → workflows loop must not repeat delegated side effects while recovering or continuing a plan. Make the plan-delegate path idempotent and CI-verifiable before adding more depth.
2. **Make AtlasCloud guarded delegation pass reliably** ([#4244](https://github.com/micro/go-micro/issues/4244)) — keep cross-provider conformance close behind the side-effect fix: AtlasCloud now gets past the Minimax request-shape 400s, but the live agent harness still misses the required guarded delegate within the retry budget.
3. **Link examples wayfinding from website getting-started path** ([#4241](https://github.com/micro/go-micro/issues/4241)) — keep adoption weighted with hardening after the examples index shipped: the repo README and CLI now point at the first-agent/0→hero map, but go-micro.dev getting-started and quickstart pages should expose the same path with a CI-guarded docs smoke check.
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
architecture-review pass._
+18
View File
@@ -16,17 +16,35 @@ next version when it ships.
## [Unreleased]
### Added
- **StreamAsk close cancellation** — agent streaming calls now cancel promptly when their runner closes, avoiding orphaned stream work. (`agent/`)
- **Agent resume pending helper** — agent durability now has a focused helper for resuming pending checkpointed runs. (`agent/`)
### Fixed
- **Plan/delegate retry idempotency** — agent retries now preserve side-effect and notification dedupe across conformance retry paths, including completion and owner-notification edge cases. (`agent/`, `internal/harness/`)
---
## [6.3.17] - July 2026
### Added
- **First-agent examples CLI wayfinding** — `micro examples` now prints the maintained provider-free first-agent examples in copy/paste order. (`cmd/micro/`)
- **0→hero CLI entrypoint** — `micro zero-to-hero` now points developers at the maintained no-secret services → agents → workflows harness and runnable examples. (`cmd/micro/`)
- **First-agent tutorial smoke harness** — the first-agent tutorial path now has smoke coverage to keep the no-secret on-ramp runnable. (`internal/harness/`)
- **No-secret agent debugging smoke** — the no-secret agent debugging path now has smoke coverage for the first-agent troubleshooting flow. (`internal/harness/`)
- **Durable checkpoint resume smoke coverage** — durable agent resume after checkpointing now has focused smoke coverage. (`agent/`, `internal/harness/`)
### Fixed
- **Plan/delegate notify replays** — duplicate and replayed plan-delegate notifications are now idempotent, so resumed runs do not duplicate completed notifications. (`agent/`, `internal/harness/`)
- **Provider conformance scheduling** — provider conformance workflow dispatches now guard their scheduling path more reliably. (`.github/workflows/`)
- **Plan/delegate notification completion** — delegated notifications now preserve plan completion state more reliably, including duplicate, paraphrased, and delegated-owner notification paths. (`agent/`, `internal/harness/`)
- **AtlasCloud tool fallback** — AtlasCloud built-in tool schemas and follow-up tool fallback handling now recover conformance delegate retries more reliably. (`ai/atlascloud/`, `agent/`)
- **Agent conformance retry completion** — conformance retry prompts and completion handling are more deterministic for delegated agent runs. (`agent/`, `internal/harness/`)
### Documentation
- **First-agent quickstart numbering** — the first-agent on-ramp numbering is consistent across the README and website docs. (`README.md`, `internal/website/docs/`)
- **First-agent inspect command** — docs now use the maintained `micro inspect agent <name>` form. (`README.md`, `internal/website/docs/`)
- **`micro loop` quickstart wayfinding** — docs now surface the loop quickstart from the public docs index and README wayfinding. (`README.md`, `internal/website/docs/`)
---
+6 -5
View File
@@ -94,15 +94,16 @@ walkable agent path in this order:
2. `micro agent demo` — print the provider-free first-agent demo command and next docs steps from the installed CLI.
3. `micro examples` — print the maintained provider-free runnable examples in copy/paste order.
4. `micro zero-to-hero` — print the maintained one-command no-secret lifecycle harness and runnable examples.
5. [Smallest first-agent example](examples/first-agent/) — run one service-backed agent with a mock model and no provider key.
6. [No-secret first-agent transcript](internal/website/docs/guides/no-secret-first-agent.md) — run the
5. [Examples wayfinding index](examples/INDEX.md) — choose the smallest no-secret first-agent, maintained [0→hero support reference](examples/support/), and next interop examples from one map.
6. [Smallest first-agent example](examples/first-agent/) — run one service-backed agent with a mock model and no provider key.
7. [No-secret first-agent transcript](internal/website/docs/guides/no-secret-first-agent.md) — run the
maintained support agent with a mock model and see services → agents → workflows succeed without a key.
7. [Your First Agent](internal/website/docs/guides/your-first-agent.md) — build a
8. [Your First Agent](internal/website/docs/guides/your-first-agent.md) — build a
service-backed agent and talk to it with `micro chat`.
8. [Debugging your agent](internal/website/docs/guides/debugging-agents.md) — use
9. [Debugging your agent](internal/website/docs/guides/debugging-agents.md) — use
`micro inspect agent <name>`, run history, memory, and provider checks when the first
conversation does something unexpected.
9. [0→hero Reference](internal/website/docs/guides/zero-to-hero.md) — complete the
10. [0→hero Reference](internal/website/docs/guides/zero-to-hero.md) — complete the
services → agents → workflows loop with scaffold, run, chat, inspect, flow
history, and deploy dry-run commands that match the maintained harness.
+25
View File
@@ -244,6 +244,31 @@ func Pending(ctx context.Context, ag Agent) ([]flow.Run, error) {
return a.pending(ctx)
}
// ResumePending resumes every checkpointed agent run that has not completed
// yet, in the same oldest-first order returned by Pending.
//
// It is a convenience for service startup and recovery loops: after recreating
// an agent with the same checkpoint store, call ResumePending to drain the
// durable backlog without listing and resuming each run manually. If any run
// fails again, ResumePending stops and returns that run id with the error so
// callers can log, alert, or retry later without hiding the failing run.
func ResumePending(ctx context.Context, ag Agent) (string, error) {
a, ok := ag.(*agentImpl)
if !ok {
return "", fmt.Errorf("agent resume pending: unsupported agent implementation %T", ag)
}
runs, err := a.pending(ctx)
if err != nil {
return "", err
}
for _, run := range runs {
if _, err := a.resume(ctx, run.ID); err != nil {
return run.ID, err
}
}
return "", nil
}
func (a *agentImpl) ask(ctx context.Context, message, parentRunID string) (*Response, error) {
a.mu.Lock()
defer a.mu.Unlock()
+46
View File
@@ -5,6 +5,7 @@ import (
"errors"
"strings"
"testing"
"time"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/client"
@@ -462,6 +463,51 @@ func countMemoryContent(messages []ai.Message, needle string) int {
return count
}
func TestResumePendingResumesOldestAgentRunsUntilFailure(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "resume-pending-agent")
base := time.Date(2026, 7, 7, 12, 0, 0, 0, time.UTC)
for _, run := range []flow.Run{
{ID: "run-ok", Flow: "resume-pending-agent", Status: "failed", State: flow.State{Stage: agentAskStep, Data: []byte("ok")}, Started: base},
{ID: "run-blocked", Flow: "resume-pending-agent", Status: "failed", State: flow.State{Stage: agentAskStep, Data: []byte("block")}, Started: base.Add(time.Minute)},
{ID: "run-later", Flow: "resume-pending-agent", Status: "failed", State: flow.State{Stage: agentAskStep, Data: []byte("later")}, Started: base.Add(2 * time.Minute)},
} {
if err := cp.Save(ctx, run); err != nil {
t.Fatalf("Save(%s): %v", run.ID, err)
}
}
var prompts []string
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
prompts = append(prompts, req.Prompt)
if req.Prompt == "block" {
return nil, errors.New("still blocked")
}
return &ai.Response{Reply: req.Prompt + " resumed"}, nil
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("resume-pending-agent"), WithCheckpoint(cp))
failedRun, err := ResumePending(ctx, a)
if err == nil {
t.Fatal("ResumePending succeeded, want blocked run error")
}
if failedRun != "run-blocked" {
t.Fatalf("failed run = %q, want run-blocked", failedRun)
}
if got, want := strings.Join(prompts, ","), "ok,block"; got != want {
t.Fatalf("prompts = %q, want %q", got, want)
}
loaded, ok, err := cp.Load(ctx, "run-ok")
if err != nil || !ok || loaded.Status != "done" {
t.Fatalf("run-ok loaded=%v err=%v status=%q, want done", ok, err, loaded.Status)
}
loaded, ok, err = cp.Load(ctx, "run-later")
if err != nil || !ok || loaded.Status != "failed" {
t.Fatalf("run-later loaded=%v err=%v status=%q, want still failed", ok, err, loaded.Status)
}
}
func TestPendingReturnsUnfinishedAgentRuns(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "pending-agent")
+26 -4
View File
@@ -36,6 +36,8 @@ const (
AttrTotalTokens = "agent.tokens.total"
AttrAttempt = "agent.model.attempt"
AttrMaxAttempts = "agent.model.max_attempts"
AttrToolAttempt = "agent.tool.attempt"
AttrToolMaxAttempts = "agent.tool.max_attempts"
AttrToolName = "agent.tool.name"
AttrDelegate = "agent.delegate"
AttrGuardrailBlock = "agent.guardrail.block"
@@ -364,7 +366,11 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
res := next(ctx, call)
dur := time.Since(start).Milliseconds()
resErr := resultError(res)
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
toolAttempts := res.Attempts
if toolAttempts <= 0 {
toolAttempts = 1
}
a.recordRunEvent(RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, Attempt: toolAttempts, MaxAttempts: a.opts.ToolMaxAttempts, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
return res
}
@@ -378,6 +384,14 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
res := next(ctx, call)
dur := time.Since(start).Milliseconds()
attrs := []attribute.KeyValue{attribute.Int64(AttrLatencyMS, dur)}
toolAttempts := res.Attempts
if toolAttempts <= 0 {
toolAttempts = 1
}
attrs = append(attrs, attribute.Int(AttrToolAttempt, toolAttempts))
if a.opts.ToolMaxAttempts > 0 {
attrs = append(attrs, attribute.Int(AttrToolMaxAttempts, a.opts.ToolMaxAttempts))
}
if res.Refused != "" {
attrs = append(attrs, attribute.Bool(AttrGuardrailBlock, true), attribute.String(AttrRefusal, res.Refused))
}
@@ -393,7 +407,7 @@ func (a *agentImpl) traceTool(next ai.ToolHandler) ai.ToolHandler {
} else {
span.SetStatus(codes.Ok, "")
}
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
a.recordSpanEvent(span, RunEvent{Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent, Kind: "tool", Name: call.Name, Attempt: toolAttempts, MaxAttempts: a.opts.ToolMaxAttempts, LatencyMS: dur, Refused: res.Refused, Error: resErr, ErrorKind: classifyToolError(resErr)})
span.End()
return res
}
@@ -461,10 +475,18 @@ func runEventAttributes(e RunEvent) []attribute.KeyValue {
attrs = append(attrs, attribute.String(AttrModel, e.Model))
}
if e.Attempt > 0 {
attrs = append(attrs, attribute.Int(AttrAttempt, e.Attempt))
if e.Kind == "tool" {
attrs = append(attrs, attribute.Int(AttrToolAttempt, e.Attempt))
} else {
attrs = append(attrs, attribute.Int(AttrAttempt, e.Attempt))
}
}
if e.MaxAttempts > 0 {
attrs = append(attrs, attribute.Int(AttrMaxAttempts, e.MaxAttempts))
if e.Kind == "tool" {
attrs = append(attrs, attribute.Int(AttrToolMaxAttempts, e.MaxAttempts))
} else {
attrs = append(attrs, attribute.Int(AttrMaxAttempts, e.MaxAttempts))
}
}
if e.LatencyMS > 0 {
attrs = append(attrs, attribute.Int64(AttrLatencyMS, e.LatencyMS))
+76
View File
@@ -142,6 +142,82 @@ func TestAgentOpenTelemetrySpans(t *testing.T) {
}
}
func TestAgentOpenTelemetryToolRetryAttempts(t *testing.T) {
exp := tracetest.NewInMemoryExporter()
tp := trace.NewTracerProvider(trace.WithSyncer(exp))
st := store.NewMemoryStore()
calls := 0
a := New(
Name("tool-retry-otel"),
Provider("oteltest"),
WithStore(st),
TraceProvider(tp),
ToolRetry(3, time.Millisecond),
WithTool("probe", "probe", nil, func(context.Context, map[string]any) (string, error) {
calls++
if calls == 1 {
return "", errors.New("rate limit exceeded")
}
return "ok", nil
}),
)
if _, err := a.Ask(context.Background(), "hello"); err != nil {
t.Fatal(err)
}
if calls != 2 {
t.Fatalf("tool calls = %d, want retry success after 2 attempts", calls)
}
var sawToolSpan bool
for _, span := range exp.GetSpans().Snapshots() {
if span.Name() != spanNameToolCall {
continue
}
attrs := spanAttributes(span.Attributes())
if attrs[AttrToolName] != "probe" {
continue
}
if attrs[AttrToolAttempt] != "2" || attrs[AttrToolMaxAttempts] != "3" {
t.Fatalf("tool retry span attempts = %#v", attrs)
}
if !spanEventHasAttr(span.Events(), "agent.tool", AttrToolAttempt, "2") || !spanEventHasAttr(span.Events(), "agent.tool", AttrToolMaxAttempts, "3") {
t.Fatalf("tool retry event missing attempt attributes: %#v", span.Events())
}
sawToolSpan = true
}
if !sawToolSpan {
t.Fatal("tool retry span not emitted")
}
summaries, err := ListRunSummaries(st, "tool-retry-otel")
if err != nil {
t.Fatal(err)
}
events, err := LoadRunEvents(st, "tool-retry-otel", summaries[0].RunID)
if err != nil {
t.Fatal(err)
}
for _, event := range events {
if event.Kind == "tool" && event.Name == "probe" && event.Attempt == 2 && event.MaxAttempts == 3 {
return
}
}
t.Fatalf("persisted tool event missing retry attempts: %#v", events)
}
func spanEventHasAttr(events []trace.Event, name, key, value string) bool {
for _, event := range events {
if event.Name != name {
continue
}
attrs := spanAttributes(event.Attributes)
if attrs[key] == value {
return true
}
}
return false
}
func TestAgentRunObservabilityRedactsInputByDefault(t *testing.T) {
secret := "deploy production with token sk-secret"
exp := tracetest.NewInMemoryExporter()
+10 -4
View File
@@ -71,14 +71,15 @@ func ResumeStreamAsk(ctx context.Context, ag Agent, runID string) (AgentStream,
// StreamAsk runs tools like Ask, emits ToolStart/ToolEnd events as they execute,
// then emits chunks of the final answer followed by a Done event.
func (a *agentImpl) StreamAsk(ctx context.Context, message string) (AgentStream, error) {
streamCtx, cancel := context.WithCancel(ctx)
events := make(chan *StreamEvent, 16)
done := make(chan struct{})
s := &agentStream{events: events, done: done}
s := &agentStream{events: events, done: done, cancel: cancel}
go func() {
defer close(events)
defer close(done)
resp, err := a.askWithStreamEvents(ctx, message, events)
resp, err := a.askWithStreamEvents(streamCtx, message, events)
if err != nil {
s.setErr(err)
return
@@ -94,14 +95,15 @@ func (a *agentImpl) StreamAsk(ctx context.Context, message string) (AgentStream,
}
func (a *agentImpl) resumeStreamAsk(ctx context.Context, runID string) (AgentStream, error) {
streamCtx, cancel := context.WithCancel(ctx)
events := make(chan *StreamEvent, 16)
done := make(chan struct{})
s := &agentStream{events: events, done: done}
s := &agentStream{events: events, done: done, cancel: cancel}
go func() {
defer close(events)
defer close(done)
resp, err := a.resumeWithStreamEvents(ctx, runID, events)
resp, err := a.resumeWithStreamEvents(streamCtx, runID, events)
if err != nil {
s.setErr(err)
return
@@ -260,6 +262,7 @@ func (a *agentImpl) streamAskAI(ctx context.Context, message string) (ai.Stream,
type agentStream struct {
events <-chan *StreamEvent
done <-chan struct{}
cancel context.CancelFunc
mu sync.Mutex
err error
}
@@ -278,6 +281,9 @@ func (s *agentStream) Recv() (*StreamEvent, error) {
}
func (s *agentStream) Close() error {
if s.cancel != nil {
s.cancel()
}
<-s.done
return nil
}
+34
View File
@@ -5,6 +5,7 @@ import (
"errors"
"io"
"testing"
"time"
"go-micro.dev/v6/ai"
"go-micro.dev/v6/flow"
@@ -75,6 +76,39 @@ func TestStreamAskEmitsToolEventsAndFinalTokens(t *testing.T) {
}
}
func TestStreamAskCloseCancelsInFlightModelCall(t *testing.T) {
started := make(chan struct{})
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
close(started)
<-ctx.Done()
return nil, ctx.Err()
}
defer func() { fakeGen = nil }()
a := newTestAgent(Name("stream-cancel"))
stream, err := a.StreamAsk(context.Background(), "cancel me")
if err != nil {
t.Fatalf("StreamAsk: %v", err)
}
select {
case <-started:
case <-time.After(time.Second):
t.Fatal("model call did not start")
}
closed := make(chan error, 1)
go func() { closed <- stream.Close() }()
select {
case err := <-closed:
if err != nil {
t.Fatalf("Close: %v", err)
}
case <-time.After(time.Second):
t.Fatal("Close did not cancel the in-flight stream")
}
}
func TestStreamAskHelperRejectsUnsupportedAgent(t *testing.T) {
_, err := StreamAsk(context.Background(), unsupportedAgent{}, "hello")
if err == nil {
+76 -3
View File
@@ -24,6 +24,7 @@ import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
@@ -94,6 +95,7 @@ func (p *Provider) String() string { return "atlascloud" }
func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.GenerateOption) (*ai.Response, error) {
tools := atlascloudTools(req.Tools)
compatTools, compatPrompt := atlascloudMinimaxCompatTools(p.opts.Model, req.Tools)
messages := []map[string]any{
{"role": "system", "content": req.SystemPrompt},
@@ -104,6 +106,9 @@ func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.Gen
if req.Prompt != "" {
messages = append(messages, map[string]any{"role": "user", "content": req.Prompt})
}
if compatPrompt != "" {
messages = append(messages, map[string]any{"role": "system", "content": compatPrompt})
}
apiReq := map[string]any{
"model": p.opts.Model,
@@ -119,7 +124,13 @@ func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.Gen
resp, rawMessage, err := p.callAPI(ctx, "chat", apiReq)
if err != nil {
return nil, err
if atlascloudShouldRetryMinimaxCompat(err, compatTools) {
apiReq["tools"] = compatTools
resp, rawMessage, err = p.callAPI(ctx, "chat-minimax-compat", apiReq)
}
if err != nil {
return nil, err
}
}
if len(resp.ToolCalls) == 0 {
@@ -161,7 +172,13 @@ func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.Gen
followUpResp, _, err := p.callAPI(ctx, "tool-follow-up", followUpReq)
if err != nil {
return nil, err
if atlascloudShouldRetryWithoutTools(err, followUpReq) {
delete(followUpReq, "tools")
followUpResp, _, err = p.callAPI(ctx, "tool-follow-up-no-tools", followUpReq)
}
if err != nil {
return nil, err
}
}
if len(followUpResp.ToolCalls) > 0 {
for i := range followUpResp.ToolCalls {
@@ -308,6 +325,18 @@ func (s *atlasStream) Close() error {
return s.body.Close()
}
type atlascloudAPIError struct {
Status string
StatusCode int
Phase string
Summary string
Body string
}
func (e *atlascloudAPIError) Error() string {
return fmt.Sprintf("API error (%s) during atlascloud %s request (%s): %s", e.Status, e.Phase, e.Summary, e.Body)
}
func (p *Provider) callAPI(ctx context.Context, phase string, req map[string]any) (*ai.Response, map[string]any, error) {
reqBody, err := json.Marshal(req)
if err != nil {
@@ -331,7 +360,7 @@ func (p *Provider) callAPI(ctx context.Context, phase string, req map[string]any
respBody, _ := io.ReadAll(httpResp.Body)
if httpResp.StatusCode != http.StatusOK {
return nil, nil, fmt.Errorf("API error (%s) during atlascloud %s request (%s): %s", httpResp.Status, phase, atlascloudRequestSummary(req), string(respBody))
return nil, nil, &atlascloudAPIError{Status: httpResp.Status, StatusCode: httpResp.StatusCode, Phase: phase, Summary: atlascloudRequestSummary(req), Body: string(respBody)}
}
var chatResp struct {
@@ -376,6 +405,50 @@ func (p *Provider) callAPI(ctx context.Context, phase string, req map[string]any
return response, rawMessage, nil
}
func atlascloudMinimaxCompatTools(model string, input []ai.Tool) ([]map[string]any, string) {
if !atlascloudIsMinimaxModel(model) || len(input) == 0 {
return nil, ""
}
var native []ai.Tool
var builtins []string
for _, tool := range input {
switch tool.Name {
case "plan", "request_input", "delegate":
builtins = append(builtins, tool.Name)
default:
native = append(native, tool)
}
}
if len(builtins) == 0 || len(native) == len(input) {
return nil, ""
}
prompt := "AtlasCloud/minimax compatibility: use native tool_calls for the listed service tools. " +
"For built-in agent tools that are not listed natively (" + strings.Join(builtins, ", ") +
"), emit exactly <tool_call name=\"tool_name\">{...}</tool_call> so the agent runtime can execute them. Do not describe those built-in tool calls in prose instead of emitting the tag."
return atlascloudTools(native), prompt
}
func atlascloudIsMinimaxModel(model string) bool {
model = strings.ToLower(model)
return strings.Contains(model, "minimax")
}
func atlascloudShouldRetryMinimaxCompat(err error, compatTools []map[string]any) bool {
if len(compatTools) == 0 {
return false
}
var apiErr *atlascloudAPIError
return errors.As(err, &apiErr) && apiErr.StatusCode == http.StatusBadRequest
}
func atlascloudShouldRetryWithoutTools(err error, req map[string]any) bool {
if _, ok := req["tools"]; !ok {
return false
}
var apiErr *atlascloudAPIError
return errors.As(err, &apiErr) && apiErr.StatusCode == http.StatusBadRequest
}
func atlascloudTools(input []ai.Tool) []map[string]any {
tools := make([]map[string]any, 0, len(input))
for _, t := range input {
+114
View File
@@ -454,6 +454,120 @@ func TestProvider_GeneratePreservesFollowUpTextToolCallInReply(t *testing.T) {
}
}
func TestProvider_GenerateRetriesMinimaxBuiltInsAsTextTools(t *testing.T) {
var bodies []map[string]any
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
t.Fatalf("decode request: %v", err)
}
bodies = append(bodies, body)
w.Header().Set("Content-Type", "application/json")
switch len(bodies) {
case 1:
http.Error(w, `{"code":400,"msg":"bad request"}`, http.StatusBadRequest)
case 2:
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"<tool_call name=\"delegate\">{\"task\":\"summarize\",\"to\":\"blocked-reviewer\"}</tool_call>"}}]}`))
default:
t.Fatalf("unexpected API call %d", len(bodies))
}
}))
defer ts.Close()
p := NewProvider(ai.WithAPIKey("test-key"), ai.WithBaseURL(ts.URL), ai.WithModel("minimaxai/minimax-m3"))
resp, err := p.Generate(context.Background(), &ai.Request{
Prompt: "plan and delegate",
Tools: []ai.Tool{
{Name: "task_TaskService_Add", Description: "add task", Properties: map[string]any{"title": map[string]any{"type": "string"}}},
{Name: "plan", Description: "record a plan", Properties: map[string]any{"steps": map[string]any{"type": "array"}}},
{Name: "request_input", Description: "request input", Properties: map[string]any{"prompt": map[string]any{"type": "string"}}},
{Name: "delegate", Description: "delegate work", Properties: map[string]any{"task": map[string]any{"type": "string"}, "to": map[string]any{"type": "string"}}},
},
})
if err != nil {
t.Fatalf("Generate returned error: %v", err)
}
if !strings.Contains(resp.Reply, `<tool_call name="delegate">`) {
t.Fatalf("Reply = %q, want text delegate fallback", resp.Reply)
}
if len(bodies) != 2 {
t.Fatalf("requests = %d, want initial plus compat retry", len(bodies))
}
initialTools := bodies[0]["tools"].([]any)
if len(initialTools) != 4 {
t.Fatalf("initial tools = %d, want all tools", len(initialTools))
}
retryTools := bodies[1]["tools"].([]any)
if len(retryTools) != 1 {
t.Fatalf("retry tools = %d, want only service tools", len(retryTools))
}
fn := retryTools[0].(map[string]any)["function"].(map[string]any)
if fn["name"] != "task_TaskService_Add" {
t.Fatalf("retry tool name = %v, want service tool only", fn["name"])
}
msgs := bodies[1]["messages"].([]any)
compat := msgs[len(msgs)-1].(map[string]any)
if compat["role"] != "system" || !strings.Contains(compat["content"].(string), `<tool_call name="tool_name">`) {
t.Fatalf("compat instruction = %#v", compat)
}
}
func TestProvider_GenerateFollowUpRetriesWithoutToolsOnBadRequest(t *testing.T) {
var bodies []map[string]any
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
t.Fatalf("decode request: %v", err)
}
bodies = append(bodies, body)
w.Header().Set("Content-Type", "application/json")
switch len(bodies) {
case 1:
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"","tool_calls":[{"id":"call-1","function":{"name":"conformance_echo","arguments":"{\"value\":\"agent-conformance\"}"}}]}}]}`))
case 2:
http.Error(w, `{"code":400,"msg":"bad request"}`, http.StatusBadRequest)
case 3:
if _, ok := body["tools"]; ok {
t.Fatalf("no-tools retry still included tools: %#v", body["tools"])
}
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"done"}}]}`))
default:
t.Fatalf("unexpected API call %d", len(bodies))
}
}))
defer ts.Close()
var toolCalls int
p := NewProvider(
ai.WithAPIKey("test-key"),
ai.WithBaseURL(ts.URL),
ai.WithModel("minimaxai/minimax-m3"),
ai.WithToolHandler(func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
toolCalls++
return ai.ToolResult{ID: call.ID, Content: `{"marker":"agent-conformance-ok"}`}
}),
)
resp, err := p.Generate(context.Background(), &ai.Request{
Prompt: "call a tool",
Tools: []ai.Tool{{Name: "conformance_echo", Description: "echo", Properties: map[string]any{"value": map[string]any{"type": "string"}}}},
})
if err != nil {
t.Fatalf("Generate returned error: %v", err)
}
if resp.Answer != "done" {
t.Fatalf("Answer = %q, want done", resp.Answer)
}
if toolCalls != 1 {
t.Fatalf("tool handler calls = %d, want one (no duplicate side effect)", toolCalls)
}
if len(bodies) != 3 {
t.Fatalf("requests = %d, want chat, failed follow-up, no-tools follow-up", len(bodies))
}
if _, ok := bodies[1]["tools"]; !ok {
t.Fatalf("first follow-up did not include tools")
}
}
func TestProvider_GenerateToolCallHTTPErrorIncludesRequestContext(t *testing.T) {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, `{"code":400,"msg":"bad request"}`, http.StatusBadRequest)
+2 -1
View File
@@ -48,7 +48,8 @@ Guide: https://go-micro.dev/docs/guides/zero-to-hero.html`
const examplesWayfinding = `First-agent examples (no provider key required)
Run these from a go-micro repository checkout in this order:
Run these from a go-micro repository checkout in this order. For the complete
examples map, open examples/INDEX.md:
1. Smallest service-backed agent
go run ./examples/first-agent
+71
View File
@@ -0,0 +1,71 @@
package main
import (
"bytes"
"os"
"path/filepath"
"strings"
"testing"
"github.com/urfave/cli/v2"
microcmd "go-micro.dev/v6/cmd"
)
func TestExamplesWayfindingIndexStaysLinked(t *testing.T) {
root := filepath.Join("..", "..")
files := map[string]string{}
for _, name := range []string{"README.md", "examples/README.md", "examples/INDEX.md"} {
b, err := os.ReadFile(filepath.Join(root, filepath.FromSlash(name)))
if err != nil {
t.Fatalf("read %s: %v", name, err)
}
files[name] = string(b)
}
for _, check := range []struct {
file string
want []string
}{
{
file: "README.md",
want: []string{"examples/INDEX.md", "examples/first-agent/", "examples/support/", "zero-to-hero.md"},
},
{
file: "examples/README.md",
want: []string{"./INDEX.md", "./first-agent/", "./support/", "./mcp/hello/", "./mcp/workflow/"},
},
{
file: "examples/INDEX.md",
want: []string{"go run ./examples/first-agent", "go run ./examples/support", "mcp/hello", "mcp/workflow", "flow-durable", "micro examples"},
},
} {
for _, want := range check.want {
if !strings.Contains(files[check.file], want) {
t.Fatalf("%s missing %q", check.file, want)
}
}
}
}
func TestExamplesCommandPointsAtWayfindingIndex(t *testing.T) {
examples := commandByName(t, "examples")
var out bytes.Buffer
app := cli.NewApp()
app.Writer = &out
if err := examples.Action(cli.NewContext(app, nil, nil)); err != nil {
t.Fatalf("micro examples failed: %v", err)
}
for _, want := range []string{
"examples/INDEX.md",
"go run ./examples/first-agent",
"go run ./examples/support",
"micro zero-to-hero",
} {
if !strings.Contains(out.String(), want) {
t.Fatalf("micro examples output missing %q:\n%s", want, out.String())
}
}
_ = microcmd.DefaultCmd // keep this test coupled to the registered command package.
}
+46
View File
@@ -0,0 +1,46 @@
# Examples wayfinding
Use this index when you want the shortest path from a first runnable agent to the
next services, agents, workflows, and interop examples. Every command below is
provider-free unless the example README says otherwise.
## Pick by goal
| Goal | Start here | Run or verify | Then try |
|------|------------|---------------|----------|
| Run the smallest no-secret agent | [`first-agent`](./first-agent/) | `go run ./examples/first-agent` | [`agent-demo`](./agent-demo/) for a larger service-backed agent |
| Prove the maintained 0→hero path | [`support`](./support/) | `go run ./examples/support` and `go test ./examples/support` | [`zero-to-hero` guide](../internal/website/docs/guides/zero-to-hero.md) |
| See planning and delegation | [`agent-plan-delegate`](./agent-plan-delegate/) | `go run ./examples/agent-plan-delegate` | [`plan-delegate` guide](../internal/website/docs/guides/plan-delegate.md) |
| Expose services through MCP | [`mcp/hello`](./mcp/hello/) | follow [`mcp`](./mcp/) setup | [`mcp/crud`](./mcp/crud/) and [`mcp/workflow`](./mcp/workflow/) |
| Try A2A or gRPC interop next | [`agent-demo`](./agent-demo/) plus gateway docs | run the example, then use the gateway docs | [`grpc-interop`](./grpc-interop/) |
| Add workflow durability | [`flow-durable`](./flow-durable/) | `go run ./examples/flow-durable` | [`flow-loop`](./flow-loop/) |
## Recommended adoption path
1. **First service:** run [`hello-world`](./hello-world/) to learn service
registration, handlers, client calls, and health checks.
2. **First agent:** run [`first-agent`](./first-agent/) with
`go run ./examples/first-agent`; it uses a deterministic mock model and needs
no provider key.
3. **0→hero reference:** run [`support`](./support/) with
`go run ./examples/support`; it keeps typed services, an agent chat loop, an
event-driven flow, and an approval gate in one maintained example.
4. **Interop next:** use [`mcp/hello`](./mcp/hello/), [`mcp/crud`](./mcp/crud/),
and [`mcp/workflow`](./mcp/workflow/) when you are ready to expose tools to
external AI clients.
5. **Workflow depth:** use [`flow-durable`](./flow-durable/) once the agent path
needs checkpointed, resumable deterministic work.
## CLI wayfinding
The installed CLI prints the same path:
```bash
micro examples
micro agent demo
micro zero-to-hero
```
Keep this file, [`README.md`](../README.md), and the `micro examples` output in
sync so new developers can find `examples/first-agent` and `examples/support`
from one documented path.
+2 -2
View File
@@ -7,8 +7,8 @@ coordinate work with workflows.
## Quick Start
Each example can be run with `go run .` from its directory unless its README says
otherwise. If you are new to the repo, follow the first-agent path below instead
of reading the directories alphabetically.
otherwise. If you are new to the repo, start with the [examples wayfinding index](./INDEX.md)
or follow the first-agent path below instead of reading the directories alphabetically.
## Recommended first-agent path
+7
View File
@@ -220,6 +220,13 @@ func AgentResume(ctx context.Context, a Agent, runID string) (*AgentResponse, er
return agent.Resume(ctx, a, runID)
}
// AgentResumePending resumes every incomplete checkpointed agent run, oldest
// first. It returns the first run id that fails again so startup recovery loops
// can leave the durable backlog visible instead of swallowing the failure.
func AgentResumePending(ctx context.Context, a Agent) (string, error) {
return agent.ResumePending(ctx, a)
}
// AgentResumeInput resumes a checkpointed agent run waiting for human input.
func AgentResumeInput(ctx context.Context, a Agent, runID, input string) (*AgentResponse, error) {
return agent.ResumeInput(ctx, a, runID, input)