Compare commits

..

19 Commits

Author SHA1 Message Date
Codex c84ebf272a Refresh planner priorities for 4622
govulncheck / govulncheck (push) Waiting to run
Harness (E2E) / Harnesses (mock LLM) (push) Waiting to run
Harness (E2E) / Provider harnesses (live LLM conformance) (push) Waiting to run
Lint / golangci-lint (push) Waiting to run
Run Tests / Unit Tests (push) Waiting to run
Run Tests / Etcd Integration Tests (push) Waiting to run
2026-07-10 22:33:44 +00:00
Asim Aslam 29d8544ce5 Harden universe A2A reachability probe (#4621)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 22:59:29 +01:00
Asim Aslam c7d510349e Refresh planner priorities for 4617 (#4619)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 22:35:20 +01:00
Asim Aslam 81f81460aa Handle AtlasCloud workspace repair fallback (#4616)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 22:15:52 +01:00
Asim Aslam ba7db2f315 Refresh planner priorities for 4613 (#4614)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 21:36:12 +01:00
Asim Aslam ed3e0e5a06 Handle AtlasCloud empty-arg text tool repair (#4612)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 21:15:35 +01:00
Asim Aslam bd433239d7 docs: refresh planner priorities for 4609 (#4610)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 20:43:17 +01:00
Asim Aslam 7b782589d3 docs: surface first-agent quickcheck (#4608)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 20:15:55 +01:00
Asim Aslam 06a4375e47 docs: refresh planner priorities for 4603 (#4604)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 19:40:51 +01:00
Asim Aslam 1b371470a9 Guard checkpointed tool result recording (#4602)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 19:21:30 +01:00
Asim Aslam 86ef6232bb docs: refresh planner priorities for 4598 (#4600)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 18:38:52 +01:00
Asim Aslam 10a5a5b235 Add first-agent quickcheck breadcrumbs (#4597)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 18:19:38 +01:00
Asim Aslam 28c411f0f7 docs: refresh planner priorities for 4590 (#4592)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 17:50:17 +01:00
Asim Aslam 99a956dec3 Add agent resume breadcrumbs (#4589)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 17:34:29 +01:00
Asim Aslam 84cb4532f5 docs: refresh planner priorities for 4586 (#4587)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 16:59:47 +01:00
Asim Aslam 3d0ea0666e Finalize universe notify after timeout (#4585)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 16:38:42 +01:00
Asim Aslam cddf85c218 docs: refresh planner priorities for 4581 (#4582)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 16:06:17 +01:00
Asim Aslam c4eec47cbc Accept completed plan delegate side effects (#4580)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 15:28:24 +01:00
Asim Aslam 87f011471b docs: refresh planner priorities for 4576 (#4577)
Co-authored-by: Codex <codex@openai.com>
2026-07-10 14:47:25 +01:00
19 changed files with 584 additions and 57 deletions
+1 -4
View File
@@ -21,10 +21,7 @@ changes, architectural rewrites. Those go to the human.
## Work queue (ranked)
1. **Complete AtlasCloud plan-delegate after successful notify timeout** ([#4573](https://github.com/micro/go-micro/issues/4573)) — The newly shipped provider-gated conformance matrix exposed the highest-value live Now failure: the 0→hero-style plan/delegate path can create the expected task and notification side effects, then keep retrying into an approval pause after AtlasCloud 408s. Put this first because it sits directly on the services → agents → workflows adoption story: side effects must finalize idempotently before developers can trust delegated agents under real provider latency.
2. **Finalize AtlasCloud universe notify after agent timeout** ([#4572](https://github.com/micro/go-micro/issues/4572)) — The same live conformance run showed the universe checkout flow can complete payment/order side effects but leave notify/tool-wrapper evidence and durable flow completion ambiguous after an agent-backed notify timeout. This is the next failure/resilience seam because it threatens the durable workflow contract before the A2A probe even runs.
3. **Make AtlasCloud universe A2A reachability probe deterministic** ([#4504](https://github.com/micro/go-micro/issues/4504)) — The universe checkout flow can complete, but the A2A reachability probe still intermittently times out under AtlasCloud. Keep this directly behind the earlier universe finalization defect: once the flow itself is durably complete, the agent must be reachable over the interop gateway without false negatives.
4. **Polish resume and inspect breadcrumbs for agent runs** ([#4569](https://github.com/micro/go-micro/issues/4569)) — Durable runs, streaming, and human-input pauses are now part of the core story, but the developer inner loop still has to make the next command obvious when an agent pauses or needs inspection. This keeps adoption pressure in the queue alongside hardening: chat → inspect → resume should feel like one walkable workflow, not an internal operations exercise.
1. **Add a no-secret first-agent chat transcript check** ([#4618](https://github.com/micro/go-micro/issues/4618)) — Developer adoption is now the highest-value open Now-phase item after #4504 shipped in #4621. README, the website, and the v6.3.15 blog point at the provider-free first-agent path, but the 0→1 agent experience still needs an expected chat/tool-call transcript that a new developer can compare against before adding provider keys. Make the smallest first-agent path more walkable and CI-verifiable.
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
architecture-review pass._
+12 -10
View File
@@ -98,18 +98,20 @@ walkable agent path in this order:
1. [Install troubleshooting](internal/website/docs/guides/install-troubleshooting.md) — verify the binary installer or `go install`, `PATH`, `micro --version`, and the no-secret smoke path before agent work.
2. `micro agent demo` — print the provider-free first-agent demo command and next docs steps from the installed CLI.
3. `micro examples` — print the maintained provider-free runnable examples in copy/paste order.
4. `micro zero-to-hero` — print the maintained one-command no-secret lifecycle harness and runnable examples.
5. [Examples wayfinding index](examples/INDEX.md) — choose the smallest no-secret first-agent, maintained [0→hero support reference](examples/support/), and next interop examples from one map.
6. [Smallest first-agent example](examples/first-agent/) — run one service-backed agent with a mock model and no provider key.
7. [No-secret first-agent transcript](internal/website/docs/guides/no-secret-first-agent.md) — run the
3. `micro agent quickcheck` (or `micro agent debug`) — when scaffold → run → chat → inspect stalls, print the short recovery map before you dive into the full debugging guide.
4. `micro examples` — print the maintained provider-free runnable examples in copy/paste order.
5. `micro zero-to-hero` — print the maintained one-command no-secret lifecycle harness and runnable examples.
6. [Examples wayfinding index](examples/INDEX.md) — choose the smallest no-secret first-agent, maintained [0→hero support reference](examples/support/), and next interop examples from one map.
7. [Smallest first-agent example](examples/first-agent/) — run one service-backed agent with a mock model and no provider key.
8. [No-secret first-agent transcript](internal/website/docs/guides/no-secret-first-agent.md) — run the
maintained support agent with a mock model and see services → agents → workflows succeed without a key.
8. [Your First Agent](internal/website/docs/guides/your-first-agent.md) — build a
9. [Your First Agent](internal/website/docs/guides/your-first-agent.md) — build a
service-backed agent and talk to it with `micro chat`.
9. [Debugging your agent](internal/website/docs/guides/debugging-agents.md) — use
`micro inspect agent <name>`, run history, memory, and provider checks when the first
conversation does something unexpected.
10. [0→hero Reference](internal/website/docs/guides/zero-to-hero.md) — complete the
10. [Debugging your agent](internal/website/docs/guides/debugging-agents.md) — use
`micro agent preflight` before `micro run`, `micro agent doctor` after `micro run`,
then `micro chat` and `micro inspect agent <name>` to recover run history, memory,
and provider checks when the first conversation does something unexpected.
11. [0→hero Reference](internal/website/docs/guides/zero-to-hero.md) — complete the
services → agents → workflows loop with scaffold, run, chat, inspect, flow
history, and deploy dry-run commands that match the maintained harness.
+16 -12
View File
@@ -235,28 +235,32 @@ func agentOperationalError(err error) error {
func (a *agentImpl) checkpointToolWrap(next ai.ToolHandler) ai.ToolHandler {
return func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
if a.opts.Checkpoint == nil || a.currentRun == nil {
run := a.currentRun
if a.opts.Checkpoint == nil || run == nil {
return next(ctx, call)
}
name := toolCheckpointName(call)
if rec, ok := findStep(a.currentRun.Steps, name); ok && rec.Status == "done" {
if rec, ok := findStep(run.Steps, name); ok && rec.Status == "done" {
return ai.ToolResult{ID: call.ID, Value: rec.Result, Content: rec.Result}
}
idx := upsertStep(&a.currentRun.Steps, flow.StepRecord{Name: name, Status: "in_progress"})
_ = a.saveRun(ctx, *a.currentRun)
idx := upsertStep(&run.Steps, flow.StepRecord{Name: name, Status: "in_progress"})
_ = a.saveRun(ctx, *run)
res := next(ctx, call)
a.currentRun.Steps[idx].Attempts++
if idx < 0 || idx >= len(run.Steps) || run.Steps[idx].Name != name {
idx = upsertStep(&run.Steps, flow.StepRecord{Name: name, Status: "in_progress"})
}
run.Steps[idx].Attempts++
if res.Refused != "" {
a.currentRun.Steps[idx].Status = "failed"
a.currentRun.Steps[idx].Error = res.Content
_ = a.saveRun(ctx, *a.currentRun)
run.Steps[idx].Status = "failed"
run.Steps[idx].Error = res.Content
_ = a.saveRun(ctx, *run)
return res
}
a.currentRun.Steps[idx].Status = "done"
a.currentRun.Steps[idx].Result = res.Content
a.currentRun.Steps[idx].Error = ""
_ = a.saveRun(ctx, *a.currentRun)
run.Steps[idx].Status = "done"
run.Steps[idx].Result = res.Content
run.Steps[idx].Error = ""
_ = a.saveRun(ctx, *run)
return res
}
}
+39
View File
@@ -154,6 +154,45 @@ func TestCheckpointSkipsDuplicateToolWithinAsk(t *testing.T) {
}
}
func TestCheckpointToolWrapSurvivesClearedCurrentRun(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "tool-cleared-run-agent")
run := flow.Run{
ID: "run-1",
Flow: "tool-cleared-run-agent",
Status: "running",
Steps: []flow.StepRecord{{Name: agentAskStep, Status: "in_progress"}},
}
a := &agentImpl{
opts: newOptions(Name("tool-cleared-run-agent"), WithCheckpoint(cp)),
currentRun: &run,
}
handler := a.checkpointToolWrap(func(context.Context, ai.ToolCall) ai.ToolResult {
a.currentRun = nil
return ai.ToolResult{ID: "call-1", Content: "created"}
})
res := handler(ctx, ai.ToolCall{ID: "call-1", Name: "external.create", Input: map[string]any{"title": "Design"}})
if res.Content != "created" {
t.Fatalf("tool result = %q, want created", res.Content)
}
loaded, ok, err := cp.Load(ctx, "run-1")
if err != nil {
t.Fatalf("load checkpoint: %v", err)
}
if !ok {
t.Fatal("checkpoint missing")
}
rec, ok := findStep(loaded.Steps, `tool:external.create:{"title":"Design"}`)
if !ok {
t.Fatalf("checkpoint steps = %#v, want completed tool step", loaded.Steps)
}
if rec.Status != "done" || rec.Result != "created" || rec.Attempts != 1 {
t.Fatalf("tool checkpoint = %#v, want done result with one attempt", rec)
}
}
func TestCheckpointContinuesRunWithUnfinishedPlanStep(t *testing.T) {
ctx := context.Background()
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "unfinished-plan-agent")
+96
View File
@@ -554,8 +554,104 @@ func atlascloudFallbackTextToolCall(toolName string, req *ai.Request) string {
case "delegate":
return atlascloudDelegateFallbackTextToolCall(req)
default:
if atlascloudToolTakesNoArguments(toolName, req.Tools) {
return atlascloudEmptyArgumentFallbackTextToolCall(toolName)
}
return atlascloudServiceFallbackTextToolCall(toolName, req)
}
}
func atlascloudServiceFallbackTextToolCall(toolName string, req *ai.Request) string {
if req == nil {
return ""
}
for _, tool := range req.Tools {
if tool.Name != toolName {
continue
}
args := atlascloudFallbackArgsForProperties(tool.Properties, atlascloudRequestText(req))
if args == nil {
return ""
}
b, err := json.Marshal(args)
if err != nil {
return ""
}
return `<tool_call name="` + toolName + `">` + string(b) + `</tool_call>`
}
return ""
}
func atlascloudFallbackArgsForProperties(properties map[string]any, ctxText string) map[string]any {
if len(properties) == 0 {
return map[string]any{}
}
args := make(map[string]any, len(properties))
for name, schema := range properties {
value, ok := atlascloudFallbackArgValue(name, schema, ctxText)
if !ok {
return nil
}
args[name] = value
}
return args
}
func atlascloudFallbackArgValue(name string, schema any, ctxText string) (any, bool) {
typeName := "string"
if m, ok := schema.(map[string]any); ok {
if t, _ := m["type"].(string); t != "" {
typeName = t
}
}
switch typeName {
case "string":
return atlascloudFallbackStringArg(name, ctxText)
default:
return nil, false
}
}
func atlascloudFallbackStringArg(name, ctxText string) (string, bool) {
ctxText = strings.TrimSpace(ctxText)
if ctxText == "" {
return "", false
}
if strings.Contains(strings.ToLower(name), "email") || strings.Contains(strings.ToLower(name), "owner") {
if email := atlascloudFirstEmail(ctxText); email != "" {
return email, true
}
}
return ctxText, true
}
func atlascloudFirstEmail(text string) string {
for _, field := range strings.FieldsFunc(text, func(r rune) bool {
return strings.ContainsRune(" \t\n\r<>\"'(),;", r)
}) {
field = strings.Trim(field, ".:")
if strings.Contains(field, "@") && strings.Contains(field, ".") {
return field
}
}
return ""
}
func atlascloudToolTakesNoArguments(toolName string, tools []ai.Tool) bool {
for _, tool := range tools {
if tool.Name != toolName {
continue
}
return len(tool.Properties) == 0
}
return false
}
func atlascloudEmptyArgumentFallbackTextToolCall(toolName string) string {
if toolName == "" {
return ""
}
return `<tool_call name="` + toolName + `">{}</tool_call>`
}
func atlascloudPlanFallbackTextToolCall(prompt string) string {
+70
View File
@@ -628,6 +628,76 @@ func TestProvider_GenerateFallsBackAfterRepeatedPartialDelegateTextToolCall(t *t
}
}
func TestProvider_GenerateFallsBackAfterRepeatedPartialNoArgumentServiceTextToolCall(t *testing.T) {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"<tool_call name=\"task_TaskService_List\">"}}]}`))
}))
defer ts.Close()
p := NewProvider(
ai.WithAPIKey("test-key"),
ai.WithBaseURL(ts.URL),
ai.WithModel("minimaxai/minimax-m3"),
)
resp, err := p.Generate(context.Background(), &ai.Request{
Prompt: "list the current launch-readiness tasks",
Tools: []ai.Tool{{
Name: "task_TaskService_List",
OriginalName: "task.TaskService.List",
Description: "List persisted launch-readiness tasks",
Properties: map[string]any{},
}},
})
if err != nil {
t.Fatalf("Generate returned error: %v", err)
}
want := `<tool_call name="task_TaskService_List">{}</tool_call>`
if resp.Reply != want {
t.Fatalf("Reply = %q, want %q", resp.Reply, want)
}
}
func TestProvider_GenerateFallsBackAfterRepeatedPartialWorkspaceServiceTextToolCall(t *testing.T) {
var bodies []map[string]any
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
t.Fatalf("decode request: %v", err)
}
bodies = append(bodies, body)
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"<tool_call name=\"workspace_WorkspaceService_Create\">"}}]}`))
}))
defer ts.Close()
p := NewProvider(
ai.WithAPIKey("test-key"),
ai.WithBaseURL(ts.URL),
ai.WithModel("minimaxai/minimax-m3"),
)
resp, err := p.Generate(context.Background(), &ai.Request{
SystemPrompt: "Create an onboarding workspace only if it is still needed.",
Prompt: "Onboard alice@acme.com. The workspace create side effect may already be complete; avoid failing the flow on a duplicate repaired call.",
Tools: []ai.Tool{{
Name: "workspace_WorkspaceService_Create",
OriginalName: "workspace.WorkspaceService.Create",
Description: "Create an onboarding workspace",
Properties: map[string]any{"owner": map[string]any{"type": "string"}},
}},
})
if err != nil {
t.Fatalf("Generate returned error: %v", err)
}
want := `<tool_call name="workspace_WorkspaceService_Create">{"owner":"alice@acme.com"}</tool_call>`
if resp.Reply != want {
t.Fatalf("Reply = %q, want %q", resp.Reply, want)
}
if len(bodies) != 2 {
t.Fatalf("requests = %d, want initial plus repair", len(bodies))
}
}
func TestProvider_GenerateRetriesMinimaxBuiltInsAsTextTools(t *testing.T) {
var bodies []map[string]any
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+68
View File
@@ -14,6 +14,32 @@ import (
"go-micro.dev/v6/store"
)
const firstAgentQuickChecksHelp = `First-agent failure-mode quick checks
Use this when scaffold -> run -> chat -> inspect stalls and you want the
smallest provider-free recovery loop before reading the full docs.
1. Confirm prerequisites before starting the gateway:
micro agent preflight
2. Start the project and keep it running in a separate terminal:
micro run
3. Check the agent is registered and the chat gateway is reachable:
micro agent doctor
4. If chat returns an answer or an error, inspect the latest run state:
micro inspect agent <name>
micro runs <name>
5. If provider chat is not configured yet, prove the no-secret path still works:
micro agent demo
go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentTranscript -count=1
Recovery docs:
https://go-micro.dev/docs/guides/debugging-agents.html
https://go-micro.dev/docs/guides/no-secret-first-agent.html`
const noSecretDemoHelp = `No-secret first-agent demo
Use this when you want the fastest provider-free agent success path before
@@ -72,6 +98,18 @@ for live-provider chat and inspect/debugging.`,
return nil
},
},
{
Name: "quickcheck",
Aliases: []string{"debug"},
Usage: "Print first-agent failure-mode quick checks",
Description: `Print provider-free recovery breadcrumbs for the scaffold -> run ->
chat -> inspect loop, including exact commands for registration, gateway, run
history, and no-secret fallback checks.`,
Action: func(c *cli.Context) error {
fmt.Fprintln(c.App.Writer, firstAgentQuickChecksHelp)
return nil
},
},
{
Name: "preflight",
Usage: "Check local prerequisites before the first provider-backed agent",
@@ -224,14 +262,44 @@ func writeRunIndex(w io.Writer, name string, runs []goagent.RunSummary, asJSON b
if run.TraceID != "" {
line += " trace=" + shortTraceID(run.TraceID)
}
if run.Checkpoint != "" {
line += " checkpoint=" + run.Checkpoint
}
if run.Stage != "" {
line += " stage=" + run.Stage
}
if run.LastError != "" {
line += " error=" + run.LastError
}
fmt.Fprintln(w, line)
writeRunIndexBreadcrumbs(w, name, run)
}
return nil
}
func writeRunIndexBreadcrumbs(w io.Writer, name string, run goagent.RunSummary) {
if run.Stage == "input-required" {
fmt.Fprintf(w, " inspect: micro agent history %s %s\n", name, run.RunID)
fmt.Fprintf(w, " input: call micro.AgentResumeInput(ctx, agent, %q, input) to continue the input-required run\n", run.RunID)
return
}
if !isResumableRunSummary(run) {
return
}
fmt.Fprintf(w, " inspect: micro agent history %s %s\n", name, run.RunID)
fmt.Fprintf(w, " resume: call micro.AgentResume(ctx, agent, %q) after recreating the agent with the same checkpoint store\n", run.RunID)
fmt.Fprintf(w, " stream: call micro.ResumeStreamAsk(ctx, agent, %q) to resume with streaming events\n", run.RunID)
}
func isResumableRunSummary(run goagent.RunSummary) bool {
switch run.Status {
case "running", "error", "failed", "refused":
return run.Checkpoint != "done" || run.Stage != ""
default:
return false
}
}
func printRunHistory(name, runID string, asJSON bool) error {
events, err := goagent.LoadRunEvents(store.DefaultStore, name, runID)
if err != nil {
+40
View File
@@ -59,6 +59,46 @@ func TestWriteRunIndexHumanIncludesStatusAndDuration(t *testing.T) {
}
}
func TestWriteRunIndexIncludesResumeBreadcrumbs(t *testing.T) {
runs := []goagent.RunSummary{{
RunID: "run-failed",
Agent: "runner",
UpdatedAt: time.Date(2026, 6, 25, 12, 34, 56, 0, time.UTC),
Events: 3,
Status: "error",
LastKind: "tool",
Checkpoint: "failed",
Stage: "ask",
}}
var out bytes.Buffer
if err := writeRunIndex(&out, "runner", runs, false); err != nil {
t.Fatal(err)
}
got := out.String()
for _, want := range []string{"checkpoint=failed", "stage=ask", `micro agent history runner run-failed`, `micro.AgentResume(ctx, agent, "run-failed")`, `micro.ResumeStreamAsk(ctx, agent, "run-failed")`} {
if !strings.Contains(got, want) {
t.Fatalf("output missing %q:\n%s", want, got)
}
}
}
func TestWriteRunIndexInputRequiredUsesResumeInput(t *testing.T) {
runs := []goagent.RunSummary{{RunID: "run-input", Agent: "runner", Status: "running", LastKind: "checkpoint", Checkpoint: "paused", Stage: "input-required"}}
var out bytes.Buffer
if err := writeRunIndex(&out, "runner", runs, false); err != nil {
t.Fatal(err)
}
got := out.String()
for _, want := range []string{`micro agent history runner run-input`, `micro.AgentResumeInput(ctx, agent, "run-input", input)`} {
if !strings.Contains(got, want) {
t.Fatalf("output missing %q:\n%s", want, got)
}
}
if strings.Contains(got, `micro.AgentResume(ctx, agent, "run-input")`) || strings.Contains(got, "ResumeStreamAsk") {
t.Fatalf("input-required run should point at ResumeInput only, got:\n%s", got)
}
}
func TestWriteRunHistoryHumanAndJSON(t *testing.T) {
events := []goagent.RunEvent{{
Time: time.Date(2026, 6, 25, 12, 34, 56, 7_000_000, time.UTC),
+21
View File
@@ -73,3 +73,24 @@ func TestRunAgentDoctorReportsActionableRecoveryFailures(t *testing.T) {
}
}
}
func TestAgentQuickcheckPrintsProviderFreeFailureModeBreadcrumbs(t *testing.T) {
got := firstAgentQuickChecksHelp
for _, want := range []string{
"First-agent failure-mode quick checks",
"scaffold -> run -> chat -> inspect",
"micro agent preflight",
"micro run",
"micro agent doctor",
"micro inspect agent <name>",
"micro runs <name>",
"micro agent demo",
"go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentTranscript -count=1",
"debugging-agents.html",
"no-secret-first-agent.html",
} {
if !strings.Contains(got, want) {
t.Fatalf("quickcheck output missing %q:\n%s", want, got)
}
}
}
+3
View File
@@ -82,6 +82,9 @@ const docsWayfinding = `First-agent and 0→hero docs:
prove service tools, mock-model chat, and inspectable run history without
configuring a provider key.
If scaffold → run → chat → inspect stalls, print the short recovery map:
micro agent quickcheck
2. No-secret first-agent transcript
https://go-micro.dev/docs/guides/no-secret-first-agent.html
Run the maintained support agent without a provider key:
+34
View File
@@ -38,6 +38,9 @@ func TestFirstAgentWalkthroughCLIBoundaries(t *testing.T) {
if !subcommands["agent"]["doctor"] {
t.Fatal("first-agent walkthrough missing recovery boundary: agent doctor")
}
if !subcommands["agent"]["quickcheck"] {
t.Fatal("first-agent walkthrough missing failure-mode boundary: agent quickcheck")
}
if !subcommands["inspect"]["agent"] {
t.Fatal("first-agent walkthrough missing inspect boundary: inspect agent")
}
@@ -118,6 +121,28 @@ func TestFirstAgentWalkthroughCLIBoundaries(t *testing.T) {
}
}
quickcheck := subcommandByName(t, agent, "quickcheck")
out.Reset()
if err := quickcheck.Action(cli.NewContext(app, nil, nil)); err != nil {
t.Fatalf("micro agent quickcheck failed: %v", err)
}
for _, want := range []string{
"First-agent failure-mode quick checks",
"scaffold -> run -> chat -> inspect",
"micro agent preflight",
"micro run",
"micro agent doctor",
"micro inspect agent <name>",
"micro runs <name>",
"micro agent demo",
"go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentTranscript -count=1",
"debugging-agents.html",
} {
if !strings.Contains(out.String(), want) {
t.Fatalf("micro agent quickcheck output missing %q:\n%s", want, out.String())
}
}
demo := subcommandByName(t, agent, "demo")
out.Reset()
if err := demo.Action(cli.NewContext(app, nil, nil)); err != nil {
@@ -150,6 +175,7 @@ func TestFirstAgentDocsMatchCLIOutput(t *testing.T) {
}
agent := commandByName(t, "agent")
outputs["micro agent demo"] = commandOutput(t, subcommandByName(t, agent, "demo"))
outputs["micro agent quickcheck"] = commandOutput(t, subcommandByName(t, agent, "quickcheck"))
contracts := []struct {
name string
@@ -161,6 +187,10 @@ func TestFirstAgentDocsMatchCLIOutput(t *testing.T) {
file: filepath.Join(root, "README.md"),
markers: []string{
"micro agent demo",
"micro agent quickcheck",
"micro agent preflight",
"micro agent doctor",
"micro inspect agent <name>",
"micro examples",
"micro zero-to-hero",
"examples/first-agent/",
@@ -176,6 +206,10 @@ func TestFirstAgentDocsMatchCLIOutput(t *testing.T) {
file: filepath.Join(root, "internal", "website", "docs", "getting-started.md"),
markers: []string{
"micro agent demo",
"micro agent quickcheck",
"micro agent preflight",
"micro agent doctor",
"micro inspect agent <name>",
"micro examples",
"micro zero-to-hero",
"github.com/micro/go-micro/tree/master/examples/first-agent",
+15 -6
View File
@@ -98,16 +98,25 @@ func writeAgentInspection(w io.Writer, name string, runs []goagent.RunSummary, a
fmt.Fprintf(w, " trace=%s", shortID(run.TraceID))
}
fmt.Fprintln(w)
if isResumableAgentRun(run) {
fmt.Fprintf(w, " resume: call micro.AgentResume(ctx, agent, %q) after recreating the agent with the same checkpoint store\n", run.RunID)
}
if run.Stage == "input-required" {
fmt.Fprintf(w, " input: call micro.AgentResumeInput(ctx, agent, %q, input) to continue the paused run\n", run.RunID)
}
writeAgentRunBreadcrumbs(w, name, run)
}
return nil
}
func writeAgentRunBreadcrumbs(w io.Writer, name string, run goagent.RunSummary) {
if run.Stage == "input-required" {
fmt.Fprintf(w, " inspect: micro agent history %s %s\n", name, run.RunID)
fmt.Fprintf(w, " input: call micro.AgentResumeInput(ctx, agent, %q, input) to continue the input-required run\n", run.RunID)
return
}
if !isResumableAgentRun(run) {
return
}
fmt.Fprintf(w, " inspect: micro agent history %s %s\n", name, run.RunID)
fmt.Fprintf(w, " resume: call micro.AgentResume(ctx, agent, %q) after recreating the agent with the same checkpoint store\n", run.RunID)
fmt.Fprintf(w, " stream: call micro.ResumeStreamAsk(ctx, agent, %q) to resume with streaming events\n", run.RunID)
}
func isResumableAgentRun(run goagent.RunSummary) bool {
switch run.Status {
case "running", "error", "failed", "refused":
+5 -2
View File
@@ -17,7 +17,7 @@ func TestWriteAgentInspectionIncludesActionableBreadcrumbs(t *testing.T) {
t.Fatal(err)
}
got := out.String()
for _, want := range []string{"Agent \"support\" runs", "run-1", "status=error", "events=4", "last=tool", "checkpoint=failed", "stage=ask", `error="boom"`, "trace=1234567890ab", `micro.AgentResume(ctx, agent, "run-1")`} {
for _, want := range []string{"Agent \"support\" runs", "run-1", "status=error", "events=4", "last=tool", "checkpoint=failed", "stage=ask", `error="boom"`, "trace=1234567890ab", `micro agent history support run-1`, `micro.AgentResume(ctx, agent, "run-1")`, `micro.ResumeStreamAsk(ctx, agent, "run-1")`} {
if !strings.Contains(got, want) {
t.Fatalf("output missing %q:\n%s", want, got)
}
@@ -31,11 +31,14 @@ func TestWriteAgentInspectionIncludesInputResumeBreadcrumb(t *testing.T) {
t.Fatal(err)
}
got := out.String()
for _, want := range []string{"checkpoint=paused", "stage=input-required", `micro.AgentResumeInput(ctx, agent, "run-input", input)`} {
for _, want := range []string{"checkpoint=paused", "stage=input-required", `micro agent history support run-input`, `micro.AgentResumeInput(ctx, agent, "run-input", input)`} {
if !strings.Contains(got, want) {
t.Fatalf("output missing %q:\n%s", want, got)
}
}
if strings.Contains(got, `micro.AgentResume(ctx, agent, "run-input")`) || strings.Contains(got, "ResumeStreamAsk") {
t.Fatalf("input-required run should point at ResumeInput only, got:\n%s", got)
}
}
func TestWriteAgentInspectionEmptyStateNamesInspectCommand(t *testing.T) {
+8 -4
View File
@@ -630,11 +630,11 @@ func waitForPlanDelegateExecution(done <-chan error, taskSvc *TaskService, notif
tasks := taskSvc.count()
notify := notifySvc.count()
if err != nil {
if hasCompletedPlanDelegateSideEffects(tasks, notify) {
fmt.Printf("\n\033[33mwarning:\033[0m flow execute returned after completed side effects: %v\n", err)
return nil
}
if isClientTimeout(err) {
if tasks > 0 && notify == 1 {
fmt.Printf("\n\033[33mwarning:\033[0m flow execute returned after completed side effects: %v\n", err)
return nil
}
return classifiedPlanDelegateTimeout(tasks, notify, err)
}
if isUnfinishedPlanError(err) && tasks > 0 && notify == 0 && recoverMissingNotify != nil {
@@ -678,6 +678,10 @@ func waitForNotifySideEffect(notifySvc *NotifyService, timeout time.Duration) (b
}
}
func hasCompletedPlanDelegateSideEffects(tasks, notify int) bool {
return tasks == 3 && notify == 1
}
func classifiedPlanDelegateTimeout(tasks, notify int, err error) error {
return fmt.Errorf("provider latency/outage during plan-delegate before required side effects completed (tasks=%d/3 notify=%d/1); retry live provider or inspect provider logs if this recurs: %w", tasks, notify, err)
}
@@ -600,6 +600,28 @@ func TestPlanDelegateExecutionAcceptsClientTimeoutAfterSideEffects(t *testing.T)
}
}
func TestPlanDelegateExecutionAcceptsApprovalPauseAfterSideEffects(t *testing.T) {
taskSvc := new(TaskService)
for _, title := range []string{"Design", "Build", "Ship"} {
var rsp AddResponse
if err := taskSvc.Add(context.Background(), &AddRequest{Title: title}, &rsp); err != nil {
t.Fatalf("Add(%q): %v", title, err)
}
}
notifySvc := new(NotifyService)
var rsp SendResponse
if err := notifySvc.Send(context.Background(), &SendRequest{To: "owner@acme.com", Message: "The launch plan is ready"}, &rsp); err != nil {
t.Fatalf("Send: %v", err)
}
done := make(chan error, 1)
done <- errors.New("agent run abc paused for approval: The comms agent is repeatedly timing out (408 errors) while retrying the launch-readiness notification")
if err := waitForPlanDelegateExecution(done, taskSvc, notifySvc, nil); err != nil {
t.Fatalf("waitForPlanDelegateExecution returned %v, want completed side effects to satisfy approval pause", err)
}
}
func TestPlanDelegateExecutionClassifiesClientTimeoutBeforeSideEffects(t *testing.T) {
done := make(chan error, 1)
done <- errors.New(`{"id":"go.micro.client","code":408,"detail":"<nil>","status":"Request Timeout"}`)
+59 -10
View File
@@ -269,13 +269,21 @@ func completeNotifyOnObservedSideEffect(ctx context.Context, in flow.State, ntf
in.Data = []byte("Buyer notified.")
return in, nil
}
wait := time.NewTimer(25 * time.Millisecond)
select {
case <-ctx.Done():
if dispatchErr != nil {
return in, dispatchErr
if !wait.Stop() {
<-wait.C
}
return in, ctx.Err()
case <-time.After(25 * time.Millisecond):
if dispatchErr == nil {
return in, ctx.Err()
}
// A timed-out Agent.Chat call can report the caller context as done
// while the remote agent is still finishing its notify tool call.
// Keep watching for the idempotent side effect until the local settle
// window expires so a post-side-effect timeout does not strand the
// durable checkout run as pending.
case <-wait.C:
}
}
if dispatchErr != nil {
@@ -349,14 +357,55 @@ func check(cond bool, format string, args ...any) {
// should not depend on a live model deciding to send another notification.
func a2aReachable(ctx context.Context, base, agent string) error {
probe := "A2A reachability probe only. Reply with the words concierge reachable. Do not call tools or send notifications."
reply, err := a2a.NewClient(base+"/agents/"+agent).Send(ctx, probe)
if err != nil {
return err
deadline, ok := ctx.Deadline()
if !ok {
deadline = time.Now().Add(10 * time.Second)
}
if strings.TrimSpace(reply) == "" {
return fmt.Errorf("empty A2A reply")
var lastErr error
for attempt := 1; ; attempt++ {
if err := ctx.Err(); err != nil {
if lastErr != nil {
return fmt.Errorf("A2A reachability probe failed after %d attempt(s): %w", attempt-1, lastErr)
}
return err
}
remaining := time.Until(deadline)
if remaining <= 0 {
if lastErr != nil {
return fmt.Errorf("A2A reachability probe failed after %d attempt(s): %w", attempt-1, lastErr)
}
return context.DeadlineExceeded
}
attemptTimeout := 4 * time.Second
if remaining < attemptTimeout {
attemptTimeout = remaining
}
attemptCtx, cancel := context.WithTimeout(ctx, attemptTimeout)
reply, err := a2a.NewClient(base+"/agents/"+agent).Send(attemptCtx, probe)
cancel()
if err == nil && strings.TrimSpace(reply) != "" {
return nil
}
if err == nil {
err = fmt.Errorf("empty A2A reply")
}
lastErr = err
if time.Until(deadline) <= 0 {
return fmt.Errorf("A2A reachability probe failed after %d attempt(s): %w", attempt, lastErr)
}
time.Sleep(minDuration(200*time.Millisecond*time.Duration(attempt), time.Until(deadline)))
}
return nil
}
func minDuration(a, b time.Duration) time.Duration {
if a < b {
return a
}
return b
}
func providerKey(provider string) string {
+64
View File
@@ -3,6 +3,9 @@ package main
import (
"context"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
@@ -25,6 +28,31 @@ func TestUniverseHarnessContract(t *testing.T) {
}
}
func TestA2AReachableRetriesTransientTimeout(t *testing.T) {
var calls int64
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if got, want := r.URL.Path, "/agents/concierge"; got != want {
t.Fatalf("path = %q, want %q", got, want)
}
if atomic.AddInt64(&calls, 1) == 1 {
time.Sleep(5 * time.Second)
return
}
w.Header().Set("Content-Type", "application/json")
fmt.Fprint(w, `{"jsonrpc":"2.0","id":1,"result":{"kind":"task","id":"task-1","contextId":"ctx-1","status":{"state":"completed"},"artifacts":[{"artifactId":"artifact-1","parts":[{"kind":"text","text":"concierge reachable"}]}]}}`)
}))
defer srv.Close()
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Second)
defer cancel()
if err := a2aReachable(ctx, srv.URL, "concierge"); err != nil {
t.Fatalf("a2aReachable returned error: %v", err)
}
if got := atomic.LoadInt64(&calls); got < 2 {
t.Fatalf("A2A calls = %d, want retry after transient timeout", got)
}
}
func TestNotifyStepCompletesAfterObservedSideEffectTimeout(t *testing.T) {
ntf := new(Notify)
before := atomic.LoadInt64(&ntf.sent)
@@ -69,6 +97,42 @@ func TestNotifyStepCompletesAfterObservedSideEffectTimeout(t *testing.T) {
}
}
func TestNotifyStepWaitsForObservedSideEffectAfterCanceledDispatchContext(t *testing.T) {
ntf := new(Notify)
before := atomic.LoadInt64(&ntf.sent)
ctx, cancel := context.WithCancel(context.Background())
cancel()
go func() {
time.Sleep(30 * time.Millisecond)
var rsp SendResponse
if err := ntf.Send(context.Background(), &SendRequest{
To: "buyer@acme.com",
Message: "Your order is confirmed.",
}, &rsp); err != nil {
t.Errorf("send notification: %v", err)
}
}()
out, err := completeNotifyOnObservedSideEffect(
ctx,
flow.State{Data: []byte(`{"order":"order-1"}`)},
ntf,
before,
time.Second,
errors.New("client observed timeout"),
)
if err != nil {
t.Fatalf("notify completion returned error: %v", err)
}
if got := out.String(); got != "Buyer notified." {
t.Fatalf("result = %q, want Buyer notified.", got)
}
if got := atomic.LoadInt64(&ntf.sent); got != 1 {
t.Fatalf("notifications sent = %d, want 1", got)
}
}
func TestNotifyStepRejectsClaimedCompletionWithoutSideEffect(t *testing.T) {
ntf := new(Notify)
before := atomic.LoadInt64(&ntf.sent)
+9 -8
View File
@@ -57,14 +57,15 @@ After this quick start, follow the agent path in order:
1. [Install troubleshooting](guides/install-troubleshooting.html) — verify the CLI install before agent work.
2. `micro agent demo` — print the provider-free first-agent demo command and next docs steps from the installed CLI.
3. `micro examples` — print the maintained provider-free runnable examples in copy/paste order.
4. `micro zero-to-hero` — print the maintained one-command no-secret lifecycle harness and runnable examples.
5. [Examples wayfinding index](https://github.com/micro/go-micro/blob/master/examples/INDEX.md) — choose the smallest no-secret first-agent, maintained [0→hero support reference](https://github.com/micro/go-micro/tree/master/examples/support), and next interop examples from one map.
6. [Smallest first-agent example](https://github.com/micro/go-micro/tree/master/examples/first-agent) — run one service-backed agent with a mock model and no provider key.
7. [No-secret first-agent transcript](guides/no-secret-first-agent.html) — run a useful support agent with a mock model before setting up a provider key.
8. [Your First Agent](guides/your-first-agent.html) — build a service-backed agent and talk to it with `micro chat`.
9. [Debugging your agent](guides/debugging-agents.html) — use `micro inspect agent <name>` to inspect service registration, tool calls, run history, memory, provider failures, and flow handoffs when the agent surprises you.
10. [0→hero reference path](guides/zero-to-hero.html) — prove the full scaffold → run → chat → inspect → deploy dry-run lifecycle with commands exercised by `make harness`.
3. `micro agent quickcheck` (or `micro agent debug`) — when scaffold → run → chat → inspect stalls, print the short recovery map before you dive into the full debugging guide.
4. `micro examples` — print the maintained provider-free runnable examples in copy/paste order.
5. `micro zero-to-hero` — print the maintained one-command no-secret lifecycle harness and runnable examples.
6. [Examples wayfinding index](https://github.com/micro/go-micro/blob/master/examples/INDEX.md) — choose the smallest no-secret first-agent, maintained [0→hero support reference](https://github.com/micro/go-micro/tree/master/examples/support), and next interop examples from one map.
7. [Smallest first-agent example](https://github.com/micro/go-micro/tree/master/examples/first-agent) — run one service-backed agent with a mock model and no provider key.
8. [No-secret first-agent transcript](guides/no-secret-first-agent.html) — run a useful support agent with a mock model before setting up a provider key.
9. [Your First Agent](guides/your-first-agent.html) — build a service-backed agent and talk to it with `micro chat`.
10. [Debugging your agent](guides/debugging-agents.html) — use `micro agent preflight` before `micro run`, `micro agent doctor` after `micro run`, then `micro chat` and `micro inspect agent <name>` to recover service registration, tool calls, run history, memory, provider failures, and flow handoffs when the agent surprises you.
11. [0→hero reference path](guides/zero-to-hero.html) — prove the full scaffold → run → chat → inspect → deploy dry-run lifecycle with commands exercised by `make harness`.
## Write a Service
@@ -24,10 +24,11 @@ Use the recovery command that matches where you are in the first-agent journey:
| Checkpoint | When to use it | Command |
| --- | --- | --- |
| Install troubleshooting | `micro` is not installed, not on `PATH`, or the shell cannot run it. | [Install troubleshooting](install-troubleshooting.html) |
| Quick recovery map | The first-agent loop stalled and you want the short scaffold → run → chat → inspect checklist before reading this full guide. | `micro agent quickcheck` (alias: `micro agent debug`) |
| Preflight before `micro run` | You have not started the local runtime yet and want to verify Go, CLI, provider-key, and gateway-port prerequisites. | `micro agent preflight` |
| Doctor after `micro run` | `micro run` is active, but chat, the `/agent` gateway, agent registration, provider settings, or inspect/run history is not behaving. | `micro agent doctor` |
`micro agent preflight` is read-only and runs before the first local run; failed
`micro agent quickcheck` is the quickest breadcrumb when you are unsure where the first-agent path failed: it prints the preflight, run, doctor, inspect, and no-secret fallback commands in one place. `micro agent preflight` is read-only and runs before the first local run; failed
checks include `Fix:` and `Next:` lines for Go, CLI installation, provider-key
setup, and the local gateway port. Once `micro run` is already up, switch to
`micro agent doctor` so the recovery output follows the live gateway, chat