Compare commits

..

14 Commits

Author SHA1 Message Date
Codex d3b7a18a25 docs: refresh planner priorities for 4440
Harness (E2E) / Harnesses (mock LLM) (push) Waiting to run
Harness (E2E) / Provider harnesses (live LLM conformance) (push) Waiting to run
Lint / golangci-lint (push) Waiting to run
Run Tests / Unit Tests (push) Waiting to run
Run Tests / Etcd Integration Tests (push) Waiting to run
2026-07-09 07:25:13 +00:00
Asim Aslam ed43db5276 test: lock zero-to-hero harness labels (#4439)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 07:02:35 +01:00
Asim Aslam 9d6d2d6c91 docs: refresh planner priorities for 4434 (#4435)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 06:26:29 +01:00
Asim Aslam 776fe1a36a Add agent stream provider conformance (#4433)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 04:56:44 +01:00
Asim Aslam 84c1ee7471 docs: refresh planner priorities for 4427 (#4429)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 04:18:01 +01:00
Asim Aslam 1a9e94219a docs: reconcile roadmap agent status (#4426)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 03:29:03 +01:00
Asim Aslam c7df280d93 docs: refresh planner queue for 4422 (#4424)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 02:38:30 +01:00
Asim Aslam 3b367975a3 Fix retry timeout test race (#4421)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 02:17:36 +01:00
Asim Aslam 0d4a101532 docs: refresh planner priorities for 4416 (#4417)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 01:44:05 +01:00
Asim Aslam ad8ff2abde Harden model call timeout enforcement (#4411)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 01:05:25 +01:00
Asim Aslam 5e6d51d261 docs: refresh planner priorities for 4407 (#4409)
Co-authored-by: Codex <codex@openai.com>
2026-07-09 00:42:18 +01:00
Asim Aslam 687a33faee Preserve checkpointed tool calls on agent resume (#4406)
goreleaser / goreleaser (push) Waiting to run
Co-authored-by: Codex <codex@openai.com>
2026-07-09 00:01:13 +01:00
Asim Aslam 48e752bbce docs: refresh planner priorities for 4403 (#4404)
Co-authored-by: Codex <codex@openai.com>
2026-07-08 23:34:57 +01:00
Asim Aslam b9e4bd171f Improve getting-started harness logs (#4402)
Co-authored-by: Codex <codex@openai.com>
2026-07-08 23:04:02 +01:00
13 changed files with 333 additions and 17 deletions
+4 -3
View File
@@ -21,9 +21,10 @@ changes, architectural rewrites. Those go to the human.
## Work queue (ranked)
1. **CI-verify the getting-started contract** ([#4399](https://github.com/micro/go-micro/issues/4399)) — The North Star says developer adoption is the gap, and the README now promises install → scaffold → run → chat → inspect → deploy-dry-run as a maintained 0→1/0→hero path. Make that path a first-class no-secret CI contract so the on-ramp cannot drift while deeper harness work continues.
2. **Resume agent runs from checkpoints** ([#4368](https://github.com/micro/go-micro/issues/4368)) — #4391 closed the OpenTelemetry RunInfo gap and #4397 documented resume checkpoint limits, so the highest-value remaining depth seam is a focused, non-breaking durable-agent resume slice that preserves completed tool calls and avoids duplicate side effects before broader durability or API design work.
3. **Broaden provider streaming conformance** ([#4386](https://github.com/micro/go-micro/issues/4386)) — The blog says Anthropic streaming shipped, but the roadmap still calls for provider-backed streaming across chat and A2A. Add a focused, provider-gated conformance slice so streaming stays end-to-end rather than becoming a one-provider success story.
1. **Handle AtlasCloud plan-delegate partial tool calls** ([#4431](https://github.com/micro/go-micro/issues/4431)) — A current live conformance failure can let AtlasCloud/minimax stop at a partial XML-style tool call, so delegation never produces the required task/notify side effects. This is the highest Now-phase risk because plan/delegate is the core services → agents → workflows seam and must fail or recover deterministically across providers.
2. **Verify first-agent docs wayfinding stays in lockstep** ([#4441](https://github.com/micro/go-micro/issues/4441)) — Developer adoption remains the current goal after the 0→1/0→hero harness landed. Keep README, website guides, examples, and CLI breadcrumbs aligned so the no-secret first-agent path stays discoverable instead of becoming a stale documentation promise.
3. **Fix AtlasCloud tool streaming capability mismatch** ([#4438](https://github.com/micro/go-micro/issues/4438)) — The provider matrix currently advertises AtlasCloud streaming capability beyond its tool-streaming implementation, causing the live agent streaming conformance path to run an unsupported assertion. Fixing the capability/implementation seam keeps streaming honest without blocking no-secret adoption work.
4. **Broaden provider streaming conformance** ([#4386](https://github.com/micro/go-micro/issues/4386)) — Once the provider-specific AtlasCloud mismatch is resolved, expand the broader streaming matrix so chat, agent RPC, and A2A streaming regressions are caught across keyed providers while local CI continues to skip cleanly without secrets.
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
architecture-review pass._
+18 -4
View File
@@ -14,8 +14,9 @@ The full, current roadmap lives at **[go-micro.dev/docs/roadmap](https://go-micr
## Where we are (v6)
Services, agents (`plan`/`delegate`, guardrails, memory, tool middleware), durable
flows, the MCP and A2A gateways (both directions, including A2A streaming,
Services, agents (`plan`/`delegate`, guardrails, memory, tool middleware,
checkpoint/resume, and OpenTelemetry run spans), durable flows, the MCP and A2A
gateways (both directions, including A2A streaming,
push notifications, and multi-turn continuation), x402 paid tools, secure by
default.
@@ -39,11 +40,24 @@ default.
propagation, retry/backoff.
- **Getting-started contract** — define and CI-verify the 0→1 and 0→hero flows.
## Shipped agent depth
- **Durable agent loop** — opt-in `Checkpoint` support lets agent `Ask` and
streaming runs persist, list pending work, and resume without replaying completed
tool calls. Human-input pauses resume through explicit input helpers.
- **Agent observability** — agent `RunInfo` now feeds OpenTelemetry spans/events
across runs, model turns, tool calls, retries, delegation lineage, and resume
checkpoints.
## Next — agentic depth
- **Durable agent loop** — resume a long run via `Checkpoint` (flows already do).
- **Streaming** — broaden provider-backed `ai.Stream` coverage and keep chat/A2A streaming end to end.
- **Agent observability** — `RunInfo` → OpenTelemetry spans.
- **Resume operations polish** — keep improving CLI/docs breadcrumbs for finding
pending agent runs and deciding whether to call resume, resume-input, or stream
resume in production.
- **Observability hardening** — keep span attributes and run inspection coherent
across agents, flows, and gateways as more providers and workflow paths are
exercised.
## Later
+5 -1
View File
@@ -432,9 +432,13 @@ func (a *agentImpl) askLocked(ctx context.Context, runID, message, parentRunID s
reply += resp.Answer
}
completedToolCalls := checkpointToolCalls(run.Steps)
if a.currentRun != nil {
completedToolCalls = checkpointToolCalls(a.currentRun.Steps)
}
res := &Response{
Reply: reply,
ToolCalls: resp.ToolCalls,
ToolCalls: mergeCheckpointToolCalls(completedToolCalls, resp.ToolCalls),
Agent: a.opts.Name,
RunID: a.runID,
ParentID: parentRunID,
+54
View File
@@ -4,6 +4,7 @@ import (
"context"
"encoding/json"
"fmt"
"strings"
"time"
"go-micro.dev/v6/ai"
@@ -288,3 +289,56 @@ func upsertStep(steps *[]flow.StepRecord, rec flow.StepRecord) int {
*steps = append(*steps, rec)
return len(*steps) - 1
}
func checkpointToolCalls(steps []flow.StepRecord) []ai.ToolCall {
calls := make([]ai.ToolCall, 0, len(steps))
for _, step := range steps {
call, ok := checkpointToolCall(step)
if !ok {
continue
}
calls = append(calls, call)
}
return calls
}
func checkpointToolCall(step flow.StepRecord) (ai.ToolCall, bool) {
if step.Status != "done" || !strings.HasPrefix(step.Name, "tool:") {
return ai.ToolCall{}, false
}
parts := strings.SplitN(strings.TrimPrefix(step.Name, "tool:"), ":", 2)
if len(parts) != 2 || parts[0] == "" {
return ai.ToolCall{}, false
}
input := map[string]any{}
if parts[1] != "null" && parts[1] != "" {
if err := json.Unmarshal([]byte(parts[1]), &input); err != nil {
return ai.ToolCall{}, false
}
}
return ai.ToolCall{Name: parts[0], Input: input, Result: step.Result}, true
}
func mergeCheckpointToolCalls(checkpointed, current []ai.ToolCall) []ai.ToolCall {
if len(checkpointed) == 0 {
return current
}
seen := make(map[string]struct{}, len(current))
for _, call := range current {
seen[toolCallKey(call.Name, call.Input)] = struct{}{}
}
merged := make([]ai.ToolCall, 0, len(checkpointed)+len(current))
for _, call := range checkpointed {
if _, ok := seen[toolCallKey(call.Name, call.Input)]; ok {
continue
}
merged = append(merged, call)
}
merged = append(merged, current...)
return merged
}
func toolCallKey(name string, input map[string]any) string {
b, _ := json.Marshal(input)
return name + ":" + string(b)
}
+6
View File
@@ -104,6 +104,12 @@ func TestResumeFailedCheckpointDoesNotReplayCompletedTool(t *testing.T) {
if resp.Reply != "finished from checkpoint" {
t.Fatalf("Resume reply = %q", resp.Reply)
}
if len(resp.ToolCalls) != 1 || resp.ToolCalls[0].Name != "external.charge" || resp.ToolCalls[0].Result != "charged" {
t.Fatalf("resumed tool calls = %#v, want preserved completed charge call", resp.ToolCalls)
}
if got := resp.ToolCalls[0].Input["order"]; got != "42" {
t.Fatalf("resumed tool input order = %#v, want 42", got)
}
if toolRuns != 1 {
t.Fatalf("tool executions after Resume = %d, want completed tool was not replayed", toolRuns)
}
+111
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"fmt"
"io"
"os"
"strings"
"testing"
@@ -45,6 +46,116 @@ func TestAgentProviderConformanceMatrix(t *testing.T) {
}
}
func TestAgentProviderStreamConformanceMatrix(t *testing.T) {
providers := []conformanceProvider{
{name: "fake"},
{name: "openai", key: "OPENAI_API_KEY", model: "GO_MICRO_CONFORMANCE_OPENAI_MODEL", live: true},
{name: "anthropic", key: "ANTHROPIC_API_KEY", model: "GO_MICRO_CONFORMANCE_ANTHROPIC_MODEL", live: true},
{name: "atlascloud", key: "ATLASCLOUD_API_KEY", model: "GO_MICRO_CONFORMANCE_ATLASCLOUD_MODEL", live: true},
{name: "groq", key: "GROQ_API_KEY", model: "GO_MICRO_CONFORMANCE_GROQ_MODEL", live: true},
{name: "minimax", key: "MINIMAX_API_KEY", model: "GO_MICRO_CONFORMANCE_MINIMAX_MODEL", live: true},
{name: "mistral", key: "MISTRAL_API_KEY", model: "GO_MICRO_CONFORMANCE_MISTRAL_MODEL", live: true},
{name: "together", key: "TOGETHER_API_KEY", model: "GO_MICRO_CONFORMANCE_TOGETHER_MODEL", live: true},
}
selected := selectedConformanceProviders(os.Getenv("GO_MICRO_AGENT_CONFORMANCE_PROVIDERS"))
for _, provider := range providers {
provider := provider
if len(selected) > 0 && !selected[provider.name] {
continue
}
t.Run(provider.name, func(t *testing.T) {
runAgentStreamConformanceScenario(t, provider)
})
}
}
func runAgentStreamConformanceScenario(t *testing.T, provider conformanceProvider) {
t.Helper()
if provider.live {
if os.Getenv(provider.key) == "" {
t.Skipf("%s not set; skipping live %s stream conformance", provider.key, provider.name)
}
if os.Getenv("GO_MICRO_AGENT_CONFORMANCE_LIVE") == "" {
t.Skipf("GO_MICRO_AGENT_CONFORMANCE_LIVE not set; skipping live %s stream conformance", provider.name)
}
if caps := ai.ProviderCapabilities(provider.name); !caps.Stream {
t.Fatalf("ProviderCapabilities(%q).Stream = false, want true for stream conformance", provider.name)
}
} else {
var sawToolSchema bool
fakeStream = func(ctx context.Context, opts ai.Options, req *ai.Request) (ai.Stream, error) {
if req.Prompt != "Stream exactly: agent-stream-conformance-ok" {
return nil, fmt.Errorf("prompt = %q", req.Prompt)
}
if len(req.Messages) == 0 || req.Messages[len(req.Messages)-1].Role != "user" || req.Messages[len(req.Messages)-1].Content != req.Prompt {
return nil, fmt.Errorf("messages = %#v, want current user turn", req.Messages)
}
for _, tool := range req.Tools {
if tool.Name == "conformance_echo" {
sawToolSchema = true
}
}
if !sawToolSchema {
return nil, errors.New("stream request omitted conformance tool schema")
}
return &sliceStream{chunks: []string{"agent-stream-", "conformance-ok"}}, nil
}
defer func() { fakeStream = nil }()
}
agentOpts := []Option{
Name("stream-conformance-" + provider.name),
Provider(provider.name),
APIKey(os.Getenv(provider.key)),
Prompt("Stream conformance: preserve the exact requested marker in the final answer."),
WithRegistry(registry.NewMemoryRegistry()),
WithStore(store.NewMemoryStore()),
WithMemory(NewInMemory(8)),
ModelCallTimeout(45 * time.Second),
WithTool("conformance_echo", "Echo a conformance value and return a deterministic marker.", map[string]any{
"value": map[string]any{"type": "string", "description": "value to echo"},
}, func(ctx context.Context, input map[string]any) (string, error) {
return `{"marker":"agent-stream-conformance-ok"}`, nil
}),
}
if provider.model != "" {
if model := os.Getenv(provider.model); model != "" {
agentOpts = append(agentOpts, Model(model))
}
}
stream, err := New(agentOpts...).Stream(context.Background(), "Stream exactly: agent-stream-conformance-ok")
if err != nil {
t.Fatalf("Stream: %v", err)
}
defer stream.Close()
var reply strings.Builder
deadline := time.After(45 * time.Second)
for {
select {
case <-deadline:
t.Fatal("timed out waiting for streamed final output")
default:
}
chunk, err := stream.Recv()
if errors.Is(err, io.EOF) {
break
}
if err != nil {
t.Fatalf("Recv: %v", err)
}
reply.WriteString(chunk.Reply)
if strings.Contains(reply.String(), "agent-stream-conformance-ok") {
return
}
}
if got := reply.String(); !strings.Contains(got, "agent-stream-conformance-ok") {
t.Fatalf("streamed reply %q does not include conformance marker", got)
}
}
func selectedConformanceProviders(csv string) map[string]bool {
out := map[string]bool{}
for _, part := range strings.Split(csv, ",") {
+49
View File
@@ -240,6 +240,55 @@ func TestAskCancellationDuringToolCallFailsRun(t *testing.T) {
}
}
func TestSlowProviderTimeoutPreventsLateToolSideEffects(t *testing.T) {
started := make(chan struct{})
release := make(chan struct{})
done := make(chan struct{})
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
close(started)
<-release
defer close(done)
if opts.ToolHandler == nil {
t.Fatal("missing tool handler")
}
res := opts.ToolHandler(ctx, ai.ToolCall{ID: "late-1", Name: "external.create", Input: map[string]any{"title": "too late"}})
if !strings.Contains(res.Content, context.DeadlineExceeded.Error()) {
t.Errorf("late tool result = %q, want deadline exceeded", res.Content)
}
return &ai.Response{Reply: "late", ToolCalls: []ai.ToolCall{{ID: "late-1", Name: "external.create", Input: map[string]any{"title": "too late"}, Result: res.Content}}}, nil
}
defer func() { fakeGen = nil }()
toolRuns := 0
a := newTestAgent(
Name("slow-provider-late-tool"),
ModelCallTimeout(10*time.Millisecond),
WithTool("external.create", "create once", nil, func(context.Context, map[string]any) (string, error) {
toolRuns++
return "created", nil
}),
)
_, err := a.Ask(context.Background(), "provider times out before tool")
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("Ask error = %v, want deadline exceeded", err)
}
select {
case <-started:
default:
t.Fatal("provider was not called")
}
close(release)
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("late provider call did not finish")
}
if toolRuns != 0 {
t.Fatalf("late tool executions = %d, want 0", toolRuns)
}
}
func TestAskCheckpointRecordsTerminalOperationalFailureStatus(t *testing.T) {
tests := []struct {
name string
+22 -1
View File
@@ -157,7 +157,7 @@ func GenerateWithRetry(ctx context.Context, m Model, req *Request, policy Genera
info.MaxAttempts = policy.MaxAttempts
callCtx = WithRunInfo(callCtx, info)
}
resp, err := m.Generate(callCtx, req, opts...)
resp, err := generateAttempt(callCtx, m, req, opts...)
cancel()
// Caller cancellation/deadline always wins and is not retried, even if
@@ -197,6 +197,27 @@ func GenerateWithRetry(ctx context.Context, m Model, req *Request, policy Genera
return nil, &RetryError{Attempts: policy.MaxAttempts, Kind: ClassifyError(last), Err: last}
}
func generateAttempt(ctx context.Context, m Model, req *Request, opts ...GenerateOption) (*Response, error) {
if err := ctx.Err(); err != nil {
return nil, err
}
type result struct {
resp *Response
err error
}
done := make(chan result, 1)
go func() {
resp, err := m.Generate(ctx, req, opts...)
done <- result{resp: resp, err: err}
}()
select {
case res := <-done:
return res.resp, res.err
case <-ctx.Done():
return nil, ctx.Err()
}
}
func retryBackoff(err error, attempt int, base time.Duration) time.Duration {
backoff := base
if backoff <= 0 {
+33 -4
View File
@@ -4,6 +4,7 @@ import (
"context"
"errors"
"net/http"
"sync/atomic"
"testing"
"time"
)
@@ -69,9 +70,9 @@ func TestGenerateWithRetryDoesNotRetryCallerCancellation(t *testing.T) {
}
func TestGenerateWithRetryHonorsPerAttemptTimeout(t *testing.T) {
attempts := 0
var attempts atomic.Int32
model := retryModel{generate: func(ctx context.Context, _ *Request, _ ...GenerateOption) (*Response, error) {
attempts++
attempts.Add(1)
<-ctx.Done()
return nil, ctx.Err()
}}
@@ -91,8 +92,8 @@ func TestGenerateWithRetryHonorsPerAttemptTimeout(t *testing.T) {
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("error = %v, want context.DeadlineExceeded", err)
}
if attempts != 2 {
t.Fatalf("attempts = %d, want 2", attempts)
if got := attempts.Load(); got != 2 {
t.Fatalf("attempts = %d, want 2", got)
}
}
@@ -135,6 +136,34 @@ func TestGenerateWithRetryAddsAttemptMetadataToRunInfo(t *testing.T) {
}
}
func TestGenerateWithRetryReturnsWhenProviderIgnoresTimeout(t *testing.T) {
started := make(chan struct{})
release := make(chan struct{})
model := retryModel{generate: func(ctx context.Context, req *Request, opts ...GenerateOption) (*Response, error) {
close(started)
<-release
return &Response{Reply: "late"}, nil
}}
defer close(release)
start := time.Now()
_, err := GenerateWithRetry(context.Background(), model, &Request{Prompt: "hi"}, GeneratePolicy{
Timeout: 10 * time.Millisecond,
MaxAttempts: 1,
})
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("GenerateWithRetry error = %v, want deadline exceeded", err)
}
if elapsed := time.Since(start); elapsed > 200*time.Millisecond {
t.Fatalf("GenerateWithRetry took %s after deadline, want prompt return", elapsed)
}
select {
case <-started:
default:
t.Fatal("provider was not called")
}
}
type statusErr int
func (e statusErr) Error() string { return "provider status" }
@@ -421,7 +421,7 @@ func runAgentConformance(provider string, timeout time.Duration) error {
if provider == "mock" {
testProvider = "fake"
}
cmd := exec.CommandContext(ctx, "go", "test", "./agent", "-run", "TestAgentProviderConformanceMatrix", "-count=1", "-v")
cmd := exec.CommandContext(ctx, "go", "test", "./agent", "-run", "TestAgentProvider(ConformanceMatrix|StreamConformanceMatrix)", "-count=1", "-v")
cmd.Dir = repoRoot()
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
@@ -50,6 +50,19 @@ func TestZeroToHeroReferenceDocs(t *testing.T) {
t.Fatalf("0→hero CI run script missing lifecycle command %q", want)
}
}
for _, want := range []string{
"scaffold:",
"run/chat/inspect:",
"deploy dry-run:",
"chat/inspect:",
"first-agent app:",
"0→hero app:",
"flow history:",
} {
if !strings.Contains(runScript, want) {
t.Fatalf("0→hero CI run script missing debuggable boundary label %q", want)
}
}
readme := readFile(t, filepath.Join(root, "README.md"))
if !strings.Contains(readme, "internal/website/docs/guides/zero-to-hero.md") {
+1 -1
View File
@@ -35,5 +35,5 @@ run_step "first-agent app: runnable provider-free example" \
go test ./examples/first-agent -run TestRunFirstAgent -count=1
run_step "0→hero app: support lifecycle smoke" \
go test ./examples/support -run 'TestRunSupportMockSmoke|TestZeroToHeroReadmeDocumentsLifecycle' -count=1
run_step "workflows: deterministic services → agents → workflows harnesses" \
run_step "flow history: deterministic services → agents → workflows harnesses" \
go test ./internal/harness/universe ./internal/harness/plan-delegate -run 'Test.*Harness|TestPlanDelegateEndToEnd|TestPlanDelegateFlowHandoff' -count=1
+16 -2
View File
@@ -11,7 +11,7 @@ Go Micro is a framework for building **agents and services** in Go. An agent is
The foundation is in place:
- **Services** — register, discover, RPC, events; every endpoint is automatically an MCP tool.
- **Agents** — a model with memory and tools that manages services, with `plan`, `delegate`, and guardrails (`MaxSteps`, `LoopLimit`, `ApproveTool`) built in, plus tool-execution middleware (`WrapTool`) and run metadata.
- **Agents** — a model with memory and tools that manages services, with `plan`, `delegate`, guardrails (`MaxSteps`, `LoopLimit`, `ApproveTool`), tool-execution middleware (`WrapTool`), run metadata, checkpoint/resume, and OpenTelemetry run spans built in.
- **Flows** — durable, event-driven workflows: ordered steps that checkpoint and resume after a crash.
- **Interop** — the MCP gateway (services as tools) and the A2A gateway (agents as agents, both directions, including A2A streaming, push notifications, and multi-turn continuation), both generated from the registry; x402 for paid tools.
- **Secure by default** — TLS verification on, state scoped per component.
@@ -34,10 +34,24 @@ The priority is that what exists works everywhere, under real conditions.
- **Failure & resilience.** Provider timeouts, rate limits, and cancellation mid-run; deadline/`context` propagation through the agent loop; retry and backoff at the model call.
- **The getting-started contract.** Define and CI-verify the 0→1 and 0→hero flows so they can't silently break.
## Shipped agent depth
- **Durable agent loop.** Opt-in `Checkpoint` support now lets agent `Ask` and
streaming runs persist, list pending work, and resume without replaying completed
tool calls. Human-input pauses resume through explicit input helpers.
- **Agent observability.** `RunInfo` now feeds OpenTelemetry spans and events for
agent runs, model turns, tool calls, retries, delegation lineage, and resume
checkpoints so production runs are traceable.
## Next — agentic depth
- **Streaming.** Broaden provider-backed `ai.Stream` coverage and keep chat plus A2A `message/stream` working end to end for real chat and long-task UX.
- **Agent observability.** Wire the new `RunInfo` into OpenTelemetry spans so a run — steps, tool calls, delegation — is traceable. This is also what anyone running it in production will need.
- **Resume operations polish.** Keep improving CLI/docs breadcrumbs for finding
pending agent runs and deciding whether to call resume, resume-input, or stream
resume in production.
- **Observability hardening.** Keep span attributes and run inspection coherent
across agents, flows, and gateways as more providers and workflow paths are
exercised.
## Later