Compare commits
22 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| d2801fcc37 | |||
| e96898d362 | |||
| c88a090d10 | |||
| f6951d2bb0 | |||
| 94286204ad | |||
| d54139c02d | |||
| 5ff3760cd4 | |||
| 69ab520360 | |||
| 0c4ea84ee3 | |||
| e6a6c72038 | |||
| 7c036459cf | |||
| 83cb4ff1a9 | |||
| 941d43bdf5 | |||
| bb3e40ade7 | |||
| dbce523437 | |||
| 98125cd770 | |||
| 3454a08079 | |||
| decfa7c63e | |||
| b6980e27a9 | |||
| d92d63943c | |||
| ec1d46526c | |||
| 40308cf779 |
@@ -21,8 +21,9 @@ changes, architectural rewrites. Those go to the human.
|
||||
|
||||
## Work queue (ranked)
|
||||
|
||||
1. **Make plan-delegate harness complete delegated notification steps** ([#4138](https://github.com/micro/go-micro/issues/4138)) — the duplicate-notification fix and first-agent tutorial harness have both shipped, but the live atlascloud plan-delegate matrix still exposed a higher-severity Now-phase contract break: a delegated notification can succeed or be reused while the persisted plan step remains incomplete. Fixing that keeps the provider conformance evaluator trustworthy without relaxing the services → agents → workflows lifecycle gate.
|
||||
2. **Verify no-secret agent debugging walkthrough** ([#4142](https://github.com/micro/go-micro/issues/4142)) — keep adoption pressure balanced with hardening by extending the maintained 0→1 path past scaffold/run/chat into inspect and debugging. A provider-free smoke check for the documented first-agent debug sequence will catch command drift at the exact seam where new developers need confidence after the first conversation behaves unexpectedly.
|
||||
1. **Stabilize AtlasCloud guarded-delegate conformance** ([#4206](https://github.com/micro/go-micro/issues/4206)) — the installed first-agent on-ramp shipped in #4207, so the top open Now-phase risk is the live provider conformance recurrence where AtlasCloud/minimax-m3 can complete the scenario without reliably exercising the guarded `delegate` path. Keep this scoped to deterministic prompt/retry/conformance coverage so the harness proves service-backed delegation works under real provider behavior without changing public APIs.
|
||||
2. **Propagate cancellation and retry signals through provider model calls** ([#4175](https://github.com/micro/go-micro/issues/4175)) — after the immediate live-conformance regression is closed, the remaining Now-phase reliability gap is failure handling under real provider conditions: cancellation/deadline propagation and retry/backoff must not duplicate tool side effects. This keeps the services → agents → workflows lifecycle dependable across providers without changing public APIs.
|
||||
3. **Resume durable agent runs from checkpoints** ([#4202](https://github.com/micro/go-micro/issues/4202)) — once the provider conformance and failure semantics are guarded, move into the Next-phase durable-agent-loop work: long-running agents should recover like flows, checkpoint progress, and avoid replaying completed tool side effects after interruption.
|
||||
|
||||
_Seeded by Claude Code from the roadmap + open issues; thereafter maintained by the
|
||||
architecture-review pass._
|
||||
|
||||
@@ -31,6 +31,7 @@ Running Go Micro in production, or building on it and want help? Paid **support,
|
||||
- [Building Agents](#building-agents) — [Plan & Delegate](#plan--delegate), [Pluggable](#batteries-included-pluggable), [Paid tools (x402)](#paid-tools-x402), [A2A](#reachable-by-other-agents-a2a)
|
||||
- [Features](#features)
|
||||
- [CLI](#cli)
|
||||
- [Autonomous improvement loop](#autonomous-improvement-loop)
|
||||
- [Multi-Service Projects](#multi-service-projects)
|
||||
- [Data Model](#data-model)
|
||||
- [AI Providers](#ai-providers)
|
||||
@@ -105,6 +106,25 @@ walkable agent path in this order:
|
||||
services → agents → workflows loop with scaffold, run, chat, inspect, flow
|
||||
history, and deploy dry-run commands that match the maintained harness.
|
||||
|
||||
### Autonomous improvement loop
|
||||
|
||||
Want the same services → agents → workflows lifecycle applied to your
|
||||
repository? `micro loop` scaffolds the autonomous improvement loop used by Go
|
||||
Micro itself: a North Star, ranked issue queue, role prompts, GitHub Actions
|
||||
workflows, and verification for CI-gated PRs.
|
||||
|
||||
```bash
|
||||
micro loop init --roles all
|
||||
micro loop verify
|
||||
```
|
||||
|
||||
Before turning on the schedule, configure a dispatch token such as
|
||||
`CODEX_TRIGGER_TOKEN`, protect the default branch with required CI checks
|
||||
(`go build ./...`, `go test ./...`, and `golangci-lint run ./...` for this
|
||||
repository), and seed `.github/loop/PRIORITIES.md` with one scoped issue per
|
||||
increment. See the [`micro loop` quickstart](internal/website/docs/guides/micro-loop.md)
|
||||
for the setup checklist and operating model.
|
||||
|
||||
### Generate from a prompt — with an LLM key
|
||||
|
||||
Set a provider key, describe what you want, and the AI designs services, writes handlers, compiles, and starts them:
|
||||
|
||||
+59
-8
@@ -2,6 +2,7 @@ package agent
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
@@ -591,6 +592,9 @@ func (a *agentImpl) handleDelegate(ctx context.Context, call ai.ToolCall) ai.Too
|
||||
return errResult(call.ID, "task is required")
|
||||
}
|
||||
to, _ := input["to"].(string)
|
||||
if cached, ok := a.cachedDelegateResult(call.ID, to, task); ok {
|
||||
return cached
|
||||
}
|
||||
|
||||
// An external agent on another framework, addressed by A2A URL.
|
||||
if strings.HasPrefix(to, "http://") || strings.HasPrefix(to, "https://") {
|
||||
@@ -598,9 +602,7 @@ func (a *agentImpl) handleDelegate(ctx context.Context, call ai.ToolCall) ai.Too
|
||||
if err != nil {
|
||||
return errResult(call.ID, "delegate to A2A agent "+to+": "+err.Error())
|
||||
}
|
||||
out := map[string]any{"agent": to, "reply": reply}
|
||||
b, _ := json.Marshal(out)
|
||||
return ai.ToolResult{ID: call.ID, Value: out, Content: string(b)}
|
||||
return a.storeDelegateResult(call.ID, to, task, map[string]any{"agent": to, "reply": reply})
|
||||
}
|
||||
|
||||
// Delegate-first: an existing agent that owns the domain handles it.
|
||||
@@ -609,9 +611,7 @@ func (a *agentImpl) handleDelegate(ctx context.Context, call ai.ToolCall) ai.Too
|
||||
if err != nil {
|
||||
return errResult(call.ID, "delegate to agent "+to+": "+err.Error())
|
||||
}
|
||||
out := map[string]any{"agent": to, "reply": reply}
|
||||
b, _ := json.Marshal(out)
|
||||
return ai.ToolResult{ID: call.ID, Value: out, Content: string(b)}
|
||||
return a.storeDelegateResult(call.ID, to, task, map[string]any{"agent": to, "reply": reply})
|
||||
}
|
||||
|
||||
// Otherwise create a focused, ephemeral sub-agent. Fresh context:
|
||||
@@ -644,9 +644,60 @@ func (a *agentImpl) handleDelegate(ctx context.Context, call ai.ToolCall) ai.Too
|
||||
if err != nil {
|
||||
return errResult(call.ID, "sub-agent: "+err.Error())
|
||||
}
|
||||
out := map[string]any{"reply": resp.Reply}
|
||||
return a.storeDelegateResult(call.ID, to, task, map[string]any{"reply": resp.Reply})
|
||||
}
|
||||
|
||||
func (a *agentImpl) cachedDelegateResult(id, to, task string) (ai.ToolResult, bool) {
|
||||
recs, err := a.stateStore().Read(delegateResultKey(to, task))
|
||||
if err != nil || len(recs) == 0 {
|
||||
return ai.ToolResult{}, false
|
||||
}
|
||||
var out map[string]any
|
||||
if err := json.Unmarshal(recs[0].Value, &out); err != nil {
|
||||
return ai.ToolResult{}, false
|
||||
}
|
||||
b, _ := json.Marshal(out)
|
||||
return ai.ToolResult{ID: call.ID, Value: out, Content: string(b)}
|
||||
return ai.ToolResult{ID: id, Value: out, Content: string(b)}, true
|
||||
}
|
||||
|
||||
func (a *agentImpl) storeDelegateResult(id, to, task string, out map[string]any) ai.ToolResult {
|
||||
b, _ := json.Marshal(out)
|
||||
_ = a.stateStore().Write(&store.Record{Key: delegateResultKey(to, task), Value: b})
|
||||
return ai.ToolResult{ID: id, Value: out, Content: string(b)}
|
||||
}
|
||||
|
||||
func delegateResultKey(to, task string) string {
|
||||
fp := normalizeDelegateTarget(to) + "\x00" + normalizeDelegateTask(task)
|
||||
sum := sha256.Sum256([]byte(fp))
|
||||
return fmt.Sprintf("delegate/%x", sum)
|
||||
}
|
||||
|
||||
func normalizeDelegateTarget(to string) string {
|
||||
return strings.Join(strings.Fields(strings.ToLower(strings.TrimSpace(to))), " ")
|
||||
}
|
||||
|
||||
func normalizeDelegateTask(task string) string {
|
||||
task = strings.ToLower(strings.TrimSpace(task))
|
||||
task = strings.Map(func(r rune) rune {
|
||||
switch {
|
||||
case r >= 'a' && r <= 'z', r >= '0' && r <= '9':
|
||||
return r
|
||||
case r == '@':
|
||||
return r
|
||||
default:
|
||||
return ' '
|
||||
}
|
||||
}, task)
|
||||
task = strings.Join(strings.Fields(task), " ")
|
||||
if strings.Contains(task, "notify") &&
|
||||
strings.Contains(task, "owner") &&
|
||||
strings.Contains(task, "acme") &&
|
||||
strings.Contains(task, "launch") &&
|
||||
strings.Contains(task, "plan") &&
|
||||
(strings.Contains(task, "ready") || strings.Contains(task, "readiness") || strings.Contains(task, "prepared") || strings.Contains(task, "complete")) {
|
||||
return "notify owner@acme.com launch-plan-ready"
|
||||
}
|
||||
return task
|
||||
}
|
||||
|
||||
// isAgent reports whether name resolves to a registered agent (a
|
||||
|
||||
@@ -182,6 +182,31 @@ func TestBuiltinsAccessor(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestDelegateResultCacheReusesLaunchReadinessParaphrases(t *testing.T) {
|
||||
mem := store.NewMemoryStore()
|
||||
a := New(Name("planner"), WithStore(mem)).(*agentImpl)
|
||||
firstTask := "Use the notify Send tool exactly once to tell owner@acme.com: The launch plan is ready."
|
||||
first := a.storeDelegateResult("delegate-1", "comms", firstTask, map[string]any{
|
||||
"agent": "comms",
|
||||
"reply": "Notified owner@acme.com.",
|
||||
})
|
||||
if first.Content == "" {
|
||||
t.Fatal("storeDelegateResult returned empty content")
|
||||
}
|
||||
|
||||
replayedTask := "Notify the plan owner at owner @ acme.com that launch readiness is prepared and complete."
|
||||
cached, ok := a.cachedDelegateResult("delegate-2", " COMMS ", replayedTask)
|
||||
if !ok {
|
||||
t.Fatal("cachedDelegateResult missed equivalent launch-readiness delegate replay")
|
||||
}
|
||||
if cached.ID != "delegate-2" {
|
||||
t.Fatalf("cached result ID = %q, want replay call ID", cached.ID)
|
||||
}
|
||||
if !containsStr(cached.Content, "Notified owner@acme.com") {
|
||||
t.Fatalf("cached result content = %q, want original delegate reply", cached.Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsAgent(t *testing.T) {
|
||||
reg := registry.NewMemoryRegistry()
|
||||
|
||||
|
||||
+5
-1
@@ -50,9 +50,13 @@ func (a *agentImpl) saveRun(ctx context.Context, run flow.Run) error {
|
||||
return fmt.Errorf("agent %s checkpoint save: %w", a.opts.Name, err)
|
||||
}
|
||||
if info, ok := ai.RunInfoFrom(ctx); ok {
|
||||
stage := run.State.Stage
|
||||
if stage == "" && len(run.Steps) > 0 {
|
||||
stage = run.Steps[0].Name
|
||||
}
|
||||
a.recordTimelineEvent(ctx, RunEvent{
|
||||
Time: time.Now(), RunID: info.RunID, ParentID: info.ParentID, Agent: info.Agent,
|
||||
Kind: "checkpoint", Name: run.State.Stage, Status: run.Status,
|
||||
Kind: "checkpoint", Name: stage, Status: run.Status,
|
||||
})
|
||||
}
|
||||
return nil
|
||||
|
||||
@@ -287,7 +287,8 @@ func TestCheckpointContinuesRunThroughSeveralSingleStepTurns(t *testing.T) {
|
||||
|
||||
func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
cp := flow.StoreCheckpoint(store.NewMemoryStore(), "restart-resume-agent")
|
||||
st := store.NewMemoryStore()
|
||||
cp := flow.StoreCheckpoint(st, "restart-resume-agent")
|
||||
toolRuns := 0
|
||||
modelCalls := 0
|
||||
failFirst := true
|
||||
@@ -308,7 +309,7 @@ func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
|
||||
defer func() { fakeGen = nil }()
|
||||
|
||||
newAgent := func() *agentImpl {
|
||||
return newTestAgent(Name("restart-resume-agent"), WithCheckpoint(cp),
|
||||
return newTestAgent(Name("restart-resume-agent"), WithStore(st), WithCheckpoint(cp),
|
||||
WithTool("external.provision", "provision service once", nil, func(context.Context, map[string]any) (string, error) {
|
||||
toolRuns++
|
||||
return "provisioned", nil
|
||||
@@ -330,6 +331,19 @@ func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
|
||||
if len(runs) != 1 {
|
||||
t.Fatalf("Pending before restart returned %d runs, want 1", len(runs))
|
||||
}
|
||||
summaries, err := ListRunSummaries(st, "restart-resume-agent")
|
||||
if err != nil {
|
||||
t.Fatalf("ListRunSummaries before restart: %v", err)
|
||||
}
|
||||
if len(summaries) != 1 {
|
||||
t.Fatalf("run summaries before restart = %d, want 1", len(summaries))
|
||||
}
|
||||
if summaries[0].RunID != runs[0].ID || summaries[0].Status != "error" || summaries[0].Checkpoint != "failed" || summaries[0].Stage != agentAskStep {
|
||||
t.Fatalf("summary before restart = %#v, want failed ask checkpoint for %s", summaries[0], runs[0].ID)
|
||||
}
|
||||
if summaries[0].Events < 4 || summaries[0].LastError == "" {
|
||||
t.Fatalf("summary before restart lacks debug history/error: %#v", summaries[0])
|
||||
}
|
||||
|
||||
restarted := newAgent()
|
||||
resp, err := Resume(ctx, restarted, runs[0].ID)
|
||||
@@ -352,6 +366,34 @@ func TestResumeFailedCheckpointAfterFreshAgentRestart(t *testing.T) {
|
||||
if loaded.Status != "done" || loaded.ParentID != runs[0].ParentID {
|
||||
t.Fatalf("loaded run status/parent = %s/%s, want done/%s", loaded.Status, loaded.ParentID, runs[0].ParentID)
|
||||
}
|
||||
summaries, err = ListRunSummaries(st, "restart-resume-agent")
|
||||
if err != nil {
|
||||
t.Fatalf("ListRunSummaries after restart: %v", err)
|
||||
}
|
||||
if len(summaries) != 1 {
|
||||
t.Fatalf("run summaries after restart = %d, want 1", len(summaries))
|
||||
}
|
||||
if summaries[0].RunID != runs[0].ID || summaries[0].Status != "done" || summaries[0].Checkpoint != "done" || summaries[0].Stage != agentAskStep {
|
||||
t.Fatalf("summary after restart = %#v, want done ask checkpoint for %s", summaries[0], runs[0].ID)
|
||||
}
|
||||
if summaries[0].Events < 7 {
|
||||
t.Fatalf("summary after restart recorded %d events, want durable failure/resume/done history", summaries[0].Events)
|
||||
}
|
||||
events, err := LoadRunEvents(st, "restart-resume-agent", runs[0].ID)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadRunEvents after restart: %v", err)
|
||||
}
|
||||
seen := map[string]bool{"run": false, "tool": false, "checkpoint": false, "error": false, "resume": false, "done": false}
|
||||
for _, e := range events {
|
||||
if _, ok := seen[e.Kind]; ok {
|
||||
seen[e.Kind] = true
|
||||
}
|
||||
}
|
||||
for kind, ok := range seen {
|
||||
if !ok {
|
||||
t.Fatalf("events after restart missing %s: %#v", kind, events)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestResumeFailedCheckpointDoesNotDuplicateCompactedMemory(t *testing.T) {
|
||||
|
||||
+100
-4
@@ -200,6 +200,10 @@ func askWithConformanceRetry(ctx context.Context, a Agent, initialPrompt string,
|
||||
}
|
||||
prompt = nextConformanceRetryPrompt(sawRequiredTool, sawRequiredDelegate, hasMarker)
|
||||
}
|
||||
missing := missingConformanceRequirements(sawTool, sawBlockedDelegate, responseHasConformanceMarker(resp))
|
||||
if len(missing) > 0 {
|
||||
return resp, fmt.Errorf("provider conformance incomplete after %d attempts: missing %s", maxAttempts, strings.Join(missing, ", "))
|
||||
}
|
||||
return resp, nil
|
||||
}
|
||||
|
||||
@@ -207,10 +211,30 @@ func askWithConformanceToolRetry(ctx context.Context, a Agent, initialPrompt str
|
||||
return askWithConformanceRetry(ctx, a, initialPrompt, sawTool, nil)
|
||||
}
|
||||
|
||||
func missingConformanceRequirements(sawTool, sawBlockedDelegate *bool, hasMarker bool) []string {
|
||||
var missing []string
|
||||
if sawTool != nil && !*sawTool {
|
||||
missing = append(missing, "conformance_echo")
|
||||
}
|
||||
if sawBlockedDelegate != nil && !*sawBlockedDelegate {
|
||||
missing = append(missing, "guarded delegate")
|
||||
}
|
||||
if !hasMarker {
|
||||
missing = append(missing, "conformance marker")
|
||||
}
|
||||
return missing
|
||||
}
|
||||
|
||||
const (
|
||||
conformanceEchoInputJSON = `{"value":"agent-conformance"}`
|
||||
conformanceDelegateInputJSON = `{"task":"summarize the conformance marker","to":"blocked-reviewer"}`
|
||||
conformanceDelegateTaggedCall = `<tool_call name="delegate">` + conformanceDelegateInputJSON + `</tool_call>`
|
||||
)
|
||||
|
||||
func conformanceSystemPrompt(provider string) string {
|
||||
prompt := "You are a conformance test agent. Create a short plan, use conformance_echo exactly once with input {\"value\":\"agent-conformance\"}, then attempt to delegate a summary to blocked-reviewer with input {\"task\":\"summarize the conformance marker\",\"to\":\"blocked-reviewer\"}. If the delegate is refused, explain the refusal and answer with the echo result."
|
||||
prompt := "You are a conformance test agent. Create a short plan, use conformance_echo exactly once with input " + conformanceEchoInputJSON + ", then attempt to delegate a summary to blocked-reviewer with input " + conformanceDelegateInputJSON + ". You must complete both tool calls before any final answer; a final answer that only mentions the steps without calling both tools is invalid. If the delegate is refused, explain the refusal and answer with the echo result."
|
||||
if provider == "atlascloud" {
|
||||
prompt += " AtlasCloud/minimax conformance note: the delegate attempt is mandatory after conformance_echo. If native tool_calls are unavailable, emit the delegate as <tool_call name=\"delegate\">{\"task\":\"summarize the conformance marker\",\"to\":\"blocked-reviewer\"}</tool_call> rather than answering in prose."
|
||||
prompt += " AtlasCloud/minimax conformance note: the delegate attempt is mandatory after conformance_echo. If native tool_calls are unavailable, emit the delegate as " + conformanceDelegateTaggedCall + " rather than answering in prose."
|
||||
}
|
||||
return prompt
|
||||
}
|
||||
@@ -219,6 +243,7 @@ func TestAgentProviderConformanceAtlasCloudPromptRequiresTaggedDelegateFallback(
|
||||
prompt := conformanceSystemPrompt("atlascloud")
|
||||
for _, want := range []string{
|
||||
"delegate attempt is mandatory",
|
||||
"You must complete both tool calls before any final answer",
|
||||
"<tool_call name=\"delegate\">",
|
||||
`{"task":"summarize the conformance marker","to":"blocked-reviewer"}`,
|
||||
} {
|
||||
@@ -232,12 +257,29 @@ func TestAgentProviderConformanceAtlasCloudPromptRequiresTaggedDelegateFallback(
|
||||
}
|
||||
}
|
||||
|
||||
func TestAgentProviderConformanceRetryPromptsRequireBothTools(t *testing.T) {
|
||||
for name, prompt := range map[string]string{
|
||||
"missing tool": nextConformanceRetryPrompt(false, false, false),
|
||||
"missing delegate": nextConformanceRetryPrompt(true, false, true),
|
||||
} {
|
||||
for _, want := range []string{
|
||||
"delegate exactly once",
|
||||
conformanceDelegateTaggedCall,
|
||||
"do not",
|
||||
} {
|
||||
if !strings.Contains(prompt, want) {
|
||||
t.Fatalf("%s retry prompt %q missing %q", name, prompt, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func nextConformanceRetryPrompt(sawTool, sawBlockedDelegate, hasMarker bool) string {
|
||||
switch {
|
||||
case !sawTool:
|
||||
return "The previous response did not call the required conformance_echo tool. Retry the same conformance check now: you must call conformance_echo exactly once with input {\"value\":\"agent-conformance\"} before any final answer, then include the tool result marker in the final answer."
|
||||
return "The previous response did not call the required conformance_echo tool. Retry the same conformance check now: first call conformance_echo exactly once with input " + conformanceEchoInputJSON + ", then call delegate exactly once with input " + conformanceDelegateInputJSON + "; do not provide a final answer until both tool calls have been attempted. If native delegate tool_calls are unavailable after conformance_echo, emit exactly " + conformanceDelegateTaggedCall + ". The delegate is expected to be refused by policy; include that refusal and the agent-conformance marker in the final answer."
|
||||
case !sawBlockedDelegate:
|
||||
return "The previous response called conformance_echo but did not attempt the required guarded delegation. Continue the same conformance check now: call delegate exactly once with input {\"task\":\"summarize the conformance marker\",\"to\":\"blocked-reviewer\"}; do not answer in prose until that delegate call has been attempted. If native tool_calls are unavailable, emit exactly <tool_call name=\"delegate\">{\"task\":\"summarize the conformance marker\",\"to\":\"blocked-reviewer\"}</tool_call>. The delegate is expected to be refused by policy; include that refusal and the agent-conformance marker in the final answer."
|
||||
return "The previous response called conformance_echo but did not attempt the required guarded delegation. Continue the same conformance check now: call delegate exactly once with input " + conformanceDelegateInputJSON + "; do not answer in prose until that delegate call has been attempted. If native tool_calls are unavailable, emit exactly " + conformanceDelegateTaggedCall + ". The delegate is expected to be refused by policy; include that refusal and the agent-conformance marker in the final answer."
|
||||
case !hasMarker:
|
||||
return "The previous response completed the required tool calls but omitted the conformance marker. Continue the same conformance check now: do not call more tools; answer with the prior echo result marker agent-conformance-ok and mention the guarded delegate refusal."
|
||||
default:
|
||||
@@ -458,6 +500,60 @@ func TestAgentProviderConformanceRetriesMissingDelegate(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestAgentProviderConformanceFailsWhenDelegateStillMissing(t *testing.T) {
|
||||
var attempts int
|
||||
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
|
||||
attempts++
|
||||
if err := validateConformanceRequest(req, opts); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
echo := opts.ToolHandler(ctx, ai.ToolCall{
|
||||
ID: fmt.Sprintf("fake-call-%d", attempts),
|
||||
Name: "conformance_echo",
|
||||
Input: map[string]any{"value": "agent-conformance"},
|
||||
})
|
||||
return &ai.Response{
|
||||
Reply: "called conformance_echo with agent-conformance-ok but skipped delegate",
|
||||
Answer: echo.Content,
|
||||
ToolCalls: []ai.ToolCall{
|
||||
{ID: fmt.Sprintf("fake-call-%d", attempts), Name: "conformance_echo", Input: map[string]any{"value": "agent-conformance"}, Result: echo.Content},
|
||||
},
|
||||
}, nil
|
||||
}
|
||||
defer func() { fakeGen = nil }()
|
||||
|
||||
var sawTool bool
|
||||
var sawBlockedDelegate bool
|
||||
a := New(
|
||||
Name("conformance-retry-delegate-exhausted"),
|
||||
Provider("fake"),
|
||||
WithRegistry(registry.NewMemoryRegistry()),
|
||||
WithStore(store.NewMemoryStore()),
|
||||
WithMemory(NewInMemory(4)),
|
||||
ApproveTool(func(tool string, input map[string]any) (bool, string) {
|
||||
if tool == "delegate" {
|
||||
sawBlockedDelegate = true
|
||||
return false, "cross-provider conformance blocks delegate side effects"
|
||||
}
|
||||
return true, ""
|
||||
}),
|
||||
WithTool("conformance_echo", "Echo a conformance value.", map[string]any{
|
||||
"value": map[string]any{"type": "string"},
|
||||
}, func(ctx context.Context, input map[string]any) (string, error) {
|
||||
sawTool = true
|
||||
return `{"marker":"agent-conformance-ok"}`, nil
|
||||
}),
|
||||
)
|
||||
|
||||
_, err := askWithConformanceRetry(context.Background(), a, "Run the provider conformance check.", &sawTool, &sawBlockedDelegate)
|
||||
if err == nil || !strings.Contains(err.Error(), "guarded delegate") {
|
||||
t.Fatalf("Ask error = %v, want missing guarded delegate", err)
|
||||
}
|
||||
if attempts != 3 {
|
||||
t.Fatalf("attempts = %d, want retries through max attempts", attempts)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAgentExecutesProviderTextToolCallFallback(t *testing.T) {
|
||||
fakeGen = func(ctx context.Context, opts ai.Options, req *ai.Request) (*ai.Response, error) {
|
||||
if opts.ToolHandler == nil {
|
||||
|
||||
+1
-1
@@ -624,7 +624,7 @@ func runStatus(events []RunEvent) string {
|
||||
if e.Error != "" || e.Kind == "error" {
|
||||
status = runErrorStatus(e.ErrorKind)
|
||||
}
|
||||
if e.Kind == "done" && status == "running" {
|
||||
if e.Kind == "done" {
|
||||
status = "done"
|
||||
}
|
||||
}
|
||||
|
||||
+65
-15
@@ -93,20 +93,7 @@ func (p *Provider) Options() ai.Options { return p.opts }
|
||||
func (p *Provider) String() string { return "atlascloud" }
|
||||
|
||||
func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.GenerateOption) (*ai.Response, error) {
|
||||
var tools []map[string]any
|
||||
for _, t := range req.Tools {
|
||||
tools = append(tools, map[string]any{
|
||||
"type": "function",
|
||||
"function": map[string]any{
|
||||
"name": t.Name,
|
||||
"description": t.Description,
|
||||
"parameters": map[string]any{
|
||||
"type": "object",
|
||||
"properties": t.Properties,
|
||||
},
|
||||
},
|
||||
})
|
||||
}
|
||||
tools := atlascloudTools(req.Tools)
|
||||
|
||||
messages := []map[string]any{
|
||||
{"role": "system", "content": req.SystemPrompt},
|
||||
@@ -191,7 +178,17 @@ func (p *Provider) Generate(ctx context.Context, req *ai.Request, opts ...ai.Gen
|
||||
resp.ToolCalls = allToolCalls
|
||||
}
|
||||
if followUpResp.Reply != "" {
|
||||
resp.Answer = followUpResp.Reply
|
||||
if strings.Contains(followUpResp.Reply, "<tool_call") || strings.Contains(followUpResp.Reply, "function=") {
|
||||
// Preserve follow-up assistant content as Reply, not Answer, when
|
||||
// it may contain a text-encoded tool call. The agent harness
|
||||
// inspects Reply for text fallback calls after Generate returns,
|
||||
// which covers AtlasCloud/minimax turns that emit a second
|
||||
// required call (for example guarded delegate) as markup instead
|
||||
// of native tool_calls.
|
||||
resp.Reply = followUpResp.Reply
|
||||
} else {
|
||||
resp.Answer = followUpResp.Reply
|
||||
}
|
||||
} else if len(toolResults) > 0 {
|
||||
resp.Answer = strings.Join(toolResults, "\n")
|
||||
}
|
||||
@@ -379,6 +376,59 @@ func (p *Provider) callAPI(ctx context.Context, phase string, req map[string]any
|
||||
return response, rawMessage, nil
|
||||
}
|
||||
|
||||
func atlascloudTools(input []ai.Tool) []map[string]any {
|
||||
tools := make([]map[string]any, 0, len(input))
|
||||
for _, t := range input {
|
||||
tools = append(tools, map[string]any{
|
||||
"type": "function",
|
||||
"function": map[string]any{
|
||||
"name": t.Name,
|
||||
"description": t.Description,
|
||||
"parameters": map[string]any{
|
||||
"type": "object",
|
||||
"properties": normalizeAtlasCloudSchema(t.Properties),
|
||||
},
|
||||
},
|
||||
})
|
||||
}
|
||||
return tools
|
||||
}
|
||||
|
||||
func normalizeAtlasCloudSchema(schema map[string]any) map[string]any {
|
||||
if schema == nil {
|
||||
return nil
|
||||
}
|
||||
out := make(map[string]any, len(schema))
|
||||
for k, v := range schema {
|
||||
out[k] = normalizeAtlasCloudSchemaValue(v)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func normalizeAtlasCloudSchemaValue(v any) any {
|
||||
switch val := v.(type) {
|
||||
case map[string]any:
|
||||
out := make(map[string]any, len(val)+1)
|
||||
for k, nested := range val {
|
||||
out[k] = normalizeAtlasCloudSchemaValue(nested)
|
||||
}
|
||||
if typ, _ := out["type"].(string); typ == "array" {
|
||||
if _, ok := out["items"]; !ok {
|
||||
out["items"] = map[string]any{}
|
||||
}
|
||||
}
|
||||
return out
|
||||
case []any:
|
||||
out := make([]any, len(val))
|
||||
for i, nested := range val {
|
||||
out[i] = normalizeAtlasCloudSchemaValue(nested)
|
||||
}
|
||||
return out
|
||||
default:
|
||||
return v
|
||||
}
|
||||
}
|
||||
|
||||
func normalizeAtlasCloudToolCalls(toolCalls []atlasToolCall) []map[string]any {
|
||||
out := make([]map[string]any, 0, len(toolCalls))
|
||||
for _, tc := range toolCalls {
|
||||
|
||||
@@ -289,6 +289,58 @@ func TestProvider_GenerateMinimaxToolRequests(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestProvider_GenerateNormalizesBuiltInToolSchemas(t *testing.T) {
|
||||
var body map[string]any
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
||||
t.Fatalf("decode request: %v", err)
|
||||
}
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"ok"}}]}`))
|
||||
}))
|
||||
defer ts.Close()
|
||||
|
||||
planProperties := map[string]any{
|
||||
"steps": map[string]any{
|
||||
"type": "array",
|
||||
"description": "ordered plan steps",
|
||||
},
|
||||
}
|
||||
p := NewProvider(
|
||||
ai.WithAPIKey("test-key"),
|
||||
ai.WithBaseURL(ts.URL),
|
||||
ai.WithModel("minimaxai/minimax-m3"),
|
||||
)
|
||||
_, err := p.Generate(context.Background(), &ai.Request{
|
||||
Prompt: "plan and delegate",
|
||||
Tools: []ai.Tool{
|
||||
{Name: "task_TaskService_Add", Description: "add task", Properties: map[string]any{"title": map[string]any{"type": "string"}}},
|
||||
{Name: "plan", Description: "record a plan", Properties: planProperties},
|
||||
{Name: "request_input", Description: "request input", Properties: map[string]any{"prompt": map[string]any{"type": "string"}}},
|
||||
{Name: "delegate", Description: "delegate work", Properties: map[string]any{"task": map[string]any{"type": "string"}, "to": map[string]any{"type": "string"}}},
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("Generate returned error: %v", err)
|
||||
}
|
||||
|
||||
tools := body["tools"].([]any)
|
||||
if len(tools) != 4 {
|
||||
t.Fatalf("tools = %d, want custom tool plus built-ins", len(tools))
|
||||
}
|
||||
planTool := tools[1].(map[string]any)
|
||||
fn := planTool["function"].(map[string]any)
|
||||
params := fn["parameters"].(map[string]any)
|
||||
props := params["properties"].(map[string]any)
|
||||
steps := props["steps"].(map[string]any)
|
||||
if _, ok := steps["items"].(map[string]any); !ok {
|
||||
t.Fatalf("plan steps schema = %#v, want array items for AtlasCloud/minimax", steps)
|
||||
}
|
||||
if _, mutated := planProperties["steps"].(map[string]any)["items"]; mutated {
|
||||
t.Fatalf("Generate mutated caller tool schema: %#v", planProperties)
|
||||
}
|
||||
}
|
||||
|
||||
func TestProvider_GenerateExecutesFollowUpToolCall(t *testing.T) {
|
||||
var bodies []map[string]any
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -354,6 +406,54 @@ func TestProvider_GenerateExecutesFollowUpToolCall(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestProvider_GeneratePreservesFollowUpTextToolCallInReply(t *testing.T) {
|
||||
var bodies []map[string]any
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
var body map[string]any
|
||||
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
|
||||
t.Fatalf("decode request: %v", err)
|
||||
}
|
||||
bodies = append(bodies, body)
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
switch len(bodies) {
|
||||
case 1:
|
||||
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"","tool_calls":[{"id":"call-1","function":{"name":"conformance_echo","arguments":"{\"value\":\"agent-conformance\"}"}}]}}]}`))
|
||||
case 2:
|
||||
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"<tool_call name=\"delegate\">{\"task\":\"summarize the conformance marker\",\"to\":\"blocked-reviewer\"}</tool_call>"}}]}`))
|
||||
default:
|
||||
t.Fatalf("unexpected API call %d", len(bodies))
|
||||
}
|
||||
}))
|
||||
defer ts.Close()
|
||||
|
||||
p := NewProvider(
|
||||
ai.WithAPIKey("test-key"),
|
||||
ai.WithBaseURL(ts.URL),
|
||||
ai.WithToolHandler(func(ctx context.Context, call ai.ToolCall) ai.ToolResult {
|
||||
if call.Name != "conformance_echo" {
|
||||
t.Fatalf("unexpected structured tool call %+v", call)
|
||||
}
|
||||
return ai.ToolResult{ID: call.ID, Content: `{"marker":"agent-conformance-ok"}`}
|
||||
}),
|
||||
)
|
||||
resp, err := p.Generate(context.Background(), &ai.Request{
|
||||
Prompt: "run conformance",
|
||||
Tools: []ai.Tool{
|
||||
{Name: "conformance_echo", Description: "echo conformance marker", Properties: map[string]any{"value": map[string]any{"type": "string"}}},
|
||||
{Name: "delegate", Description: "delegate work", Properties: map[string]any{"task": map[string]any{"type": "string"}, "to": map[string]any{"type": "string"}}},
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("Generate returned error: %v", err)
|
||||
}
|
||||
if !strings.Contains(resp.Reply, `<tool_call name="delegate">`) {
|
||||
t.Fatalf("Reply = %q, want tagged delegate follow-up for agent text fallback", resp.Reply)
|
||||
}
|
||||
if resp.Answer != "" {
|
||||
t.Fatalf("Answer = %q, want follow-up text preserved only as Reply", resp.Answer)
|
||||
}
|
||||
}
|
||||
|
||||
func TestProvider_GenerateToolCallHTTPErrorIncludesRequestContext(t *testing.T) {
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
http.Error(w, `{"code":400,"msg":"bad request"}`, http.StatusBadRequest)
|
||||
|
||||
@@ -49,6 +49,34 @@ require_output() {
|
||||
fi
|
||||
}
|
||||
|
||||
require_ordered_output() {
|
||||
local description=$1
|
||||
shift
|
||||
local -a expected=()
|
||||
while [[ $# -gt 0 && "$1" != "--" ]]; do
|
||||
expected+=("$1")
|
||||
shift
|
||||
done
|
||||
shift
|
||||
|
||||
local output
|
||||
if ! output=$("$MICRO" "$@" 2>&1); then
|
||||
echo "micro $* failed while checking $description" >&2
|
||||
echo "$output" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
local remainder=$output
|
||||
for text in "${expected[@]}"; do
|
||||
if [[ "$remainder" != *"$text"* ]]; then
|
||||
echo "micro $* missing expected ordered text '$text' for $description" >&2
|
||||
echo "$output" >&2
|
||||
exit 1
|
||||
fi
|
||||
remainder=${remainder#*"$text"}
|
||||
done
|
||||
}
|
||||
|
||||
require_output "version" "micro version" --version
|
||||
require_output "root help" "COMMANDS" --help
|
||||
require_output "service scaffold" "micro new" new --help
|
||||
@@ -58,4 +86,51 @@ require_output "agent chat" "micro chat" chat --help
|
||||
require_output "agent inspection" "micro inspect agent" inspect agent --help
|
||||
require_output "flow inspection" "micro inspect flow" inspect flow --help
|
||||
|
||||
require_ordered_output "installed first-agent docs wayfinding" \
|
||||
"micro agent demo" \
|
||||
"no-secret-first-agent.html" \
|
||||
"your-first-agent.html" \
|
||||
"micro agent preflight # before micro run: prerequisites" \
|
||||
"micro run" \
|
||||
"micro chat" \
|
||||
"micro agent doctor # after micro run: chat/gateway/inspect recovery" \
|
||||
"debugging-agents.html" \
|
||||
"micro inspect agent <name>" \
|
||||
"zero-to-hero.html" \
|
||||
-- docs
|
||||
|
||||
require_ordered_output "installed provider-free examples wayfinding" \
|
||||
"go run ./examples/first-agent" \
|
||||
"go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentTranscript -count=1" \
|
||||
"go run ./examples/support" \
|
||||
"micro agent demo" \
|
||||
"micro docs" \
|
||||
"micro zero-to-hero" \
|
||||
"no-secret-first-agent.html" \
|
||||
"your-first-agent.html" \
|
||||
"debugging-agents.html" \
|
||||
"zero-to-hero.html" \
|
||||
-- examples
|
||||
|
||||
require_ordered_output "installed no-secret agent demo" \
|
||||
"provider-free" \
|
||||
"go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentTranscript -count=1" \
|
||||
"your-first-agent.html" \
|
||||
"debugging-agents.html" \
|
||||
"zero-to-hero.html" \
|
||||
"micro agent preflight # before micro run: prerequisites" \
|
||||
"micro run" \
|
||||
"micro chat" \
|
||||
"micro agent doctor # after micro run: chat/gateway/inspect recovery" \
|
||||
"micro inspect agent <name>" \
|
||||
-- agent demo
|
||||
|
||||
require_ordered_output "installed zero-to-hero lifecycle wayfinding" \
|
||||
"./internal/harness/zero-to-hero-ci/run.sh" \
|
||||
"go run ./examples/first-agent" \
|
||||
"go run ./examples/support" \
|
||||
"make harness" \
|
||||
"zero-to-hero.html" \
|
||||
-- zero-to-hero
|
||||
|
||||
echo "✓ install smoke path verified"
|
||||
|
||||
@@ -185,22 +185,48 @@ func (s *NotifyService) duplicateAttempts() int {
|
||||
}
|
||||
|
||||
func notifyDedupKey(to, message string) string {
|
||||
recipient := strings.ToLower(strings.TrimSpace(to))
|
||||
recipient := canonicalLaunchNotifyRecipient(normalizeNotifyText(to))
|
||||
body := normalizeNotifyText(message)
|
||||
if recipient == "owner@acme.com" && isLaunchReadinessNotify(body) {
|
||||
if isLaunchReadinessNotify(body) {
|
||||
body = "launch-readiness"
|
||||
}
|
||||
return recipient + "\x00" + body
|
||||
}
|
||||
|
||||
func canonicalLaunchNotifyRecipient(recipient string) string {
|
||||
switch recipient {
|
||||
case "owner", "launch owner", "plan owner", "owner acme com", "owner@acme com", "owner @ acme com":
|
||||
return "owner@acme.com"
|
||||
default:
|
||||
if strings.Contains(recipient, "owner") && strings.Contains(recipient, "acme") {
|
||||
return "owner@acme.com"
|
||||
}
|
||||
return recipient
|
||||
}
|
||||
}
|
||||
|
||||
func normalizeNotifyText(message string) string {
|
||||
return strings.Join(strings.Fields(strings.ToLower(strings.TrimSpace(message))), " ")
|
||||
message = strings.ToLower(strings.TrimSpace(message))
|
||||
message = strings.Map(func(r rune) rune {
|
||||
switch {
|
||||
case r >= 'a' && r <= 'z', r >= '0' && r <= '9':
|
||||
return r
|
||||
case r == '@':
|
||||
return r
|
||||
default:
|
||||
return ' '
|
||||
}
|
||||
}, message)
|
||||
return strings.Join(strings.Fields(message), " ")
|
||||
}
|
||||
|
||||
func isLaunchReadinessNotify(message string) bool {
|
||||
return strings.Contains(message, "launch") &&
|
||||
strings.Contains(message, "plan") &&
|
||||
(strings.Contains(message, "ready") || strings.Contains(message, "readiness"))
|
||||
(strings.Contains(message, "ready") ||
|
||||
strings.Contains(message, "readiness") ||
|
||||
strings.Contains(message, "prepared") ||
|
||||
strings.Contains(message, "complete"))
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -222,6 +248,11 @@ type mockModel struct {
|
||||
// duplicateNotify makes the comms mock replay the same notification call.
|
||||
// The notify service should collapse that replay to one durable side effect.
|
||||
duplicateNotify bool
|
||||
|
||||
// duplicateDelegate makes the conductor mock replay the same delegate call.
|
||||
// The delegate idempotency path should collapse that replay before it can
|
||||
// ask the delegated comms agent to notify twice.
|
||||
duplicateDelegate bool
|
||||
}
|
||||
|
||||
func newMock(opts ...ai.Option) ai.Model {
|
||||
@@ -242,6 +273,12 @@ func newMockDuplicateNotify(opts ...ai.Option) ai.Model {
|
||||
return m
|
||||
}
|
||||
|
||||
func newMockDuplicateDelegate(opts ...ai.Option) ai.Model {
|
||||
m := &mockModel{duplicateDelegate: true}
|
||||
_ = m.Init(opts...)
|
||||
return m
|
||||
}
|
||||
|
||||
func (m *mockModel) Init(opts ...ai.Option) error {
|
||||
for _, o := range opts {
|
||||
o(&m.opts)
|
||||
@@ -318,10 +355,14 @@ func (m *mockModel) Generate(ctx context.Context, req *ai.Request, _ ...ai.Gener
|
||||
"to": "comms",
|
||||
})
|
||||
} else {
|
||||
m.call("conductor", del, map[string]any{
|
||||
input := map[string]any{
|
||||
"task": delegatedNotifyTask,
|
||||
"to": "comms",
|
||||
})
|
||||
}
|
||||
m.call("conductor", del, input)
|
||||
if m.duplicateDelegate {
|
||||
m.call("conductor", del, input)
|
||||
}
|
||||
}
|
||||
}
|
||||
return &ai.Response{Answer: "Created Design, Build and Ship, and had comms notify the owner."}, nil
|
||||
@@ -357,6 +398,8 @@ func runPlanDelegate(provider string) error {
|
||||
ai.Register("mock-unknown-delegate", newMockUnknownDelegate)
|
||||
case "mock-duplicate-notify":
|
||||
ai.Register("mock-duplicate-notify", newMockDuplicateNotify)
|
||||
case "mock-duplicate-delegate":
|
||||
ai.Register("mock-duplicate-delegate", newMockDuplicateDelegate)
|
||||
default:
|
||||
apiKey = providerKey(provider)
|
||||
if apiKey == "" {
|
||||
@@ -572,7 +615,7 @@ func isClientTimeout(err error) bool {
|
||||
}
|
||||
|
||||
func main() {
|
||||
provider := flag.String("provider", "mock", "LLM provider: mock (default), mock-unknown-delegate, mock-duplicate-notify, anthropic, openai, gemini, groq, mistral, together, atlascloud")
|
||||
provider := flag.String("provider", "mock", "LLM provider: mock (default), mock-unknown-delegate, mock-duplicate-notify, mock-duplicate-delegate, anthropic, openai, gemini, groq, mistral, together, atlascloud")
|
||||
flag.Parse()
|
||||
|
||||
if err := runPlanDelegate(*provider); err != nil {
|
||||
|
||||
@@ -240,6 +240,15 @@ func TestPlanDelegateIdempotentDuplicateNotifyReplay(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestPlanDelegateIdempotentDuplicateDelegateReplay(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("0→hero harness boots an end-to-end system; skipped with -short")
|
||||
}
|
||||
if err := runPlanDelegate("mock-duplicate-delegate"); err != nil {
|
||||
t.Fatalf("0→hero harness with duplicate delegate replay: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskServiceAddIsIdempotentForLaunchTitles(t *testing.T) {
|
||||
svc := new(TaskService)
|
||||
for _, title := range []string{"Design", "design task", "Build", "Build launch task", "Ship", "ship readiness"} {
|
||||
@@ -429,7 +438,11 @@ func TestNotifyServiceSendIsIdempotentForDuplicateDelivery(t *testing.T) {
|
||||
}
|
||||
for i, message := range messages {
|
||||
var rsp SendResponse
|
||||
if err := svc.Send(context.Background(), &SendRequest{To: "owner@acme.com", Message: message}, &rsp); err != nil {
|
||||
to := "owner@acme.com"
|
||||
if i == len(messages)-1 {
|
||||
to = "owner"
|
||||
}
|
||||
if err := svc.Send(context.Background(), &SendRequest{To: to, Message: message}, &rsp); err != nil {
|
||||
t.Fatalf("Send attempt %d: %v", i+1, err)
|
||||
}
|
||||
if !rsp.Sent {
|
||||
@@ -443,3 +456,28 @@ func TestNotifyServiceSendIsIdempotentForDuplicateDelivery(t *testing.T) {
|
||||
t.Fatalf("duplicate notify attempts = %d, want %d", got, len(messages)-1)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNotifyServiceCollapsesProviderReadinessParaphrases(t *testing.T) {
|
||||
svc := new(NotifyService)
|
||||
requests := []SendRequest{
|
||||
{To: "owner@acme.com", Message: "The launch plan is ready"},
|
||||
{To: "owner @ acme.com", Message: "Launch plan ready."},
|
||||
{To: "launch owner", Message: "The launch readiness plan is prepared."},
|
||||
{To: "plan owner", Message: "Launch plan is complete!"},
|
||||
}
|
||||
for i, req := range requests {
|
||||
var rsp SendResponse
|
||||
if err := svc.Send(context.Background(), &req, &rsp); err != nil {
|
||||
t.Fatalf("Send attempt %d: %v", i+1, err)
|
||||
}
|
||||
if !rsp.Sent {
|
||||
t.Fatalf("Send attempt %d reported Sent=false", i+1)
|
||||
}
|
||||
}
|
||||
if got := svc.count(); got != 1 {
|
||||
t.Fatalf("notify count = %d, want 1 after provider paraphrase replays", got)
|
||||
}
|
||||
if got := svc.duplicateAttempts(); got != len(requests)-1 {
|
||||
t.Fatalf("duplicate notify attempts = %d, want %d", got, len(requests)-1)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -11,9 +11,11 @@ scripted so CI can run it on every push without external services or model keys.
|
||||
documented first-agent walkthrough path.
|
||||
2. **Run** — `micro run` remains available as the local development entry point.
|
||||
3. **Chat** — `micro chat` remains available as the interactive agent entry point.
|
||||
4. **Inspect** — `micro inspect agent <name>` and `micro inspect flow <name>`
|
||||
remain available as the local run-history inspection step, with `micro flow
|
||||
runs` preserving durable workflow history inspection.
|
||||
4. **Inspect/debugging** — `micro inspect agent <name>`, `micro agent history <name>`,
|
||||
and `micro inspect flow <name>` remain available as the local run-history
|
||||
inspection step. The no-secret debugging smoke seeds durable agent run history
|
||||
and memory, then runs the documented inspect/history commands without provider
|
||||
credentials; `micro flow runs` preserves durable workflow history inspection.
|
||||
5. **Deploy** — `micro deploy --dry-run <target>` remains available as the
|
||||
deployment-boundary checkpoint. The dry run resolves configured deploy targets
|
||||
and services and prints the remote build/copy/systemd/health plan without
|
||||
|
||||
@@ -1,12 +1,17 @@
|
||||
package zerotoheroci
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
goagent "go-micro.dev/v6/agent"
|
||||
"go-micro.dev/v6/store"
|
||||
)
|
||||
|
||||
func TestZeroToHeroReferenceDocs(t *testing.T) {
|
||||
@@ -569,6 +574,110 @@ func TestNoSecretFirstAgentTranscript(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestNoSecretFirstAgentDebuggingSmoke(t *testing.T) {
|
||||
root := filepath.Clean(filepath.Join("..", "..", ".."))
|
||||
home := t.TempDir()
|
||||
storeDir := filepath.Join(home, "micro", "store")
|
||||
st := store.NewFileStore(store.DirOption(storeDir))
|
||||
|
||||
seedNoSecretAgentDebuggingState(t, st)
|
||||
if err := st.Close(); err != nil {
|
||||
t.Fatalf("close seeded store: %v", err)
|
||||
}
|
||||
|
||||
micro := buildMicroBinary(t, root)
|
||||
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
args []string
|
||||
want []string
|
||||
}{
|
||||
{
|
||||
name: "demo advertises provider-free debug path",
|
||||
args: []string{"agent", "demo"},
|
||||
want: []string{"No-secret first-agent demo", "provider-free", "run history", "micro inspect agent <name>"},
|
||||
},
|
||||
{
|
||||
name: "inspect shows seeded run history",
|
||||
args: []string{"inspect", "agent", "assistant", "--limit", "1"},
|
||||
want: []string{`Agent "assistant" runs`, "run-debug-smoke", "status=done", "events=3", "last=done", "trace=trace-debug-"},
|
||||
},
|
||||
{
|
||||
name: "inspect filters documented statuses",
|
||||
args: []string{"inspect", "agent", "--status", "done", "--json", "assistant"},
|
||||
want: []string{"run-debug-smoke", `"status": "done"`, `"trace_id": "trace-debug-smoke"`},
|
||||
},
|
||||
{
|
||||
name: "agent history shows memory and run index",
|
||||
args: []string{"agent", "history", "assistant"},
|
||||
want: []string{"user:", "Triage ticket-1", "assistant:", "ticket-1 is ready", "Runs:", "run-debug-smoke", "status=done"},
|
||||
},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
out := runMicroCLIWithHome(t, micro, home, tc.args...)
|
||||
for _, want := range tc.want {
|
||||
if !strings.Contains(out, want) {
|
||||
t.Fatalf("micro %s output missing %q:\n%s", strings.Join(tc.args, " "), want, out)
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func seedNoSecretAgentDebuggingState(t *testing.T, st store.Store) {
|
||||
t.Helper()
|
||||
scoped := store.Scope(st, "agent", "assistant")
|
||||
runID := "run-debug-smoke"
|
||||
events := []goagent.RunEvent{
|
||||
{Time: time.Unix(1700000000, 0), RunID: runID, Agent: "assistant", TraceID: "trace-debug-smoke", Kind: "run", Name: "ask"},
|
||||
{Time: time.Unix(1700000001, 0), RunID: runID, Agent: "assistant", TraceID: "trace-debug-smoke", Kind: "model", Provider: "mock", Model: "first-agent-mock"},
|
||||
{Time: time.Unix(1700000002, 0), RunID: runID, Agent: "assistant", TraceID: "trace-debug-smoke", Kind: "done", Name: "answer"},
|
||||
}
|
||||
for _, event := range events {
|
||||
b, err := json.Marshal(event)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
key := "runs/" + event.RunID + "/" + event.Time.Format("20060102150405.000000000") + "-" + event.Kind
|
||||
if err := scoped.Write(&store.Record{Key: key, Value: b}); err != nil {
|
||||
t.Fatalf("seed run event: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
mem := goagent.NewMemory(scoped, "history", 10)
|
||||
mem.Add("user", "Triage ticket-1 for Alice")
|
||||
mem.Add("assistant", "ticket-1 is ready for Alice without provider secrets")
|
||||
}
|
||||
|
||||
func buildMicroBinary(t *testing.T, root string) string {
|
||||
t.Helper()
|
||||
bin := filepath.Join(t.TempDir(), "micro")
|
||||
cmd := exec.Command("go", "build", "-o", bin, "./cmd/micro")
|
||||
cmd.Dir = root
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("build micro CLI failed: %v\n%s", err, out)
|
||||
}
|
||||
return bin
|
||||
}
|
||||
|
||||
func runMicroCLIWithHome(t *testing.T, micro, home string, args ...string) string {
|
||||
t.Helper()
|
||||
cmd := exec.Command(micro, args...)
|
||||
cmd.Env = append(os.Environ(),
|
||||
"HOME="+home,
|
||||
"MICRO_AI_API_KEY=",
|
||||
"OPENAI_API_KEY=",
|
||||
"ANTHROPIC_API_KEY=",
|
||||
"GEMINI_API_KEY=",
|
||||
)
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("micro %s failed: %v\n%s", strings.Join(args, " "), err, out)
|
||||
}
|
||||
return string(out)
|
||||
}
|
||||
|
||||
func TestFirstAgentWayfindingTargetsExist(t *testing.T) {
|
||||
root := filepath.Clean(filepath.Join("..", "..", ".."))
|
||||
for _, target := range []string{
|
||||
|
||||
@@ -8,7 +8,7 @@ cd "$ROOT"
|
||||
# without secrets or long-running daemons.
|
||||
go test ./cmd/micro -run 'TestFirstAgentWalkthroughCLIBoundaries|TestZeroToHeroCLIBoundaries' -count=1
|
||||
go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1
|
||||
go test ./internal/harness/zero-to-hero-ci -run 'TestNoSecretFirstAgentTranscript|TestZeroToHeroReferenceDocs|TestYourFirstAgentTutorialSmoke' -count=1
|
||||
go test ./internal/harness/zero-to-hero-ci -run 'TestNoSecretFirstAgentTranscript|TestNoSecretFirstAgentDebuggingSmoke|TestZeroToHeroReferenceDocs|TestYourFirstAgentTutorialSmoke' -count=1
|
||||
|
||||
# Deterministic no-secret reference scenarios. These use the real Go Micro
|
||||
# runtime and mock only the LLM provider. The support example is the maintained
|
||||
|
||||
@@ -42,6 +42,8 @@ examples:
|
||||
guides:
|
||||
- title: Debugging your agent
|
||||
url: /docs/guides/debugging-agents.html
|
||||
- title: micro loop quickstart
|
||||
url: /docs/guides/micro-loop.html
|
||||
- title: Plan & Delegate
|
||||
url: /docs/guides/plan-delegate.html
|
||||
- title: Agent Guardrails
|
||||
@@ -86,6 +88,7 @@ search_order:
|
||||
- /docs/guides/your-first-agent.html
|
||||
- /docs/guides/zero-to-hero.html
|
||||
- /docs/guides/debugging-agents.html
|
||||
- /docs/guides/micro-loop.html
|
||||
- /docs/getting-started.html
|
||||
- /docs/mcp.html
|
||||
- /docs/architecture.html
|
||||
|
||||
@@ -259,4 +259,5 @@ The flow discovers all services as tools and lets the LLM decide which RPCs to c
|
||||
- [Agent Design](https://github.com/micro/go-micro/blob/master/internal/docs/AGENT_DESIGN.md) — the full agent interface specification
|
||||
- [MCP & AI Agents](mcp.html) — MCP gateway, tool discovery, and auth
|
||||
- [Data Model](model.html) — typed persistence with CRUD and queries
|
||||
- [`micro loop` quickstart](guides/micro-loop.html) — scaffold a CI-gated autonomous improvement loop for a repository
|
||||
- [Deployment](deployment.html) — deploy via SSH + systemd
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
---
|
||||
layout: default
|
||||
---
|
||||
|
||||
# `micro loop` quickstart
|
||||
|
||||
`micro loop` scaffolds the autonomous improvement loop that Go Micro uses on
|
||||
this repository: GitHub Actions workflows for planning, building, evaluation
|
||||
feedback, coherence, security, and release. Use it when you want a repository to
|
||||
continuously turn a ranked queue into small PRs while CI remains the merge gate.
|
||||
|
||||
## 1. Initialize the loop
|
||||
|
||||
Run the default loop from the repository root:
|
||||
|
||||
```bash
|
||||
micro loop init
|
||||
```
|
||||
|
||||
For every role used by Go Micro itself, scaffold all workflows:
|
||||
|
||||
```bash
|
||||
micro loop init --roles all
|
||||
```
|
||||
|
||||
The command writes:
|
||||
|
||||
- `.github/loop/NORTH_STAR.md` — the direction every increment should optimize.
|
||||
- `.github/loop/PRIORITIES.md` — the ranked queue; the builder takes the top open issue.
|
||||
- `.github/loop/prompts/*.md` — editable policy for planner, builder, triage, coherence, and security roles.
|
||||
- `.github/workflows/loop-*.yml` — generated GitHub Actions mechanics.
|
||||
|
||||
Edit the files under `.github/loop/` to steer the loop. Re-run
|
||||
`micro loop init --roles all --force` only when you want to regenerate workflow
|
||||
mechanics from the installed CLI.
|
||||
|
||||
## 2. Configure the dispatch token
|
||||
|
||||
The scheduled builder needs a repository secret containing a token from a user
|
||||
account that the coding agent will answer. Go Micro names that secret
|
||||
`CODEX_TRIGGER_TOKEN` by default. If you use another secret name, pass it when
|
||||
you initialize the loop:
|
||||
|
||||
```bash
|
||||
micro loop init --agent @codex --token-secret LOOP_TOKEN --roles all
|
||||
```
|
||||
|
||||
The token needs enough repository permission to open issues, comment, push
|
||||
branches, create pull requests, and enable auto-merge. Run `gh auth setup-git` in
|
||||
the environment that will push branches so `git push` uses the same credentials
|
||||
as `gh`.
|
||||
|
||||
## 3. Make CI the gate
|
||||
|
||||
The loop should not be its own reviewer. Protect the default branch so PRs merge
|
||||
only after the required checks pass. At minimum, require the same commands the
|
||||
Go Micro loop verifies locally and in CI:
|
||||
|
||||
```bash
|
||||
go build ./...
|
||||
go test ./...
|
||||
golangci-lint run ./...
|
||||
```
|
||||
|
||||
If your repository has a harness or end-to-end grader, make that required too.
|
||||
Keep human approval requirements out of the autonomous path unless you intend the
|
||||
loop to pause for review.
|
||||
|
||||
## 4. Verify the wiring
|
||||
|
||||
After editing the North Star, queue, prompts, token secret, and branch
|
||||
protection, run:
|
||||
|
||||
```bash
|
||||
micro loop verify
|
||||
```
|
||||
|
||||
`micro loop verify` checks that the loop direction, queue, prompts, role
|
||||
workflows, and non-loop CI gate are present. Fix any reported missing items
|
||||
before relying on scheduled increments.
|
||||
|
||||
## 5. Operate the queue
|
||||
|
||||
Keep one ranked list in `.github/loop/PRIORITIES.md`. Each item should link a
|
||||
scoped issue and be small enough for one PR. The builder closes both the priority
|
||||
issue and the per-run tracker issue in the PR body, for example:
|
||||
|
||||
```text
|
||||
Closes #1234
|
||||
Closes #5678
|
||||
```
|
||||
|
||||
Use the North Star to keep the queue honest: favor small improvements that move
|
||||
developers through the services → agents → workflows lifecycle, and surface
|
||||
breaking API or brand/positioning decisions for humans instead of auto-merging
|
||||
them.
|
||||
@@ -100,6 +100,7 @@ CI keeps those CLI boundaries present with:
|
||||
|
||||
```sh
|
||||
go test ./cmd/micro -run TestFirstAgentWalkthroughCLIBoundaries -count=1
|
||||
go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentDebuggingSmoke -count=1
|
||||
```
|
||||
|
||||
## Debug transcript checkpoint
|
||||
|
||||
@@ -22,7 +22,7 @@ cloud credentials?"
|
||||
| First agent | `micro new`, `micro agent preflight`, `micro run`, `micro chat`, and `micro inspect agent <name>` stay available for the documented first-agent walkthrough. | `go test ./cmd/micro -run TestFirstAgentWalkthroughCLIBoundaries -count=1` |
|
||||
| Run | `micro run` remains the local development entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
|
||||
| Chat | `micro chat` remains the interactive agent entry point. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
|
||||
| Inspect | `micro inspect agent <name>`, `micro inspect flow <flow>`, and `micro flow runs <flow>` remain discoverable for run history. | `go test ./cmd/micro -run TestZeroToHeroCLIBoundaries -count=1` |
|
||||
| Inspect | `micro inspect agent <name>`, `micro agent history <name>`, `micro inspect flow <flow>`, and `micro flow runs <flow>` remain discoverable for run history; the no-secret debugging smoke seeds durable agent history and runs the documented inspect/history commands without provider keys. | `go test ./internal/harness/zero-to-hero-ci -run TestNoSecretFirstAgentDebuggingSmoke -count=1` |
|
||||
| Deploy | `micro deploy --dry-run` resolves deploy targets without touching remote infrastructure. | `go test ./cmd/micro/cli/deploy -run TestDeployDryRun -count=1` |
|
||||
| Smallest first agent | `examples/first-agent` runs one service-backed agent with a deterministic mock model and no provider key. | `go test ./examples/first-agent -run TestRunFirstAgent -count=1` |
|
||||
| Runtime reference app | `examples/support` runs typed services, an agent using those services as tools, an event-driven flow handoff, and an approval gate with only the model mocked. | `go test ./examples/support -run 'TestRunSupportMockSmoke|TestZeroToHeroReadmeDocumentsLifecycle' -count=1` |
|
||||
|
||||
@@ -30,6 +30,7 @@ Otherwise continue to read the docs for more information about the framework.
|
||||
- [Your First Agent](guides/your-first-agent.html) - Build a service-backed agent end to end
|
||||
- [MCP & AI Agents](mcp.html) - Turn services into AI-callable tools with the Model Context Protocol
|
||||
- [CLI & Gateway Guide](guides/cli-gateway.html) - Development vs Production modes
|
||||
- [`micro loop` quickstart](guides/micro-loop.html) - Scaffold an autonomous CI-gated improvement loop
|
||||
- [Quick Start](quickstart.html)
|
||||
- [Architecture](architecture.html)
|
||||
- [Configuration](config.html)
|
||||
@@ -64,5 +65,6 @@ Otherwise continue to read the docs for more information about the framework.
|
||||
- [Real-World Examples](examples/realworld/)
|
||||
- [Migration Guides](guides/migration/)
|
||||
- [Observability](observability.html)
|
||||
- [`micro loop` quickstart](guides/micro-loop.html)
|
||||
- [Contributing](contributing.html)
|
||||
- [Roadmap](roadmap.html)
|
||||
|
||||
Reference in New Issue
Block a user