[Agent Architecture] AI 主流的八种Agent架构的理解和简单工程实践

[Agent Architecture] AI 主流的八种Agent架构的理解和简单工程实践

LLM本身只是推理引擎。真正让系统具备“感知—决策—行动—反馈”能力的,是围绕模型构建的 Agent 架构。
趁着十一假期梳理了一下业界最常被验证的八种 Agent 架构模式和实现思路。它们不是互斥选项,而是可以像积木一样组合的工程机制,没有绝对完美的架构。

以下结合业务场景对常见几种 Agent 架构进行简单拆解。个人习惯使用Golang实现Agent。Python暂不做实现。

一、先统一抽象:Go 里的 Agent 骨架

在 Go 中,Agent 系统适合用“小接口 + 组合”表达。下面这组接口贯穿全文:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
type Message struct {
Role string // system / user / assistant
Content string
}

type Model interface {
Generate(ctx context.Context, msgs []Message) (string, error)
}

type Tool interface {
Name() string
Run(ctx context.Context, arg string) (string, error)
}

type Agent interface {
Run(ctx context.Context, input string) (string, error)
}

context.Context 负责超时、取消和链路追踪;Tool 统一工具注册、权限校验和 Mock 测试;Agent 让不同架构可以互相嵌套。以下八种架构都围绕这套骨架展开。


1. ReAct:推理与行动的最小闭环

核心思想
ReAct = Reasoning + Acting。模型先根据历史推理下一步,再选择工具并生成参数,工具返回真实观察结果后写回上下文,模型继续推理,直到输出最终答案。

流程结构图

graph TD
    A[用户输入] --> B[构建上下文]
    B --> C{模型推理}
    C -->|输出 FINAL| D[返回最终答案]
    C -->|输出 TOOL| E[解析工具名与参数]
    E --> F{工具是否注册}
    F -->|否| G[写入错误观察]
    F -->|是| H[执行工具]
    H --> I[写入观察结果]
    G --> B
    I --> B
    C -->|其他输出| D
    C -->|超过最大轮次| J[强制终止]

适用场景
路径不确定、需要边执行边判断的任务,例如查资料、调 API、多步信息收集。

工程风险
循环成本、步数失控、工具误调用。必须设置 maxTurns、超时、工具白名单和终止条件。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
type ReActAgent struct {
Model Model
Tools map[string]Tool
MaxTurns int
}

func (a *ReActAgent) Run(ctx context.Context, input string) (string, error) {
max := a.MaxTurns
if max <= 0 {
max = 6
}

msgs := []Message{{Role: "user", Content: input}}

for turn := 0; turn < max; turn++ {
out, err := a.Model.Generate(ctx, msgs)
if err != nil {
return "", err
}
msgs = append(msgs, Message{Role: "assistant", Content: out})

switch {
case strings.HasPrefix(out, "FINAL:"):
return strings.TrimSpace(strings.TrimPrefix(out, "FINAL:")), nil

case strings.HasPrefix(out, "TOOL:"):
body := strings.TrimSpace(strings.TrimPrefix(out, "TOOL:"))
name, arg, ok := strings.Cut(body, "|")
if !ok {
msgs = append(msgs, Message{
Role: "user",
Content: "OBSERVE: 工具调用格式应为 TOOL:name|arg",
})
continue
}

tool, ok := a.Tools[strings.TrimSpace(name)]
if !ok {
msgs = append(msgs, Message{
Role: "user",
Content: "OBSERVE: 未注册工具 " + name,
})
continue
}

res, err := tool.Run(ctx, strings.TrimSpace(arg))
if err != nil {
res = "ERROR: " + err.Error()
}
msgs = append(msgs, Message{Role: "user", Content: "OBSERVE: " + res})

default:
return out, nil
}
}

return "", fmt.Errorf("ReAct 超过最大轮次 %d", max)
}

2. Plan-and-Execute:先规划,后执行

核心思想
把“做什么”和“怎么做”拆开。Planner 先把目标拆成有序子任务,Executor 再逐项执行。执行失败时可以触发重规划,而不是推翻全部结果。

流程结构图

graph TD
    A[用户目标] --> B[Planner 拆解目标]
    B --> C[生成子任务列表]
    C --> D{还有任务未执行?}
    D -->|是| E[取出下一个任务]
    E --> F[Executor 执行]
    F --> G{执行成功?}
    G -->|是| H[记录结果]
    G -->|否| I[触发重规划]
    I --> B
    H --> D
    D -->|否| J[汇总所有结果]
    J --> K[返回最终输出]

适用场景
步骤多、依赖明确、需要审计执行过程的复杂任务,例如季度分析、报告生成、多阶段数据处理。

工程风险
规划质量是瓶颈;对简单任务,规划开销是净浪费。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
type Planner interface {
Plan(ctx context.Context, goal string) ([]string, error)
}

type Executor interface {
Do(ctx context.Context, step string) (string, error)
}

type PlanExecuteAgent struct {
Planner Planner
Executor Executor
}

func (a *PlanExecuteAgent) Run(ctx context.Context, goal string) (string, error) {
steps, err := a.Planner.Plan(ctx, goal)
if err != nil {
return "", err
}

var outputs []string
for i, step := range steps {
res, err := a.Executor.Do(ctx, step)
if err != nil {
newGoal := fmt.Sprintf("原目标:%s;第 %d 步失败:%s;请重新规划", goal, i+1, step)
steps, err = a.Planner.Plan(ctx, newGoal)
if err != nil {
return "", err
}
continue
}
outputs = append(outputs, res)
}

return strings.Join(outputs, "\n"), nil
}

3. Multi-Agent:角色分工与协作拓扑

核心思想
Multi-Agent 不是“多开几个模型”,而是定义角色、消息传递和协作拓扑。常见拓扑有三类:层级式(Supervisor 拆任务)、并行式(独立子任务同时执行)、流水线式(A 的输出交给 B)。

流程结构图

graph TD
    A[用户输入] --> B[Supervisor 路由]
    B --> C{选择 Worker}
    C -->|搜索类| D[搜索 Agent]
    C -->|分析类| E[分析 Agent]
    C -->|写作类| F[写作 Agent]
    D --> G[并行/串行收集结果]
    E --> G
    F --> G
    G --> H[合并输出]
    H --> I[返回最终结果]

适用场景
任务超出单 Agent 上下文窗口,或需要搜索、分析、写作、质检等专精角色并行。

工程风险
协调开销大、调试困难。单 Agent 未验证可靠前就上 Multi-Agent,是昂贵且常见的错误。

Go 实现:Supervisor 路由

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
type Supervisor struct {
Workers map[string]Agent
Router func(input string) string
}

func (s *Supervisor) Run(ctx context.Context, input string) (string, error) {
name := "writer"
if s.Router != nil {
name = s.Router(input)
}

worker, ok := s.Workers[name]
if !ok {
return "", fmt.Errorf("Supervisor: 没有可用 Worker %s", name)
}

return worker.Run(ctx, input)
}

Go 实现:并行 Worker

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
type ParallelAgent struct {
Agents []Agent
}

func (p *ParallelAgent) Run(ctx context.Context, input string) (string, error) {
type result struct {
out string
err error
}

ch := make(chan result, len(p.Agents))
for _, ag := range p.Agents {
go func(a Agent) {
out, err := a.Run(ctx, input)
ch <- result{out: out, err: err}
}(ag)
}

var parts []string
for range p.Agents {
r := <-ch
if r.err != nil {
return "", r.err
}
parts = append(parts, r.out)
}

return strings.Join(parts, "\n"), nil
}

4. Reflective Agent:生成后自我审查

核心思想
Agent 生成候选答案后,不直接返回,而是由 Critic 检查质量。若不合格,把反馈交回 Agent 重新生成。类似人类“写完—审查—修订”的过程。

流程结构图

graph TD
    A[用户输入] --> B[Agent 生成初稿]
    B --> C[Reviewer 审查]
    C --> D{是否合格?}
    D -->|是| E[返回最终结果]
    D -->|否| F{是否达到最大轮次?}
    F -->|是| E
    F -->|否| G[附反馈重新生成]
    G --> C

适用场景
代码生成、数学推理、复杂决策验证等精度优先于延迟的场景。

工程风险
调用次数和延迟显著增加,需要限制审查轮数。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
type Reviewer interface {
Check(ctx context.Context, draft string) (ok bool, feedback string)
}

type ReflectiveAgent struct {
Inner Agent
Reviewer Reviewer
Rounds int
}

func (a *ReflectiveAgent) Run(ctx context.Context, input string) (string, error) {
draft, err := a.Inner.Run(ctx, input)
if err != nil {
return "", err
}

rounds := a.Rounds
if rounds <= 0 {
rounds = 1
}

for i := 0; i < rounds; i++ {
ok, feedback := a.Reviewer.Check(ctx, draft)
if ok {
return draft, nil
}

draft, err = a.Inner.Run(ctx, input+"\n审查反馈:"+feedback)
if err != nil {
return "", err
}
}

return draft, nil
}

5. Tool-Augmented Agent:工具是 Agent 的手脚

核心思想
Agent 的价值不只在于“知道什么”,更在于“能做什么”。工具增强要解决:工具描述、参数校验、结果压缩、权限控制、审批流程。

流程结构图

graph TD
    A[Agent 决策调用工具] --> B[ToolBox 查找工具]
    B --> C{工具是否注册?}
    C -->|否| D[返回未知工具错误]
    C -->|是| E[Policy 权限校验]
    E --> F{是否通过?}
    F -->|否| G[拒绝调用并返回]
    F -->|是| H[执行工具]
    H --> I[结果压缩]
    I --> J[写回上下文]

适用场景
所有需要外部操作的任务:工单系统、搜索、数据库、Email、代码执行、支付等等。

工程风险
权限失控。生产系统通常把工具分为“可直接调用”和“需要审批”两类常用场景。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
type ToolBox struct {
Tools map[string]Tool
Policy func(toolName, arg string) error
}

func (b *ToolBox) Call(ctx context.Context, name, arg string) (string, error) {
tool, ok := b.Tools[name]
if !ok {
return "", fmt.Errorf("ToolBox: 未注册工具 %s", name)
}

if b.Policy != nil {
if err := b.Policy(name, arg); err != nil {
return "", err
}
}

return tool.Run(ctx, arg)
}

用函数适配器快速注册工具:

1
2
3
4
5
6
7
8
9
10
type ToolFunc struct {
NameValue string
Fn func(context.Context, string) (string, error)
}

func (t ToolFunc) Name() string { return t.NameValue }

func (t ToolFunc) Run(ctx context.Context, arg string) (string, error) {
return t.Fn(ctx, arg)
}

6. Memory-Augmented Agent:跨会话上下文管理

核心思想
上下文窗口是稀缺资源。Memory-Augmented 把记忆分层:工作记忆(当前会话短期上下文)、会话记忆(本次对话持久化记录)、长期记忆(跨会话知识、用户偏好、历史模式)。Agent 需要决定何时写入、何时检索、何时压缩或丢弃。

流程结构图

graph TD
    A[用户输入] --> B[检索长期记忆]
    B --> C{是否有相关记忆?}
    C -->|否| D[直接使用原始输入]
    C -->|是| E[注入记忆到上下文]
    D --> F[Agent 生成回复]
    E --> F
    F --> G[写入会话/长期记忆]
    G --> H[返回结果]

适用场景
长对话、跨会话助手、个性化服务、需要记住历史决策的系统。

工程风险
记忆污染、检索偏差、隐私合规。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
type Memory interface {
Put(ctx context.Context, key, value string) error
Recall(ctx context.Context, query string, limit int) ([]string, error)
}

type MemoryAgent struct {
Inner Agent
Mem Memory
}

func (a *MemoryAgent) Run(ctx context.Context, input string) (string, error) {
recalls, _ := a.Mem.Recall(ctx, input, 3)

prompt := input
if len(recalls) > 0 {
prompt = "相关记忆:\n" + strings.Join(recalls, "\n") + "\n\n" + input
}

out, err := a.Inner.Run(ctx, prompt)
if err != nil {
return "", err
}

_ = a.Mem.Put(ctx, input, out)
return out, nil
}

7. RAG Agent:检索优先的知识工作流

核心思想
RAG Agent 是 Memory-Augmented 的知识特化版。链路通常是:文档解析 → 分块 → 嵌入 → 向量存储 → 检索 → 注入上下文 → 生成回答。

流程结构图

graph TD
    A[知识文档] --> B[文档解析]
    B --> C[分块]
    C --> D[向量嵌入]
    D --> E[写入向量库]
    F[用户问题] --> G[生成查询向量]
    E --> H[相似度检索 TopK]
    G --> H
    H --> I{是否命中?}
    I -->|否| J[直接回答或拒答]
    I -->|是| K[注入上下文]
    K --> L[LLM 生成回答]
    L --> M[返回结果]

适用场景
知识库问答、文档分析、企业检索助手、客服知识库。

工程风险
质量瓶颈往往不在检索,而在文档解析和分块策略。解析错了,再好的向量检索也救不回来。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
type Retriever interface {
TopK(ctx context.Context, query string, k int) ([]string, error)
}

type RAGAgent struct {
Inner Agent
Retriever Retriever
TopK int
}

func (a *RAGAgent) Run(ctx context.Context, input string) (string, error) {
k := a.TopK
if k <= 0 {
k = 4
}

docs, err := a.Retriever.TopK(ctx, input, k)
if err != nil {
return "", err
}

prompt := "请优先根据以下资料回答;若资料不足,请明确说明。\n\n资料:\n" +
strings.Join(docs, "\n---\n") +
"\n\n问题:" + input

return a.Inner.Run(ctx, prompt)
}

8. Autonomous Loop:围绕长期目标持续运行

核心思想
Agent 不被单次用户输入驱动,而是围绕一个长期目标持续执行:感知 → 规划 → 执行 → 验证 → 继续,直到目标达成或人工终止。

流程结构图

graph TD
    A[设定长期目标] --> B{预算与循环检查}
    B -->|超预算/超轮次| C[终止并报告]
    B -->|正常| D[执行 Agent]
    D --> E{目标是否达成?}
    E -->|是| F[返回结果并结束]
    E -->|否| G[更新目标状态与上下文]
    G --> H{是否通过人工检查点?}
    H -->|否| B
    H -->|是| I[等待人工干预]
    I --> B

适用场景
研究性任务、监控类工作流、需要数小时甚至数天持续执行的任务。

工程风险
资源失控、错误循环、无终止。必须配备持久化状态、成本熔断、人工检查点和最大循环次数。

Go 实现

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
type LoopAgent struct {
Inner Agent
Goal string
MaxCycles int
Budget int
}

func (a *LoopAgent) Run(ctx context.Context) (string, error) {
goal := a.Goal

for i := 0; i < a.MaxCycles; i++ {
select {
case <-ctx.Done():
return "", ctx.Err()
default:
}

out, err := a.Inner.Run(ctx, goal)
if err != nil {
return "", err
}

a.Budget -= len(out)
if a.Budget <= 0 {
return "", fmt.Errorf("Autonomous Loop: 预算耗尽")
}

if strings.Contains(out, "DONE") {
return out, nil
}

goal = "继续推进目标:" + a.Goal + "\n上一轮结果:" + out
}

return "", fmt.Errorf("Autonomous Loop: 达到最大循环次数")
}

三、架构选择速查表

架构解决的问题典型场景主要代价
ReAct路径未知,边走边看工具调用、多步信息收集循环成本、步数失控
Plan-and-Execute步骤多,依赖明确复杂多步任务、可审计流程规划质量瓶颈
Multi-Agent单 Agent 能力或上下文不足并行专精、角色分工协调开销、调试困难
Reflective输出精度优先代码生成、数学推理延迟和成本翻倍
Tool-AugmentedAgent 能做什么所有外部操作类任务权限控制、安全审批
Memory-Augmented上下文窗口不够用长对话、跨会话助手记忆污染、检索偏差
RAG Agent需要外部知识知识库问答、文档分析解析和分块质量
Autonomous Loop长期目标持续执行研究、监控、长周期任务资源失控、无终止

四、组合视角:真实 Agent 是多种架构的叠加

八大架构并不是“八选一”。一个生产级 Agent 常见组合方式如下:

graph TD
    A[用户请求] --> B[Memory-Agent 检索历史]
    B --> C[Plan-Agent 规划步骤]
    C --> D[ReAct-Agent 执行每步]
    D --> E[ToolBox 调用工具]
    E --> F[Reflective-Agent 审查结果]
    F --> G{是否合格?}
    G -->|否| D
    G -->|是| H[RAG-Agent 补充知识]
    H --> I[Multi-Agent 汇总输出]
    I --> J[Autonomous Loop 持续跟进]

建议的落地顺序

  1. 先用 ReAct 跑通端到端;
  2. 加入 ToolBox 管理工具权限;
  3. 遇到上下文瓶颈时加 Memory 或 RAG;
  4. 精度不够时加 Reflective;
  5. 任务并行度不够时加 Multi-Agent;
  6. 需要长周期运行时再加 Autonomous Loop。

不要一开始就全都上。架构越多,循环路径越复杂,调试成本呈指数上升。


五、完整可运行示例:Go 实现 ReAct Agent

保存为 main.go,执行 go run main.go。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
package main

import (
"context"
"fmt"
"log"
"strings"
)

type Message struct {
Role string
Content string
}

type Model interface {
Generate(ctx context.Context, msgs []Message) (string, error)
}

type Tool interface {
Name() string
Run(ctx context.Context, arg string) (string, error)
}

type ToolFunc struct {
name string
fn func(context.Context, string) (string, error)
}

func (t ToolFunc) Name() string { return t.name }

func (t ToolFunc) Run(ctx context.Context, arg string) (string, error) {
return t.fn(ctx, arg)
}

type Agent interface {
Run(ctx context.Context, input string) (string, error)
}

type ReActAgent struct {
Model Model
Tools map[string]Tool
MaxTurns int
}

func (a *ReActAgent) Run(ctx context.Context, input string) (string, error) {
max := a.MaxTurns
if max <= 0 {
max = 6
}

msgs := []Message{{Role: "user", Content: input}}

for turn := 0; turn < max; turn++ {
out, err := a.Model.Generate(ctx, msgs)
if err != nil {
return "", err
}
msgs = append(msgs, Message{Role: "assistant", Content: out})

switch {
case strings.HasPrefix(out, "FINAL:"):
return strings.TrimSpace(strings.TrimPrefix(out, "FINAL:")), nil

case strings.HasPrefix(out, "TOOL:"):
body := strings.TrimSpace(strings.TrimPrefix(out, "TOOL:"))
name, arg, ok := strings.Cut(body, "|")
if !ok {
msgs = append(msgs, Message{
Role: "user",
Content: "OBSERVE: 工具调用格式应为 TOOL:name|arg",
})
continue
}

tool, ok := a.Tools[strings.TrimSpace(name)]
if !ok {
msgs = append(msgs, Message{
Role: "user",
Content: "OBSERVE: 未注册工具 " + name,
})
continue
}

res, err := tool.Run(ctx, strings.TrimSpace(arg))
if err != nil {
res = "ERROR: " + err.Error()
}
msgs = append(msgs, Message{Role: "user", Content: "OBSERVE: " + res})

default:
return out, nil
}
}

return "", fmt.Errorf("ReAct 超过最大轮次 %d", max)
}

type ScriptModel struct{}

func (ScriptModel) Generate(ctx context.Context, msgs []Message) (string, error) {
toolCalls := 0
for _, m := range msgs {
if m.Role == "assistant" && strings.HasPrefix(m.Content, "TOOL:") {
toolCalls++
}
}

switch toolCalls {
case 0:
return "TOOL:天气|成都", nil
case 1:
return "TOOL:百科|Go语言", nil
default:
return "FINAL:成都未来三天多云转晴,18-26℃;Go 语言由 Google 设计,原生支持 goroutine 与 channel,适合构建高并发 Agent 服务。", nil
}
}

func main() {
tools := map[string]Tool{
"天气": ToolFunc{
name: "天气",
fn: func(ctx context.Context, arg string) (string, error) {
return arg + ":未来三天多云转晴,18-26℃", nil
},
},
"百科": ToolFunc{
name: "百科",
fn: func(ctx context.Context, arg string) (string, error) {
return arg + ":静态强类型编译语言,原生支持 goroutine 与 channel。", nil
},
},
}

agent := &ReActAgent{
Model: ScriptModel{},
Tools: tools,
MaxTurns: 5,
}

out, err := agent.Run(context.Background(), "查成都天气并介绍 Go 语言")
if err != nil {
log.Fatal(err)
}

fmt.Println(out)
}

运行结果类似:

1
成都未来三天多云转晴,18-26℃;Go 语言由 Google 设计,原生支持 goroutine 与 channel,适合构建高并发 Agent 服务。

六、结语

主流的这八类 Agent 架构不是八个互斥答案,而是八种可组合的工程机制。但如果是非常简单的场景,切记不要生搬硬套框架,尤其只需要一个接口就能解决的事情,不要套一整个Agent逻辑!

一个生产级 Agent 可能同时具备 ReAct 的循环、Plan-and-Execute 的规划层、Reflective 的审查层、Memory-Augmented 的跨会话记忆,以及 Tool-Augmented 的权限控制,根据实际情况选型即可,不要过度设计⚠️

Go 的接口组合、context、goroutine 和 channel,让这些机制可以像积木一样拼装。
选择架构的第一原则不是“哪个最强”,而是“哪个最小够用”。从 ReAct 开始,遇到上下文、精度、并行度或长期运行瓶颈时,再逐步叠加其他模式。这样比一开始堆满架构、然后花几倍时间调试失控循环,要划算得多。

[Agent Architecture] AI 主流的八种Agent架构的理解和简单工程实践

https://www.wdft.com/c568e978.html

Author

Jaco Liu

Posted on

2026-10-05

Updated on

2026-10-05

Licensed under