Compare commits
20 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| fad4722324 | |||
| e2335877a8 | |||
| 0b62ef35b4 | |||
| a27732dca8 | |||
| 3f1e24f1d6 | |||
| 6026e7a9c6 | |||
| b4161badc6 | |||
| a6b3571077 | |||
| abb9eb28a1 | |||
| 28a4a44420 | |||
| d307ee9a8d | |||
| a890058eb8 | |||
| 2a437a2cd9 | |||
| 14c15d96d0 | |||
| caa7ec8e16 | |||
| d82387244a | |||
| 1dc96af1c8 | |||
| 8a1c0ce89c | |||
| 2b9e0ba673 | |||
| b0dfe3b224 |
@@ -1,176 +1,176 @@
|
||||
<p align="center">
|
||||
<img src="assets/logo.jpg" width="200"/>
|
||||
</p>
|
||||
|
||||
English | [中文](README_zh.md) | [한국어](README_ko.md) | [日本語](README_ja.md)
|
||||
|
||||
[](https://github.com/mannaandpoem/OpenManus/stargazers)
|
||||
 
|
||||
[](https://opensource.org/licenses/MIT)  
|
||||
[](https://discord.gg/DYn29wFk9z)
|
||||
|
||||
# 👋 OpenManus
|
||||
|
||||
Manus is incredible, but OpenManus can achieve any idea without an *Invite Code* 🛫!
|
||||
|
||||
Our team members [@Xinbin Liang](https://github.com/mannaandpoem) and [@Jinyu Xiang](https://github.com/XiangJinyu) (core authors), along with [@Zhaoyang Yu](https://github.com/MoshiQAQ), [@Jiayi Zhang](https://github.com/didiforgithub), and [@Sirui Hong](https://github.com/stellaHSR), we are from [@MetaGPT](https://github.com/geekan/MetaGPT). The prototype is launched within 3 hours and we are keeping building!
|
||||
|
||||
It's a simple implementation, so we welcome any suggestions, contributions, and feedback!
|
||||
|
||||
Enjoy your own agent with OpenManus!
|
||||
|
||||
We're also excited to introduce [OpenManus-RL](https://github.com/OpenManus/OpenManus-RL), an open-source project dedicated to reinforcement learning (RL)- based (such as GRPO) tuning methods for LLM agents, developed collaboratively by researchers from UIUC and OpenManus.
|
||||
|
||||
## Project Demo
|
||||
|
||||
<video src="https://private-user-images.githubusercontent.com/61239030/420168772-6dcfd0d2-9142-45d9-b74e-d10aa75073c6.mp4?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NDEzMTgwNTksIm5iZiI6MTc0MTMxNzc1OSwicGF0aCI6Ii82MTIzOTAzMC80MjAxNjg3NzItNmRjZmQwZDItOTE0Mi00NWQ5LWI3NGUtZDEwYWE3NTA3M2M2Lm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTAzMDclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwMzA3VDAzMjIzOVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTdiZjFkNjlmYWNjMmEzOTliM2Y3M2VlYjgyNDRlZDJmOWE3NWZhZjE1MzhiZWY4YmQ3NjdkNTYwYTU5ZDA2MzYmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.UuHQCgWYkh0OQq9qsUWqGsUbhG3i9jcZDAMeHjLt5T4" data-canonical-src="https://private-user-images.githubusercontent.com/61239030/420168772-6dcfd0d2-9142-45d9-b74e-d10aa75073c6.mp4?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NDEzMTgwNTksIm5iZiI6MTc0MTMxNzc1OSwicGF0aCI6Ii82MTIzOTAzMC80MjAxNjg3NzItNmRjZmQwZDItOTE0Mi00NWQ5LWI3NGUtZDEwYWE3NTA3M2M2Lm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTAzMDclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwMzA3VDAzMjIzOVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTdiZjFkNjlmYWNjMmEzOTliM2Y3M2VlYjgyNDRlZDJmOWE3NWZhZjE1MzhiZWY4YmQ3NjdkNTYwYTU5ZDA2MzYmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.UuHQCgWYkh0OQq9qsUWqGsUbhG3i9jcZDAMeHjLt5T4" controls="controls" muted="muted" class="d-block rounded-bottom-2 border-top width-fit" style="max-height:640px; min-height: 200px"></video>
|
||||
|
||||
## Installation
|
||||
|
||||
We provide two installation methods. Method 2 (using uv) is recommended for faster installation and better dependency management.
|
||||
|
||||
### Method 1: Using conda
|
||||
|
||||
1. Create a new conda environment:
|
||||
|
||||
```bash
|
||||
conda create -n open_manus python=3.12
|
||||
conda activate open_manus
|
||||
```
|
||||
|
||||
2. Clone the repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/mannaandpoem/OpenManus.git
|
||||
cd OpenManus
|
||||
```
|
||||
|
||||
3. Install dependencies:
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Method 2: Using uv (Recommended)
|
||||
|
||||
1. Install uv (A fast Python package installer and resolver):
|
||||
|
||||
```bash
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
```
|
||||
|
||||
2. Clone the repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/mannaandpoem/OpenManus.git
|
||||
cd OpenManus
|
||||
```
|
||||
|
||||
3. Create a new virtual environment and activate it:
|
||||
|
||||
```bash
|
||||
uv venv --python 3.12
|
||||
source .venv/bin/activate # On Unix/macOS
|
||||
# Or on Windows:
|
||||
# .venv\Scripts\activate
|
||||
```
|
||||
|
||||
4. Install dependencies:
|
||||
|
||||
```bash
|
||||
uv pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Browser Automation Tool (Optional)
|
||||
```bash
|
||||
playwright install
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
OpenManus requires configuration for the LLM APIs it uses. Follow these steps to set up your configuration:
|
||||
|
||||
1. Create a `config.toml` file in the `config` directory (you can copy from the example):
|
||||
|
||||
```bash
|
||||
cp config/config.example.toml config/config.toml
|
||||
```
|
||||
|
||||
2. Edit `config/config.toml` to add your API keys and customize settings:
|
||||
|
||||
```toml
|
||||
# Global LLM configuration
|
||||
[llm]
|
||||
model = "gpt-4o"
|
||||
base_url = "https://api.openai.com/v1"
|
||||
api_key = "sk-..." # Replace with your actual API key
|
||||
max_tokens = 4096
|
||||
temperature = 0.0
|
||||
|
||||
# Optional configuration for specific LLM models
|
||||
[llm.vision]
|
||||
model = "gpt-4o"
|
||||
base_url = "https://api.openai.com/v1"
|
||||
api_key = "sk-..." # Replace with your actual API key
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
One line for run OpenManus:
|
||||
|
||||
```bash
|
||||
python main.py
|
||||
```
|
||||
|
||||
Then input your idea via terminal!
|
||||
|
||||
For MCP tool version, you can run:
|
||||
```bash
|
||||
python run_mcp.py
|
||||
```
|
||||
|
||||
For unstable multi-agent version, you also can run:
|
||||
|
||||
```bash
|
||||
python run_flow.py
|
||||
```
|
||||
|
||||
## How to contribute
|
||||
|
||||
We welcome any friendly suggestions and helpful contributions! Just create issues or submit pull requests.
|
||||
|
||||
Or contact @mannaandpoem via 📧email: mannaandpoem@gmail.com
|
||||
|
||||
**Note**: Before submitting a pull request, please use the pre-commit tool to check your changes. Run `pre-commit run --all-files` to execute the checks.
|
||||
|
||||
## Community Group
|
||||
Join our networking group on Feishu and share your experience with other developers!
|
||||
|
||||
<div align="center" style="display: flex; gap: 20px;">
|
||||
<img src="assets/community_group.jpg" alt="OpenManus 交流群" width="300" />
|
||||
</div>
|
||||
|
||||
## Star History
|
||||
|
||||
[](https://star-history.com/#mannaandpoem/OpenManus&Date)
|
||||
|
||||
## Acknowledgement
|
||||
|
||||
Thanks to [anthropic-computer-use](https://github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo)
|
||||
and [browser-use](https://github.com/browser-use/browser-use) for providing basic support for this project!
|
||||
|
||||
Additionally, we are grateful to [AAAJ](https://github.com/metauto-ai/agent-as-a-judge), [MetaGPT](https://github.com/geekan/MetaGPT), [OpenHands](https://github.com/All-Hands-AI/OpenHands) and [SWE-agent](https://github.com/SWE-agent/SWE-agent).
|
||||
|
||||
OpenManus is built by contributors from MetaGPT. Huge thanks to this agent community!
|
||||
|
||||
## Cite
|
||||
```bibtex
|
||||
@misc{openmanus2025,
|
||||
author = {Xinbin Liang and Jinyu Xiang and Zhaoyang Yu and Jiayi Zhang and Sirui Hong},
|
||||
title = {OpenManus: An open-source framework for building general AI agents},
|
||||
year = {2025},
|
||||
publisher = {GitHub},
|
||||
journal = {GitHub repository},
|
||||
howpublished = {\url{https://github.com/mannaandpoem/OpenManus}},
|
||||
}
|
||||
```
|
||||
<p align="center">
|
||||
<img src="assets/logo.jpg" width="200"/>
|
||||
</p>
|
||||
|
||||
English | [中文](README_zh.md) | [한국어](README_ko.md) | [日本語](README_ja.md)
|
||||
|
||||
[](https://github.com/mannaandpoem/OpenManus/stargazers)
|
||||
 
|
||||
[](https://opensource.org/licenses/MIT)  
|
||||
[](https://discord.gg/DYn29wFk9z)
|
||||
|
||||
# 👋 OpenManus
|
||||
|
||||
Manus is incredible, but OpenManus can achieve any idea without an *Invite Code* 🛫!
|
||||
|
||||
Our team members [@Xinbin Liang](https://github.com/mannaandpoem) and [@Jinyu Xiang](https://github.com/XiangJinyu) (core authors), along with [@Zhaoyang Yu](https://github.com/MoshiQAQ), [@Jiayi Zhang](https://github.com/didiforgithub), and [@Sirui Hong](https://github.com/stellaHSR), we are from [@MetaGPT](https://github.com/geekan/MetaGPT). The prototype is launched within 3 hours and we are keeping building!
|
||||
|
||||
It's a simple implementation, so we welcome any suggestions, contributions, and feedback!
|
||||
|
||||
Enjoy your own agent with OpenManus!
|
||||
|
||||
We're also excited to introduce [OpenManus-RL](https://github.com/OpenManus/OpenManus-RL), an open-source project dedicated to reinforcement learning (RL)- based (such as GRPO) tuning methods for LLM agents, developed collaboratively by researchers from UIUC and OpenManus.
|
||||
|
||||
## Project Demo
|
||||
|
||||
<video src="https://private-user-images.githubusercontent.com/61239030/420168772-6dcfd0d2-9142-45d9-b74e-d10aa75073c6.mp4?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NDEzMTgwNTksIm5iZiI6MTc0MTMxNzc1OSwicGF0aCI6Ii82MTIzOTAzMC80MjAxNjg3NzItNmRjZmQwZDItOTE0Mi00NWQ5LWI3NGUtZDEwYWE3NTA3M2M2Lm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTAzMDclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwMzA3VDAzMjIzOVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTdiZjFkNjlmYWNjMmEzOTliM2Y3M2VlYjgyNDRlZDJmOWE3NWZhZjE1MzhiZWY4YmQ3NjdkNTYwYTU5ZDA2MzYmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.UuHQCgWYkh0OQq9qsUWqGsUbhG3i9jcZDAMeHjLt5T4" data-canonical-src="https://private-user-images.githubusercontent.com/61239030/420168772-6dcfd0d2-9142-45d9-b74e-d10aa75073c6.mp4?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NDEzMTgwNTksIm5iZiI6MTc0MTMxNzc1OSwicGF0aCI6Ii82MTIzOTAzMC80MjAxNjg3NzItNmRjZmQwZDItOTE0Mi00NWQ5LWI3NGUtZDEwYWE3NTA3M2M2Lm1wND9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTAzMDclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwMzA3VDAzMjIzOVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTdiZjFkNjlmYWNjMmEzOTliM2Y3M2VlYjgyNDRlZDJmOWE3NWZhZjE1MzhiZWY4YmQ3NjdkNTYwYTU5ZDA2MzYmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.UuHQCgWYkh0OQq9qsUWqGsUbhG3i9jcZDAMeHjLt5T4" controls="controls" muted="muted" class="d-block rounded-bottom-2 border-top width-fit" style="max-height:640px; min-height: 200px"></video>
|
||||
|
||||
## Installation
|
||||
|
||||
We provide two installation methods. Method 2 (using uv) is recommended for faster installation and better dependency management.
|
||||
|
||||
### Method 1: Using conda
|
||||
|
||||
1. Create a new conda environment:
|
||||
|
||||
```bash
|
||||
conda create -n open_manus python=3.12
|
||||
conda activate open_manus
|
||||
```
|
||||
|
||||
2. Clone the repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/mannaandpoem/OpenManus.git
|
||||
cd OpenManus
|
||||
```
|
||||
|
||||
3. Install dependencies:
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Method 2: Using uv (Recommended)
|
||||
|
||||
1. Install uv (A fast Python package installer and resolver):
|
||||
|
||||
```bash
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
```
|
||||
|
||||
2. Clone the repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/mannaandpoem/OpenManus.git
|
||||
cd OpenManus
|
||||
```
|
||||
|
||||
3. Create a new virtual environment and activate it:
|
||||
|
||||
```bash
|
||||
uv venv --python 3.12
|
||||
source .venv/bin/activate # On Unix/macOS
|
||||
# Or on Windows:
|
||||
# .venv\Scripts\activate
|
||||
```
|
||||
|
||||
4. Install dependencies:
|
||||
|
||||
```bash
|
||||
uv pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Browser Automation Tool (Optional)
|
||||
```bash
|
||||
playwright install
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
OpenManus requires configuration for the LLM APIs it uses. Follow these steps to set up your configuration:
|
||||
|
||||
1. Create a `config.toml` file in the `config` directory (you can copy from the example):
|
||||
|
||||
```bash
|
||||
cp config/config.example.toml config/config.toml
|
||||
```
|
||||
|
||||
2. Edit `config/config.toml` to add your API keys and customize settings:
|
||||
|
||||
```toml
|
||||
# Global LLM configuration
|
||||
[llm]
|
||||
model = "gpt-4o"
|
||||
base_url = "https://api.openai.com/v1"
|
||||
api_key = "sk-..." # Replace with your actual API key
|
||||
max_tokens = 4096
|
||||
temperature = 0.0
|
||||
|
||||
# Optional configuration for specific LLM models
|
||||
[llm.vision]
|
||||
model = "gpt-4o"
|
||||
base_url = "https://api.openai.com/v1"
|
||||
api_key = "sk-..." # Replace with your actual API key
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
One line for run OpenManus:
|
||||
|
||||
```bash
|
||||
python main.py
|
||||
```
|
||||
|
||||
Then input your idea via terminal!
|
||||
|
||||
For MCP tool version, you can run:
|
||||
```bash
|
||||
python run_mcp.py
|
||||
```
|
||||
|
||||
For unstable multi-agent version, you also can run:
|
||||
|
||||
```bash
|
||||
python run_flow.py
|
||||
```
|
||||
|
||||
## How to contribute
|
||||
|
||||
We welcome any friendly suggestions and helpful contributions! Just create issues or submit pull requests.
|
||||
|
||||
Or contact @mannaandpoem via 📧email: mannaandpoem@gmail.com
|
||||
|
||||
**Note**: Before submitting a pull request, please use the pre-commit tool to check your changes. Run `pre-commit run --all-files` to execute the checks.
|
||||
|
||||
## Community Group
|
||||
Join our networking group on Feishu and share your experience with other developers!
|
||||
|
||||
<div align="center" style="display: flex; gap: 20px;">
|
||||
<img src="assets/community_group.jpg" alt="OpenManus 交流群" width="300" />
|
||||
</div>
|
||||
|
||||
## Star History
|
||||
|
||||
[](https://star-history.com/#mannaandpoem/OpenManus&Date)
|
||||
|
||||
## Acknowledgement
|
||||
|
||||
Thanks to [anthropic-computer-use](https://github.com/anthropics/anthropic-quickstarts/tree/main/computer-use-demo)
|
||||
and [browser-use](https://github.com/browser-use/browser-use) for providing basic support for this project!
|
||||
|
||||
Additionally, we are grateful to [AAAJ](https://github.com/metauto-ai/agent-as-a-judge), [MetaGPT](https://github.com/geekan/MetaGPT), [OpenHands](https://github.com/All-Hands-AI/OpenHands) and [SWE-agent](https://github.com/SWE-agent/SWE-agent).
|
||||
|
||||
OpenManus is built by contributors from MetaGPT. Huge thanks to this agent community!
|
||||
|
||||
## Cite
|
||||
```bibtex
|
||||
@misc{openmanus2025,
|
||||
author = {Xinbin Liang and Jinyu Xiang and Zhaoyang Yu and Jiayi Zhang and Sirui Hong},
|
||||
title = {OpenManus: An open-source framework for building general AI agents},
|
||||
year = {2025},
|
||||
publisher = {GitHub},
|
||||
journal = {GitHub repository},
|
||||
howpublished = {\url{https://github.com/mannaandpoem/OpenManus}},
|
||||
}
|
||||
```
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
from app.agent.base import BaseAgent
|
||||
from app.agent.browser import BrowserAgent
|
||||
from app.agent.cot import CoTAgent
|
||||
from app.agent.mcp import MCPAgent
|
||||
from app.agent.planning import PlanningAgent
|
||||
from app.agent.react import ReActAgent
|
||||
@@ -11,7 +10,6 @@ from app.agent.toolcall import ToolCallAgent
|
||||
__all__ = [
|
||||
"BaseAgent",
|
||||
"BrowserAgent",
|
||||
"CoTAgent",
|
||||
"PlanningAgent",
|
||||
"ReActAgent",
|
||||
"SWEAgent",
|
||||
|
||||
@@ -1,47 +0,0 @@
|
||||
from typing import Optional
|
||||
|
||||
from pydantic import Field
|
||||
|
||||
from app.agent.base import BaseAgent
|
||||
from app.llm import LLM
|
||||
from app.logger import logger
|
||||
from app.prompt.cot import NEXT_STEP_PROMPT, SYSTEM_PROMPT
|
||||
from app.schema import AgentState, Message
|
||||
|
||||
|
||||
class CoTAgent(BaseAgent):
|
||||
"""Chain of Thought Agent - Focuses on demonstrating the thinking process of large language models without executing tools"""
|
||||
|
||||
name: str = "cot"
|
||||
description: str = "An agent that uses Chain of Thought reasoning"
|
||||
|
||||
system_prompt: str = SYSTEM_PROMPT
|
||||
next_step_prompt: Optional[str] = NEXT_STEP_PROMPT
|
||||
|
||||
llm: LLM = Field(default_factory=LLM)
|
||||
|
||||
max_steps: int = 1 # CoT typically only needs one step to complete reasoning
|
||||
|
||||
async def step(self) -> str:
|
||||
"""Execute one step of chain of thought reasoning"""
|
||||
logger.info(f"🧠 {self.name} is thinking...")
|
||||
|
||||
# If next_step_prompt exists and this isn't the first message, add it to user messages
|
||||
if self.next_step_prompt and len(self.messages) > 1:
|
||||
self.memory.add_message(Message.user_message(self.next_step_prompt))
|
||||
|
||||
# Use system prompt and user messages
|
||||
response = await self.llm.ask(
|
||||
messages=self.messages,
|
||||
system_msgs=[Message.system_message(self.system_prompt)]
|
||||
if self.system_prompt
|
||||
else None,
|
||||
)
|
||||
|
||||
# Record assistant's response
|
||||
self.memory.add_message(Message.assistant_message(response))
|
||||
|
||||
# Set state to finished after completion
|
||||
self.state = AgentState.FINISHED
|
||||
|
||||
return response
|
||||
@@ -0,0 +1,64 @@
|
||||
import json
|
||||
import uuid
|
||||
|
||||
from openai.types.chat.chat_completion_message_tool_call import (
|
||||
ChatCompletionMessageToolCall,
|
||||
Function,
|
||||
)
|
||||
|
||||
from app.agent.toolcall import ToolCallAgent
|
||||
from app.logger import logger
|
||||
from app.tool import LatexGenerator, ToolCollection, Validator
|
||||
|
||||
|
||||
class PPTAgent(ToolCallAgent):
|
||||
"""
|
||||
Agent that executes a fixed sequence of tools, potentially terminating
|
||||
early if the validator tool indicates completion.
|
||||
"""
|
||||
|
||||
name: str = "fixed_toolcall"
|
||||
description: str = (
|
||||
"an agent that executes a fixed sequence of tools in predefined order, "
|
||||
"potentially terminating early based on validator feedback."
|
||||
)
|
||||
|
||||
available_tools: ToolCollection = ToolCollection(LatexGenerator(), Validator())
|
||||
|
||||
max_steps: int = 7
|
||||
curr_step: int = 0
|
||||
|
||||
async def think(self) -> bool:
|
||||
"""Process current state and decide next actions using tools"""
|
||||
# pick which of your tools to call
|
||||
tool_idx = self.curr_step % len(self.available_tools.tools)
|
||||
tool_meta = self.available_tools.tools[tool_idx]
|
||||
|
||||
payload = {
|
||||
"request": self.memory.messages[0].content,
|
||||
"history": str(self.memory.messages),
|
||||
}
|
||||
arg_str = json.dumps(payload)
|
||||
|
||||
# build the Function descriptor
|
||||
func_call = Function(
|
||||
name=tool_meta.name,
|
||||
arguments=arg_str,
|
||||
)
|
||||
|
||||
# generate a proper call ID
|
||||
call_id = f"call_{uuid.uuid4().hex}"
|
||||
|
||||
# wrap it up in a ChatCompletionMessageToolCall
|
||||
tool_call = ChatCompletionMessageToolCall(
|
||||
id=call_id, function=func_call, type="function"
|
||||
)
|
||||
|
||||
# assign to self.tool_calls just like the SDK would
|
||||
self.tool_calls = tool_calls = [tool_call]
|
||||
logger.info(
|
||||
f"🛠️ {self.name} selected {len(tool_calls) if tool_calls else 0} tools to use"
|
||||
)
|
||||
self.curr_step += 1
|
||||
|
||||
return True
|
||||
@@ -37,30 +37,6 @@ class ProxySettings(BaseModel):
|
||||
|
||||
class SearchSettings(BaseModel):
|
||||
engine: str = Field(default="Google", description="Search engine the llm to use")
|
||||
fallback_engines: List[str] = Field(
|
||||
default_factory=lambda: ["DuckDuckGo", "Baidu"],
|
||||
description="Fallback search engines to try if the primary engine fails",
|
||||
)
|
||||
retry_delay: int = Field(
|
||||
default=60,
|
||||
description="Seconds to wait before retrying all engines again after they all fail",
|
||||
)
|
||||
max_retries: int = Field(
|
||||
default=3,
|
||||
description="Maximum number of times to retry all engines when all fail",
|
||||
)
|
||||
api_key: Optional[str] = Field(
|
||||
None,
|
||||
description="API key for the search engine's official API (currently used for Google)",
|
||||
)
|
||||
cx: Optional[str] = Field(
|
||||
None,
|
||||
description="Custom Search Engine ID for search APIs that require it (currently used for Google)",
|
||||
)
|
||||
use_fallback: bool = Field(
|
||||
True,
|
||||
description="Whether to fall back to web scraping when the API fails or is not configured",
|
||||
)
|
||||
|
||||
|
||||
class BrowserSettings(BaseModel):
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
from abc import ABC, abstractmethod
|
||||
from enum import Enum
|
||||
from typing import Dict, List, Optional, Union
|
||||
|
||||
from pydantic import BaseModel
|
||||
@@ -6,6 +7,10 @@ from pydantic import BaseModel
|
||||
from app.agent.base import BaseAgent
|
||||
|
||||
|
||||
class FlowType(str, Enum):
|
||||
PLANNING = "planning"
|
||||
|
||||
|
||||
class BaseFlow(BaseModel, ABC):
|
||||
"""Base class for execution flows supporting multiple agents"""
|
||||
|
||||
@@ -55,3 +60,32 @@ class BaseFlow(BaseModel, ABC):
|
||||
@abstractmethod
|
||||
async def execute(self, input_text: str) -> str:
|
||||
"""Execute the flow with given input"""
|
||||
|
||||
|
||||
class PlanStepStatus(str, Enum):
|
||||
"""Enum class defining possible statuses of a plan step"""
|
||||
|
||||
NOT_STARTED = "not_started"
|
||||
IN_PROGRESS = "in_progress"
|
||||
COMPLETED = "completed"
|
||||
BLOCKED = "blocked"
|
||||
|
||||
@classmethod
|
||||
def get_all_statuses(cls) -> list[str]:
|
||||
"""Return a list of all possible step status values"""
|
||||
return [status.value for status in cls]
|
||||
|
||||
@classmethod
|
||||
def get_active_statuses(cls) -> list[str]:
|
||||
"""Return a list of values representing active statuses (not started or in progress)"""
|
||||
return [cls.NOT_STARTED.value, cls.IN_PROGRESS.value]
|
||||
|
||||
@classmethod
|
||||
def get_status_marks(cls) -> Dict[str, str]:
|
||||
"""Return a mapping of statuses to their marker symbols"""
|
||||
return {
|
||||
cls.COMPLETED.value: "[✓]",
|
||||
cls.IN_PROGRESS.value: "[→]",
|
||||
cls.BLOCKED.value: "[!]",
|
||||
cls.NOT_STARTED.value: "[ ]",
|
||||
}
|
||||
|
||||
@@ -1,15 +1,10 @@
|
||||
from enum import Enum
|
||||
from typing import Dict, List, Union
|
||||
|
||||
from app.agent.base import BaseAgent
|
||||
from app.flow.base import BaseFlow
|
||||
from app.flow.base import BaseFlow, FlowType
|
||||
from app.flow.planning import PlanningFlow
|
||||
|
||||
|
||||
class FlowType(str, Enum):
|
||||
PLANNING = "planning"
|
||||
|
||||
|
||||
class FlowFactory:
|
||||
"""Factory for creating different types of flows with support for multiple agents"""
|
||||
|
||||
|
||||
+1
-31
@@ -1,47 +1,17 @@
|
||||
import json
|
||||
import time
|
||||
from enum import Enum
|
||||
from typing import Dict, List, Optional, Union
|
||||
|
||||
from pydantic import Field
|
||||
|
||||
from app.agent.base import BaseAgent
|
||||
from app.flow.base import BaseFlow
|
||||
from app.flow.base import BaseFlow, PlanStepStatus
|
||||
from app.llm import LLM
|
||||
from app.logger import logger
|
||||
from app.schema import AgentState, Message, ToolChoice
|
||||
from app.tool import PlanningTool
|
||||
|
||||
|
||||
class PlanStepStatus(str, Enum):
|
||||
"""Enum class defining possible statuses of a plan step"""
|
||||
|
||||
NOT_STARTED = "not_started"
|
||||
IN_PROGRESS = "in_progress"
|
||||
COMPLETED = "completed"
|
||||
BLOCKED = "blocked"
|
||||
|
||||
@classmethod
|
||||
def get_all_statuses(cls) -> list[str]:
|
||||
"""Return a list of all possible step status values"""
|
||||
return [status.value for status in cls]
|
||||
|
||||
@classmethod
|
||||
def get_active_statuses(cls) -> list[str]:
|
||||
"""Return a list of values representing active statuses (not started or in progress)"""
|
||||
return [cls.NOT_STARTED.value, cls.IN_PROGRESS.value]
|
||||
|
||||
@classmethod
|
||||
def get_status_marks(cls) -> Dict[str, str]:
|
||||
"""Return a mapping of statuses to their marker symbols"""
|
||||
return {
|
||||
cls.COMPLETED.value: "[✓]",
|
||||
cls.IN_PROGRESS.value: "[→]",
|
||||
cls.BLOCKED.value: "[!]",
|
||||
cls.NOT_STARTED.value: "[ ]",
|
||||
}
|
||||
|
||||
|
||||
class PlanningFlow(BaseFlow):
|
||||
"""A flow that manages planning and execution of tasks using agents."""
|
||||
|
||||
|
||||
+23
-7
@@ -1,19 +1,30 @@
|
||||
import logging
|
||||
import sys
|
||||
|
||||
|
||||
logging.basicConfig(level=logging.INFO, handlers=[logging.StreamHandler(sys.stderr)])
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import atexit
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import sys
|
||||
from inspect import Parameter, Signature
|
||||
from typing import Any, Dict, Optional
|
||||
|
||||
from mcp.server.fastmcp import FastMCP
|
||||
|
||||
from app.logger import logger
|
||||
|
||||
# Add directories to Python path (needed for proper importing)
|
||||
current_dir = os.path.dirname(os.path.abspath(__file__))
|
||||
parent_dir = os.path.dirname(current_dir)
|
||||
root_dir = os.path.dirname(parent_dir)
|
||||
sys.path.insert(0, parent_dir)
|
||||
sys.path.insert(0, current_dir)
|
||||
sys.path.insert(0, root_dir)
|
||||
|
||||
# Configure logging (using the same format as original)
|
||||
logging.basicConfig(
|
||||
level=logging.INFO, format="%(asctime)s - %(name)s - %(levelname)s - %(message)s"
|
||||
)
|
||||
logger = logging.getLogger("mcp-server")
|
||||
|
||||
from app.tool.base import BaseTool
|
||||
from app.tool.bash import Bash
|
||||
from app.tool.browser_use_tool import BrowserUseTool
|
||||
@@ -34,6 +45,11 @@ class MCPServer:
|
||||
self.tools["editor"] = StrReplaceEditor()
|
||||
self.tools["terminate"] = Terminate()
|
||||
|
||||
from app.logger import logger as app_logger
|
||||
|
||||
global logger
|
||||
logger = app_logger
|
||||
|
||||
def register_tool(self, tool: BaseTool, method_name: Optional[str] = None) -> None:
|
||||
"""Register a tool with parameter validation and documentation."""
|
||||
tool_name = method_name or tool.name
|
||||
|
||||
@@ -1,15 +0,0 @@
|
||||
SYSTEM_PROMPT = """You are an assistant focused on Chain of Thought reasoning. For each question, please follow these steps:
|
||||
|
||||
1. Break down the problem: Divide complex problems into smaller, more manageable parts
|
||||
2. Think step by step: Think through each part in detail, showing your reasoning process
|
||||
3. Synthesize conclusions: Integrate the thinking from each part into a complete solution
|
||||
4. Provide an answer: Give a final concise answer
|
||||
|
||||
Your response should follow this format:
|
||||
Thinking: [Detailed thought process, including problem decomposition, reasoning for each step, and analysis]
|
||||
Answer: [Final answer based on the thought process, clear and concise]
|
||||
|
||||
Remember, the thinking process is more important than the final answer, as it demonstrates how you reached your conclusion.
|
||||
"""
|
||||
|
||||
NEXT_STEP_PROMPT = "Please continue your thinking based on the conversation above. If you've reached a conclusion, provide your final answer."
|
||||
@@ -0,0 +1,49 @@
|
||||
SYSTEM_PROMPT = """
|
||||
You are a LaTeX Beamer Presentation Generator. Your task is to generate a complete, informative, and ready-to-compile Beamer slide deck in LaTeX, based on the task description and any past drafts or feedback.
|
||||
|
||||
## Goals:
|
||||
- Each slide must be **self-contained**, meaning the audience should understand the slide without external explanations.
|
||||
- The presentation must **teach** or **explain** the topic in sufficient detail using structured LaTeX slides.
|
||||
- Each slide must contribute meaningfully to the overall structure and flow of the presentation.
|
||||
|
||||
|
||||
|
||||
## Requirements:
|
||||
|
||||
1. Preamble & Setup
|
||||
- Start with `\\documentclass{beamer}`.
|
||||
- Use packages such as `amsmath`, `amsfonts`, and `graphicx`.
|
||||
- Use the `Madrid` theme unless otherwise specified.
|
||||
- Include full metadata: `\\title{}`, `\\author{}`, and `\\date{\\today}`.
|
||||
|
||||
2. Slide Design
|
||||
- MUST mark each slide with a comment indicating its number, `% Slide 1`, `% Slide 2`.
|
||||
- - Slides must follow a **logical order** that ensures smooth flow and coherence.
|
||||
- AIM for a **minimum of 300 words per slide* Contain **enough detail** (text, bullets, equations, definitions, or examples)
|
||||
|
||||
3. Depth of Content
|
||||
- For important concept, include motivation, problem, intuitive explanation, mathematical formulation or equation (if applicable)
|
||||
- practical example or application can also be included
|
||||
|
||||
4. Completeness & Validity
|
||||
- Reflect all provided feedback and correct deficiencies from past versions.
|
||||
- MUST No placeholders or incomplete content.
|
||||
- Your output will be used directly. Therefore, it must be a ready-to-use result.
|
||||
- Include `\\end{document}`.
|
||||
- Ensure valid LaTeX syntax.
|
||||
|
||||
5. Style & Clarity
|
||||
- Maintain consistent formatting and indentation.
|
||||
- Use bullet points or short paragraphs for clarity.
|
||||
- Keep math readable and contextualized with supporting text.
|
||||
|
||||
**Only output the final LaTeX source code. Do not include explanations, notes, or comments.**
|
||||
"""
|
||||
|
||||
USER_CONTENT = """
|
||||
## Task
|
||||
{request}
|
||||
|
||||
## Past Drafts & Feedback
|
||||
{history}
|
||||
"""
|
||||
@@ -0,0 +1,42 @@
|
||||
TEXT_VALIDATION_PROMPT = """
|
||||
You are a task result evaluator responsible for determining whether a task result meets the task requirements, if not, you need to improve it.
|
||||
|
||||
# Objective and Steps
|
||||
1. **Completeness and Quality Check:**
|
||||
- Verify that the result includes all required elements of the task.
|
||||
- Evaluate whether the output meets overall quality criteria (accuracy, clarity, formatting, and completeness).
|
||||
|
||||
2. **Change Detection:**
|
||||
- If this is a subsequent result, compare it with previous iterations.
|
||||
- If the differences are minimal or the result has not significantly improved, consider it "good enough" for finalization.
|
||||
|
||||
3. **Feedback and Escalation:**
|
||||
- If the result meets the criteria or the improvements are negligible compared to previous iterations, return **"No further feedback"**.
|
||||
- Otherwise, provide **direct and precise feedback** and **output the improved result in the required format** for finalization.
|
||||
|
||||
4. **Ensure Completeness:**
|
||||
- Your output must meet all requirements of the task.
|
||||
- Include all necessary details so that the output is self-contained and can be directly used as input for downstream tasks.
|
||||
|
||||
5. **Do NOT:**
|
||||
- Leave any section with placeholders (e.g., "TODO", "Add content here").
|
||||
- Include any commentary or reminders to the writer or user (e.g., "We can add more later").
|
||||
- Output partial slides or omit essential details assuming future input.
|
||||
|
||||
- **If the result meets the standard:**
|
||||
- Return **"No further feedback."**.
|
||||
|
||||
- **If the result does not meet the standard:**
|
||||
- add detailed jusification for the change start with "here are some feedbacks" and directly write an improved new result start with "here are the changes".
|
||||
|
||||
# Note that: Any output containing incomplete sections, placeholders is not allowed.
|
||||
"""
|
||||
USER_CONTENT = """
|
||||
## Current Task Requirement:
|
||||
{request}
|
||||
|
||||
---
|
||||
|
||||
## Current Task Latest Result:
|
||||
{history}
|
||||
"""
|
||||
@@ -96,6 +96,35 @@ class Message(BaseModel):
|
||||
message["base64_image"] = self.base64_image
|
||||
return message
|
||||
|
||||
def to_string(self) -> str:
|
||||
"""
|
||||
Convert the Message instance into a human-readable string format.
|
||||
|
||||
Returns:
|
||||
str: A formatted string representing the message.
|
||||
"""
|
||||
# Format the header with role and name
|
||||
role_str = self.role.upper() if self.role else "UNKNOWN"
|
||||
name_str = f" ({self.name})" if self.name else ""
|
||||
header = f"[{role_str}{name_str}]"
|
||||
|
||||
# Start with the message content
|
||||
content = self.content or ""
|
||||
|
||||
# Append tool call details if available
|
||||
if self.tool_calls:
|
||||
tool_calls_str = "\n".join(
|
||||
f" ↳ ToolCall: {tc.function.name}({tc.function.arguments}) [id={tc.id}]"
|
||||
for tc in self.tool_calls
|
||||
)
|
||||
content += "\n" + tool_calls_str
|
||||
|
||||
# Append information about attached image if any
|
||||
if self.base64_image:
|
||||
content += "\n ↳ [Image Attached: base64 content hidden]"
|
||||
|
||||
return f"{header}\n{content.strip()}"
|
||||
|
||||
@classmethod
|
||||
def user_message(
|
||||
cls, content: str, base64_image: Optional[str] = None
|
||||
@@ -182,3 +211,12 @@ class Memory(BaseModel):
|
||||
def to_dict_list(self) -> List[dict]:
|
||||
"""Convert messages to list of dicts"""
|
||||
return [msg.to_dict() for msg in self.messages]
|
||||
|
||||
def to_string(self) -> str:
|
||||
"""
|
||||
Convert the memory's list of messages to a readable string format.
|
||||
|
||||
Returns:
|
||||
str: A formatted string representing the entire conversation history.
|
||||
"""
|
||||
return "\n\n".join(msg.to_string() for msg in self.messages)
|
||||
|
||||
@@ -2,10 +2,12 @@ from app.tool.base import BaseTool
|
||||
from app.tool.bash import Bash
|
||||
from app.tool.browser_use_tool import BrowserUseTool
|
||||
from app.tool.create_chat_completion import CreateChatCompletion
|
||||
from app.tool.latex_generator import LatexGenerator
|
||||
from app.tool.planning import PlanningTool
|
||||
from app.tool.str_replace_editor import StrReplaceEditor
|
||||
from app.tool.terminate import Terminate
|
||||
from app.tool.tool_collection import ToolCollection
|
||||
from app.tool.validator import Validator
|
||||
|
||||
|
||||
__all__ = [
|
||||
@@ -13,8 +15,11 @@ __all__ = [
|
||||
"Bash",
|
||||
"BrowserUseTool",
|
||||
"Terminate",
|
||||
"ValiTerminate",
|
||||
"StrReplaceEditor",
|
||||
"ToolCollection",
|
||||
"CreateChatCompletion",
|
||||
"PlanningTool",
|
||||
"Validator",
|
||||
"LatexGenerator",
|
||||
]
|
||||
|
||||
@@ -448,22 +448,6 @@ Page content:
|
||||
"extracted_content": {
|
||||
"type": "object",
|
||||
"description": "The content extracted from the page according to the goal",
|
||||
"properties": {
|
||||
"text": {
|
||||
"type": "string",
|
||||
"description": "Text content extracted from the page",
|
||||
},
|
||||
"metadata": {
|
||||
"type": "object",
|
||||
"description": "Additional metadata about the extracted content",
|
||||
"properties": {
|
||||
"source": {
|
||||
"type": "string",
|
||||
"description": "Source of the extracted content",
|
||||
}
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
},
|
||||
"required": ["extracted_content"],
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
from pydantic import Field
|
||||
|
||||
from app.llm import LLM
|
||||
from app.prompt.latex_generator import SYSTEM_PROMPT, USER_CONTENT
|
||||
from app.tool.base import BaseTool
|
||||
|
||||
|
||||
_Latex_Generator_DESCRIPTION = """
|
||||
This agent generates complete, high-quality LaTeX documents with a focus on Beamer presentations. It accepts topic-specific input and produces fully self-contained LaTeX source code, including all required packages, structures, and rich content elements such as equations, figures, and formatted text. The agent ensures completeness by avoiding any placeholders or incomplete sections.
|
||||
|
||||
In addition to generation, the agent supports iterative refinement: it evaluates and improves the generated LaTeX code based on validation feedback to ensure correctness, formatting quality, and logical structure. The final output is ready for immediate compilation and professional presentation use.
|
||||
"""
|
||||
|
||||
|
||||
class LatexGenerator(BaseTool):
|
||||
llm: LLM = Field(default_factory=LLM, description="Language model instance")
|
||||
name: str = "latexgenerator"
|
||||
description: str = _Latex_Generator_DESCRIPTION
|
||||
parameters: dict = {}
|
||||
|
||||
async def generate(self, request: str, history: str = ""):
|
||||
"""Abstract method for result validate logic.
|
||||
|
||||
Args:
|
||||
step_result: The result string to validate.
|
||||
"""
|
||||
system_content = SYSTEM_PROMPT
|
||||
user_content = USER_CONTENT.format(request=request, history=history)
|
||||
|
||||
feedback = await self.llm.ask(
|
||||
messages=[{"role": "user", "content": user_content}],
|
||||
system_msgs=[{"role": "system", "content": system_content}],
|
||||
)
|
||||
return feedback
|
||||
|
||||
async def execute(self, request: str, history: str = "") -> str:
|
||||
"""Finish the current execution"""
|
||||
return await self.generate(request, history)
|
||||
@@ -1,156 +1,9 @@
|
||||
from typing import List
|
||||
|
||||
import requests
|
||||
from googlesearch import search
|
||||
|
||||
from app.config import config
|
||||
from app.logger import logger
|
||||
from app.tool.search.base import WebSearchEngine
|
||||
|
||||
|
||||
class GoogleSearchEngine(WebSearchEngine):
|
||||
def perform_search(
|
||||
self, query: str, num_results: int = 10, *args, **kwargs
|
||||
) -> List[str]:
|
||||
"""
|
||||
Google search engine using the official Google Custom Search API when configured,
|
||||
falling back to web scraping if not configured or if the API call fails.
|
||||
|
||||
Args:
|
||||
query (str): The search query to submit to the search engine.
|
||||
num_results (int, optional): The number of search results to return. Default is 10.
|
||||
*args: Additional positional arguments.
|
||||
**kwargs: Additional keyword arguments.
|
||||
|
||||
Returns:
|
||||
List[str]: A list of URLs matching the search query.
|
||||
"""
|
||||
# Check for API configuration in the search settings
|
||||
search_config = getattr(config, "search_config", None)
|
||||
api_key = getattr(search_config, "api_key", None) if search_config else None
|
||||
cx = getattr(search_config, "cx", None) if search_config else None
|
||||
use_fallback = (
|
||||
getattr(search_config, "use_fallback", True) if search_config else True
|
||||
)
|
||||
|
||||
# If API is configured, try using the Google Search API
|
||||
if api_key and cx:
|
||||
try:
|
||||
logger.info("Using Google Custom Search API for search")
|
||||
return self._api_search(query, api_key, cx, num_results)
|
||||
except requests.RequestException as e:
|
||||
# More specific error handling for HTTP-related errors
|
||||
status_code = (
|
||||
getattr(e.response, "status_code", None)
|
||||
if hasattr(e, "response")
|
||||
else None
|
||||
)
|
||||
if status_code == 429:
|
||||
logger.warning("Google API rate limit exceeded")
|
||||
elif status_code and 400 <= status_code < 500:
|
||||
logger.warning(
|
||||
f"Google API client error: {e} (status code: {status_code})"
|
||||
)
|
||||
elif status_code and 500 <= status_code < 600:
|
||||
logger.warning(
|
||||
f"Google API server error: {e} (status code: {status_code})"
|
||||
)
|
||||
else:
|
||||
logger.warning(f"Google API request error: {e}")
|
||||
|
||||
if not use_fallback:
|
||||
logger.warning(
|
||||
"Fallback to scraping is disabled. Returning empty results."
|
||||
)
|
||||
return []
|
||||
logger.info("Falling back to web scraping search")
|
||||
except Exception as e:
|
||||
# General error handling for other types of exceptions
|
||||
logger.warning(f"Google API error: {e}")
|
||||
if not use_fallback:
|
||||
logger.warning(
|
||||
"Fallback to scraping is disabled. Returning empty results."
|
||||
)
|
||||
return []
|
||||
logger.info("Falling back to web scraping search")
|
||||
|
||||
# Use web scraping if API is not configured or if API call failed and fallback is enabled
|
||||
return self._scraping_search(query, num_results)
|
||||
|
||||
@staticmethod
|
||||
def _api_search(
|
||||
query: str, api_key: str, cx: str, num_results: int = 10
|
||||
) -> List[str]:
|
||||
"""
|
||||
Perform a search using Google's Custom Search JSON API.
|
||||
|
||||
Args:
|
||||
query (str): The search query.
|
||||
api_key (str): The API key for Google Custom Search.
|
||||
cx (str): The Custom Search Engine ID.
|
||||
num_results (int, optional): The number of results to return. Default is 10.
|
||||
|
||||
Returns:
|
||||
List[str]: A list of URLs matching the search query.
|
||||
|
||||
Raises:
|
||||
requests.RequestException: If there's an issue with the HTTP request.
|
||||
ValueError: If the response cannot be parsed as JSON.
|
||||
"""
|
||||
base_url = "https://www.googleapis.com/customsearch/v1"
|
||||
results = []
|
||||
|
||||
# API allows max 10 results per request, so we need to paginate
|
||||
for start_index in range(
|
||||
1, min(num_results + 1, 101), 10
|
||||
): # Google API limits to 100 results max
|
||||
params = {
|
||||
"q": query,
|
||||
"key": api_key,
|
||||
"cx": cx,
|
||||
"start": start_index,
|
||||
"num": min(
|
||||
10, num_results - len(results)
|
||||
), # Can't request more than 10 at once
|
||||
}
|
||||
|
||||
response = requests.get(base_url, params=params, timeout=10) # Add timeout
|
||||
response.raise_for_status() # Raise exception for 4XX/5XX responses
|
||||
data = response.json()
|
||||
|
||||
if "items" in data:
|
||||
for item in data["items"]:
|
||||
if "link" in item:
|
||||
results.append(item["link"])
|
||||
if len(results) >= num_results:
|
||||
return results
|
||||
else:
|
||||
# No more results or empty result set
|
||||
if (
|
||||
"searchInformation" in data
|
||||
and "totalResults" in data["searchInformation"]
|
||||
):
|
||||
logger.info(
|
||||
f"Total results: {data['searchInformation']['totalResults']}"
|
||||
)
|
||||
break
|
||||
|
||||
return results
|
||||
|
||||
@staticmethod
|
||||
def _scraping_search(query: str, num_results: int = 10) -> List[str]:
|
||||
"""
|
||||
Perform a search using web scraping as a fallback method.
|
||||
|
||||
Args:
|
||||
query (str): The search query.
|
||||
num_results (int, optional): The number of results to return. Default is 10.
|
||||
|
||||
Returns:
|
||||
List[str]: A list of URLs matching the search query.
|
||||
"""
|
||||
try:
|
||||
return list(search(query, num_results=num_results))
|
||||
except Exception as e:
|
||||
logger.warning(f"Web scraping search failed: {e}")
|
||||
return []
|
||||
def perform_search(self, query, num_results=10, *args, **kwargs):
|
||||
"""Google search engine."""
|
||||
return search(query, num_results=num_results)
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
from pydantic import Field
|
||||
|
||||
from app.llm import LLM
|
||||
from app.prompt.validator import TEXT_VALIDATION_PROMPT, USER_CONTENT
|
||||
from app.tool.base import BaseTool
|
||||
|
||||
|
||||
_VALIDATE_DESCRIPTION = """
|
||||
This tool evaluates the quality and completeness of a subtask result against a set of predefined criteria.
|
||||
It checks whether the result fully satisfies task requirements, maintains high quality in terms of clarity, accuracy, and formatting,
|
||||
and determines whether improvements have been made in comparison to prior versions.
|
||||
If the result is satisfactory or improvements are minimal, it returns "The step result has already reached the requirement.".
|
||||
Otherwise, it provides detailed feedback and a revised version of the result that meets all requirements and is ready for downstream use.
|
||||
"""
|
||||
|
||||
|
||||
class Validator(BaseTool):
|
||||
llm: LLM = Field(default_factory=LLM, description="Language model instance")
|
||||
name: str = "validator"
|
||||
description: str = _VALIDATE_DESCRIPTION
|
||||
parameters: dict = {}
|
||||
|
||||
async def validate(self, request: str, history: str):
|
||||
"""Abstract method for result validate logic.
|
||||
|
||||
Args:
|
||||
step_result: The result string to validate.
|
||||
"""
|
||||
system_content = TEXT_VALIDATION_PROMPT
|
||||
user_content = USER_CONTENT.format(request=request, history=history)
|
||||
|
||||
feedback = await self.llm.ask(
|
||||
messages=[{"role": "user", "content": user_content}],
|
||||
system_msgs=[{"role": "system", "content": system_content}],
|
||||
)
|
||||
return feedback
|
||||
|
||||
async def execute(self, request: str, history: str) -> str:
|
||||
"""Finish the current execution"""
|
||||
return await self.validate(request, history)
|
||||
+7
-82
@@ -4,7 +4,6 @@ from typing import List
|
||||
from tenacity import retry, stop_after_attempt, wait_exponential
|
||||
|
||||
from app.config import config
|
||||
from app.logger import logger
|
||||
from app.tool.base import BaseTool
|
||||
from app.tool.search import (
|
||||
BaiduSearchEngine,
|
||||
@@ -45,8 +44,6 @@ class WebSearch(BaseTool):
|
||||
async def execute(self, query: str, num_results: int = 10) -> List[str]:
|
||||
"""
|
||||
Execute a Web search and return a list of URLs.
|
||||
Tries engines in order based on configuration, falling back if an engine fails with errors.
|
||||
If all engines fail, it will wait and retry up to the configured number of times.
|
||||
|
||||
Args:
|
||||
query (str): The search query to submit to the search engine.
|
||||
@@ -55,109 +52,37 @@ class WebSearch(BaseTool):
|
||||
Returns:
|
||||
List[str]: A list of URLs matching the search query.
|
||||
"""
|
||||
# Get retry settings from config
|
||||
retry_delay = 60 # Default to 60 seconds
|
||||
max_retries = 3 # Default to 3 retries
|
||||
|
||||
if config.search_config:
|
||||
retry_delay = getattr(config.search_config, "retry_delay", 60)
|
||||
max_retries = getattr(config.search_config, "max_retries", 3)
|
||||
|
||||
# Try searching with retries when all engines fail
|
||||
for retry_count in range(
|
||||
max_retries + 1
|
||||
): # +1 because first try is not a retry
|
||||
links = await self._try_all_engines(query, num_results)
|
||||
if links:
|
||||
return links
|
||||
|
||||
if retry_count < max_retries:
|
||||
# All engines failed, wait and retry
|
||||
logger.warning(
|
||||
f"All search engines failed. Waiting {retry_delay} seconds before retry {retry_count + 1}/{max_retries}..."
|
||||
)
|
||||
await asyncio.sleep(retry_delay)
|
||||
else:
|
||||
logger.error(
|
||||
f"All search engines failed after {max_retries} retries. Giving up."
|
||||
)
|
||||
|
||||
return []
|
||||
|
||||
async def _try_all_engines(self, query: str, num_results: int) -> List[str]:
|
||||
"""
|
||||
Try all search engines in the configured order.
|
||||
|
||||
Args:
|
||||
query (str): The search query to submit to the search engine.
|
||||
num_results (int): The number of search results to return.
|
||||
|
||||
Returns:
|
||||
List[str]: A list of URLs matching the search query, or empty list if all engines fail.
|
||||
"""
|
||||
engine_order = self._get_engine_order()
|
||||
failed_engines = []
|
||||
|
||||
for engine_name in engine_order:
|
||||
engine = self._search_engine[engine_name]
|
||||
try:
|
||||
logger.info(f"🔎 Attempting search with {engine_name.capitalize()}...")
|
||||
links = await self._perform_search_with_engine(
|
||||
engine, query, num_results
|
||||
)
|
||||
if links:
|
||||
if failed_engines:
|
||||
logger.info(
|
||||
f"Search successful with {engine_name.capitalize()} after trying: {', '.join(failed_engines)}"
|
||||
)
|
||||
return links
|
||||
except Exception as e:
|
||||
failed_engines.append(engine_name.capitalize())
|
||||
is_rate_limit = "429" in str(e) or "Too Many Requests" in str(e)
|
||||
|
||||
if is_rate_limit:
|
||||
logger.warning(
|
||||
f"⚠️ {engine_name.capitalize()} search engine rate limit exceeded, trying next engine..."
|
||||
)
|
||||
else:
|
||||
logger.warning(
|
||||
f"⚠️ {engine_name.capitalize()} search failed with error: {e}"
|
||||
)
|
||||
|
||||
if failed_engines:
|
||||
logger.error(f"All search engines failed: {', '.join(failed_engines)}")
|
||||
print(f"Search engine '{engine_name}' failed with error: {e}")
|
||||
return []
|
||||
|
||||
def _get_engine_order(self) -> List[str]:
|
||||
"""
|
||||
Determines the order in which to try search engines.
|
||||
Preferred engine is first (based on configuration), followed by fallback engines,
|
||||
and then the remaining engines.
|
||||
Preferred engine is first (based on configuration), followed by the remaining engines.
|
||||
|
||||
Returns:
|
||||
List[str]: Ordered list of search engine names.
|
||||
"""
|
||||
preferred = "google"
|
||||
fallbacks = []
|
||||
|
||||
if config.search_config:
|
||||
if config.search_config.engine:
|
||||
preferred = config.search_config.engine.lower()
|
||||
if config.search_config.fallback_engines:
|
||||
fallbacks = [
|
||||
engine.lower() for engine in config.search_config.fallback_engines
|
||||
]
|
||||
if config.search_config and config.search_config.engine:
|
||||
preferred = config.search_config.engine.lower()
|
||||
|
||||
engine_order = []
|
||||
# Add preferred engine first
|
||||
if preferred in self._search_engine:
|
||||
engine_order.append(preferred)
|
||||
|
||||
# Add configured fallback engines in order
|
||||
for fallback in fallbacks:
|
||||
if fallback in self._search_engine and fallback not in engine_order:
|
||||
engine_order.append(fallback)
|
||||
|
||||
for key in self._search_engine:
|
||||
if key not in engine_order:
|
||||
engine_order.append(key)
|
||||
return engine_order
|
||||
|
||||
@retry(
|
||||
|
||||
@@ -73,21 +73,6 @@ temperature = 0.0 # Controls randomness for vision mod
|
||||
# [search]
|
||||
# Search engine for agent to use. Default is "Google", can be set to "Baidu" or "DuckDuckGo".
|
||||
#engine = "Google"
|
||||
# Fallback engine order. Default is ["DuckDuckGo", "Baidu"] - will try in this order after primary engine fails.
|
||||
#fallback_engines = ["DuckDuckGo", "Baidu"]
|
||||
# Seconds to wait before retrying all engines again when they all fail due to rate limits. Default is 60.
|
||||
#retry_delay = 60
|
||||
# Maximum number of times to retry all engines when all fail. Default is 3.
|
||||
#max_retries = 3
|
||||
# API key for the search engine's official API (currently used for Google)
|
||||
# For Google, create an API key at https://console.cloud.google.com/apis/credentials
|
||||
#api_key = ""
|
||||
# Custom Search Engine ID for search APIs that require it (currently used for Google)
|
||||
# For Google, create a Custom Search Engine at https://programmablesearchengine.google.com/
|
||||
#cx = ""
|
||||
# Whether to fall back to web scraping when the API fails or is not configured. Default is true.
|
||||
#use_fallback = true
|
||||
|
||||
|
||||
## Sandbox configuration
|
||||
#[sandbox]
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
import asyncio
|
||||
|
||||
from app.agent.ppt import PPTAgent
|
||||
from app.logger import logger
|
||||
|
||||
|
||||
async def main():
|
||||
agent = PPTAgent()
|
||||
try:
|
||||
prompt = """
|
||||
1. Lecture slide:
|
||||
I am a lecturer. I am teaching the machine learning coure for research students. Please generate latex code for lecture slide for different reinforcement learning algorithms.
|
||||
Note that:
|
||||
1). Note that the lecture duration is 2 hour, so we need to generate 30 pages.
|
||||
2). for each reinforcement learning algorithms, the slide should include motivation, problem and intuitive solution and detailed math equations.
|
||||
3). Please make sure the the lecture have a good self-contain.
|
||||
"""
|
||||
if not prompt.strip():
|
||||
logger.warning("Empty prompt provided.")
|
||||
return
|
||||
|
||||
logger.warning("Processing your request...")
|
||||
await agent.run(prompt)
|
||||
logger.info("Request processing completed.")
|
||||
except KeyboardInterrupt:
|
||||
logger.warning("Operation interrupted.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
|
Before Width: | Height: | Size: 164 KiB After Width: | Height: | Size: 164 KiB |
|
Before Width: | Height: | Size: 36 KiB After Width: | Height: | Size: 36 KiB |
@@ -7,7 +7,6 @@ numpy
|
||||
datasets~=3.2.0
|
||||
fastapi~=0.115.11
|
||||
tiktoken~=0.9.0
|
||||
requests~=2.31.0
|
||||
|
||||
html2text~=2024.2.26
|
||||
gymnasium~=1.0.0
|
||||
|
||||
+2
-1
@@ -2,7 +2,8 @@ import asyncio
|
||||
import time
|
||||
|
||||
from app.agent.manus import Manus
|
||||
from app.flow.flow_factory import FlowFactory, FlowType
|
||||
from app.flow.base import FlowType
|
||||
from app.flow.flow_factory import FlowFactory
|
||||
from app.logger import logger
|
||||
|
||||
|
||||
|
||||
+6
-15
@@ -13,14 +13,10 @@ class MCPRunner:
|
||||
|
||||
def __init__(self):
|
||||
self.root_path = config.root_path
|
||||
self.server_reference = "app.mcp.server"
|
||||
self.server_script = self.root_path / "app" / "mcp" / "server.py"
|
||||
self.agent = MCPAgent()
|
||||
|
||||
async def initialize(
|
||||
self,
|
||||
connection_type: str,
|
||||
server_url: str | None = None,
|
||||
) -> None:
|
||||
async def initialize(self, connection_type: str, server_url: str = None) -> None:
|
||||
"""Initialize the MCP agent with the appropriate connection."""
|
||||
logger.info(f"Initializing MCPAgent with {connection_type} connection...")
|
||||
|
||||
@@ -28,7 +24,7 @@ class MCPRunner:
|
||||
await self.agent.initialize(
|
||||
connection_type="stdio",
|
||||
command=sys.executable,
|
||||
args=["-m", self.server_reference],
|
||||
args=[str(self.server_script)],
|
||||
)
|
||||
else: # sse
|
||||
await self.agent.initialize(connection_type="sse", server_url=server_url)
|
||||
@@ -51,14 +47,9 @@ class MCPRunner:
|
||||
|
||||
async def run_default(self) -> None:
|
||||
"""Run the agent in default mode."""
|
||||
prompt = input("Enter your prompt: ")
|
||||
if not prompt.strip():
|
||||
logger.warning("Empty prompt provided.")
|
||||
return
|
||||
|
||||
logger.warning("Processing your request...")
|
||||
await self.agent.run(prompt)
|
||||
logger.info("Request processing completed.")
|
||||
await self.agent.run(
|
||||
"Hello, what tools are available to me? Terminate after you have listed the tools."
|
||||
)
|
||||
|
||||
async def cleanup(self) -> None:
|
||||
"""Clean up agent resources."""
|
||||
|
||||
@@ -1,11 +0,0 @@
|
||||
# coding: utf-8
|
||||
# A shortcut to launch OpenManus MCP server, where its introduction also solves other import issues.
|
||||
from app.mcp.server import MCPServer, parse_args
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
args = parse_args()
|
||||
|
||||
# Create and run server (maintaining original flow)
|
||||
server = MCPServer()
|
||||
server.run(transport=args.transport)
|
||||
Reference in New Issue
Block a user