|
1 | | -# AI Tool Calling from Scratch |
| 1 | +# AI Tool Calling |
2 | 2 |
|
3 | 3 | Provider-agnostic Python library for TOOL_CALL-style tool execution across multiple LLM/vision backends. The library exposes typed primitives (`Agent`, `Tool`, `Message`, `Role`), pluggable providers (OpenAI adapter + Anthropic/Gemini/Local), streaming support, a hardened TOOL_CALL parser, prompt template, CLI entrypoint, bounding-box demo tool, and helper examples. |
4 | 4 |
|
| 5 | +## Why this library |
| 6 | + |
| 7 | +- Provider-agnostic tool calling with schema validation and a hardened parser (handles fenced/mixed JSON). |
| 8 | +- Robustness controls: retries with rate-limit backoff, request timeouts, per-tool execution timeouts, iteration caps. |
| 9 | +- Streaming or one-shot responses; vision supported where providers allow. |
| 10 | +- Simple ergonomics: `@tool`/`ToolRegistry`, CLI for one-offs or chat, ready-made examples and tests. |
| 11 | + |
5 | 12 | ## What's Included |
6 | 13 |
|
7 | 14 | - Core package at `src/toolcalling/` with agent loop, parser, prompt builder, and provider adapters |
8 | | -- OpenAI provider implementation plus Anthropic/Gemini/Local providers sharing the same interface |
9 | | -- CLI (`toolcalling`) to list tools, run one-offs, or chat interactively with streaming |
10 | | -- Bounding-box detection tool (OpenAI Vision) and demo runner in `scripts/chat.py` |
11 | | -- Tests with fake providers to validate schemas, parsing, and agent wiring |
| 15 | +- Providers: OpenAI plus Anthropic/Gemini/Local sharing the same interface |
| 16 | +- Library-first examples (see below) and tests with fake providers for schemas, parsing, agent wiring |
12 | 17 | - PyPI-ready metadata (`pyproject.toml`) using a src-layout package |
13 | 18 |
|
14 | 19 | ## Install |
@@ -41,50 +46,173 @@ response = agent.run([Message(role=Role.USER, content="Search for Backtrack")]) |
41 | 46 | print(response.content) |
42 | 47 | ``` |
43 | 48 |
|
44 | | -## CLI |
45 | | - |
46 | | -```bash |
47 | | -# List available tools (echo + detect_bounding_box) |
48 | | -toolcalling list-tools |
49 | | - |
50 | | -# Run a one-off prompt (supports --stream, --dry-run) |
51 | | -toolcalling run --provider openai --model gpt-4o --prompt "Say hello" --tool echo |
| 49 | +## Common ways to use it (library-first) |
52 | 50 |
|
53 | | -# Interactive chat with history (streams tokens if enabled) |
54 | | -toolcalling chat --provider local --stream |
| 51 | +- Define tools (`Tool` or `@tool`/`ToolRegistry`), pick a provider, run `Agent.run([...])`. |
| 52 | +- Add vision by supplying `image_path` on `Message` when the provider supports it. |
| 53 | +- For offline/testing: use the Local provider and/or `TOOLCALLING_BBOX_MOCK_JSON=tests/fixtures/bbox_mock.json`. |
| 54 | +- Optional dev helpers (not required for library use): `scripts/smoke_cli.py` for quick provider smokes; `scripts/chat.py` for the vision demo. |
55 | 55 |
|
56 | | -# Vision run with an image |
57 | | -toolcalling run --prompt "Find the object" --image assets/environment.png --tool detect_bounding_box |
58 | | -``` |
| 56 | +## Providers (incl. vision & limits) |
59 | 57 |
|
60 | | -## Bounding Box Demo |
61 | | - |
62 | | -``` |
63 | | -python scripts/chat.py # processes assets/environment.png |
64 | | -python scripts/chat.py --interactive |
65 | | -``` |
66 | | - |
67 | | -The demo uses the bounding-box tool backed by OpenAI Vision and writes `*_with_bbox.png` alongside the input image. |
68 | | -For offline tests, set `TOOLCALLING_BBOX_MOCK_JSON=tests/fixtures/bbox_mock.json` to bypass the network. |
| 58 | +- OpenAI: streaming; vision via Chat Completions `image_url` (e.g., `gpt-5`); request timeout default 30s; retries/backoff via `AgentConfig`. |
| 59 | +- Anthropic: streaming; vision model-dependent; set `ANTHROPIC_API_KEY`. |
| 60 | +- Gemini: streaming; vision model-dependent; set `GEMINI_API_KEY`. |
| 61 | +- Local: no network; echoes latest user text; no vision. |
| 62 | +- Rate limits: agent detects `rate limit`/`429` and backs off + retries. |
| 63 | +- Timeouts: `AgentConfig.request_timeout` (provider) and `tool_timeout_seconds` (per tool). |
| 64 | + |
| 65 | +## Agent config at a glance |
| 66 | + |
| 67 | +- Core: `model`, `temperature`, `max_tokens`, `max_iterations`. |
| 68 | +- Reliability: `max_retries`, `retry_backoff_seconds`, rate-limit backoff, `request_timeout`. |
| 69 | +- Execution safety: `tool_timeout_seconds` to bound tool runtime. |
| 70 | +- Streaming: `stream=True` to stream provider deltas; optional `stream_handler` callback. |
| 71 | + |
| 72 | +## Library examples |
| 73 | + |
| 74 | +- Simple search tool (text): |
| 75 | + |
| 76 | + ```python |
| 77 | + from toolcalling import Agent, AgentConfig, Message, Role, tool |
| 78 | + from toolcalling.providers.openai_provider import OpenAIProvider |
| 79 | + |
| 80 | + @tool(description="Echo input") |
| 81 | + def echo(text: str) -> str: |
| 82 | + return text |
| 83 | + |
| 84 | + agent = Agent(tools=[echo], provider=OpenAIProvider(), config=AgentConfig(max_iterations=3)) |
| 85 | + resp = agent.run([Message(role=Role.USER, content="Hello!")]) |
| 86 | + print(resp.content) |
| 87 | + ``` |
| 88 | + |
| 89 | +- Vision bounding box (uses your image): |
| 90 | + |
| 91 | + ```python |
| 92 | + from toolcalling import Agent, AgentConfig, Message, Role |
| 93 | + from toolcalling.examples.bbox import create_bounding_box_tool |
| 94 | + |
| 95 | + bbox_tool = create_bounding_box_tool() |
| 96 | + agent = Agent(tools=[bbox_tool], config=AgentConfig(max_iterations=5, model="gpt-4o")) |
| 97 | + resp = agent.run([ |
| 98 | + Message(role=Role.USER, content="Find the object", image_path="path/to/your/image.png") |
| 99 | + ]) |
| 100 | + print(resp.content) |
| 101 | + ``` |
| 102 | + |
| 103 | +- Multi-tool workflow (search + summarize): |
| 104 | + |
| 105 | + ```python |
| 106 | + from toolcalling import Agent, AgentConfig, Message, Role, tool |
| 107 | + from toolcalling.providers.gemini_provider import GeminiProvider |
| 108 | + |
| 109 | + @tool(description="Search the web") |
| 110 | + def search(query: str) -> str: |
| 111 | + return f"results for {query}" |
| 112 | + |
| 113 | + @tool(description="Summarize text") |
| 114 | + def summarize(text: str) -> str: |
| 115 | + return f"summary: {text[:200]}" |
| 116 | + |
| 117 | + agent = Agent( |
| 118 | + tools=[search, summarize], |
| 119 | + provider=GeminiProvider(), |
| 120 | + config=AgentConfig(max_iterations=4, stream=False), |
| 121 | + ) |
| 122 | + resp = agent.run([Message(role=Role.USER, content="Find and summarize AI tool-calling docs")]) |
| 123 | + print(resp.content) |
| 124 | + ``` |
| 125 | + |
| 126 | +- Local/offline dry-run (no network): |
| 127 | + |
| 128 | + ```python |
| 129 | + from toolcalling import Agent, AgentConfig, Message, Role, tool |
| 130 | + from toolcalling.providers.stubs import LocalProvider |
| 131 | + |
| 132 | + @tool(description="Format a todo item") |
| 133 | + def todo(item: str, priority: str = "medium") -> str: |
| 134 | + return f"[{priority}] {item}" |
| 135 | + |
| 136 | + agent = Agent( |
| 137 | + tools=[todo], |
| 138 | + provider=LocalProvider(), |
| 139 | + config=AgentConfig(max_iterations=2, model="local"), |
| 140 | + ) |
| 141 | + resp = agent.run([Message(role=Role.USER, content="Add buy milk to my list")]) |
| 142 | + print(resp.content) |
| 143 | + ``` |
| 144 | + |
| 145 | +- Streaming responses: |
| 146 | + |
| 147 | + ```python |
| 148 | + from toolcalling import Agent, AgentConfig, Message, Role, tool |
| 149 | + from toolcalling.providers.openai_provider import OpenAIProvider |
| 150 | + |
| 151 | + @tool(description="Generate a short outline") |
| 152 | + def outline(topic: str) -> str: |
| 153 | + return f"Outline for {topic}" |
| 154 | + |
| 155 | + chunks = [] |
| 156 | + def on_chunk(text: str): |
| 157 | + chunks.append(text) |
| 158 | + |
| 159 | + agent = Agent( |
| 160 | + tools=[outline], |
| 161 | + provider=OpenAIProvider(), |
| 162 | + config=AgentConfig(stream=True, max_iterations=2), |
| 163 | + ) |
| 164 | + resp = agent.run([Message(role=Role.USER, content="Draft an outline about LLM safety")], stream_handler=on_chunk) |
| 165 | + print("".join(chunks)) |
| 166 | + ``` |
| 167 | + |
| 168 | +- Per-tool config injection (e.g., API keys per tool): |
| 169 | + |
| 170 | + ```python |
| 171 | + from toolcalling import Agent, AgentConfig, Message, Role, Tool |
| 172 | + from toolcalling.tools import ToolParameter |
| 173 | + from toolcalling.providers.openai_provider import OpenAIProvider |
| 174 | + |
| 175 | + def call_weather_api(city: str, api_key: str) -> str: |
| 176 | + # placeholder for real HTTP call |
| 177 | + return f"Weather for {city} using key {api_key[:4]}***" |
| 178 | + |
| 179 | + weather_tool = Tool( |
| 180 | + name="weather_lookup", |
| 181 | + description="Gets weather for a city", |
| 182 | + parameters=[ToolParameter(name="city", param_type=str, description="City name")], |
| 183 | + function=call_weather_api, |
| 184 | + injected_kwargs={"api_key": "YOUR_WEATHER_API_KEY"}, |
| 185 | + ) |
| 186 | + |
| 187 | + agent = Agent( |
| 188 | + tools=[weather_tool], |
| 189 | + provider=OpenAIProvider(), |
| 190 | + config=AgentConfig(max_iterations=2), |
| 191 | + ) |
| 192 | + resp = agent.run([Message(role=Role.USER, content="What's the weather in Paris?")]) |
| 193 | + print(resp.content) |
| 194 | + ``` |
69 | 195 |
|
70 | 196 | ## Tool ergonomics |
71 | 197 |
|
72 | 198 | - Use `ToolRegistry` or the `@tool` decorator to infer schemas from function signatures and register tools. |
73 | 199 | - Inject per-tool config or auth using `injected_kwargs` or `config_injector` when constructing a `Tool`. |
| 200 | +- Type hints map to JSON schema; defaults make parameters optional. |
74 | 201 |
|
75 | 202 | ## Tests |
76 | 203 |
|
77 | 204 | ```bash |
78 | 205 | python tests/test_framework.py |
79 | 206 | ``` |
80 | 207 |
|
| 208 | +- Covers parsing (mixed/fenced), agent loop (retries/streaming), provider mocks (Anthropic/Gemini), CLI streaming, bbox mock path, and tool schema basics. |
| 209 | + |
81 | 210 | ## Packaging |
82 | 211 |
|
83 | 212 | The project ships a `pyproject.toml` with console scripts and a src layout. Adjust version/metadata before publishing to PyPI. |
84 | 213 | CI workflow (`.github/workflows/ci.yml`) runs tests, build, and twine check. Tags matching `v*` attempt TestPyPI/PyPI publishes when tokens are provided. |
85 | 214 |
|
86 | 215 | ## More docs |
87 | 216 |
|
88 | | -- See `docs/USER_GUIDE.md` for provider/tool/agent/CLI/demo/testing/release details. |
89 | | -- Smoke the CLI across providers (skips if env vars missing): `python scripts/smoke_cli.py`. |
90 | | -- Extra example (no network): `python examples/search_weather.py`. |
| 217 | +- Single source of truth is this README. |
| 218 | +- Optional dev helpers: `python scripts/smoke_cli.py` (skips providers missing keys), `python scripts/chat.py` (vision demo), `python examples/search_weather.py` (local mock tools). |
0 commit comments