If you’ve used Claude Code, Cursor, or Aider and wondered what’s inside, the answer is smaller than you’d think: a while loop around one HTTP request. This post builds a working coding agent in a single Python file of about 300 lines, with requests as the only dependency. By the end you’ll understand every line, and you’ll have an agent that reads your code, edits it, and runs your tests.
The full code is on GitHub: Fazel94/agent_500_lines (agent.py).
What you’ll build
Here’s the finished agent fixing a bug. The scratch directory holds a fizzbuzz.py whose loop is range(1, n) instead of range(1, n + 1), and a two-assertion pytest file that catches it:
$ python ../agent.py --yes "run the tests, fix the bug, run them again"
→ list_files({"path": "tests"})
→ list_files({"path": "."})
→ bash({"command": "pytest"})
→ read_file({"path": "fizzbuzz.py"})
→ edit_file({"path": "fizzbuzz.py", "old": "for i in range(1, n):", "new": "for i in range(1)
→ bash({"command": "pytest"})
I fixed the bug in the `fizzbuzz` function by changing the loop range from
`range(1, n)` to `range(1, n + 1)`, ensuring it includes the number `n`. After
making this change, I ran the tests again, and all of them passed successfully.
That took seventeen seconds and six tool calls with gpt-4o-mini, one of the cheapest models available. Each → line is the agent telling you which tool the model asked for. The first one is a wrong guess: there is no tests/ directory. The model got an error back, looked at . instead, and kept going. (The edit_file line is cut off because the agent truncates its own trace at 80 characters.)
The agent talks to any OpenAI-compatible API. OpenAI’s chat request format has become a de facto standard, so OpenRouter, Together, vLLM, llama.cpp’s server, and Ollama all accept the same request at /v1/chat/completions. Switching providers means changing one environment variable.
The whole idea is a loop
Modern chat APIs support tool calling: alongside your messages, you send the model a list of functions it’s allowed to call. Instead of answering in prose, the model can reply “call read_file with path=fizzbuzz.py”. Your code runs that function and sends the result back as a new message. The model never runs anything itself; it only asks.
An agent is that exchange on repeat:
messages = [system_prompt, user_prompt]
loop:
reply = POST /chat/completions with messages + tool schemas
append reply to messages
if reply has no tool_calls: print reply.content and stop
for each call in reply.tool_calls:
result = run the matching python function
append {"role": "tool", "tool_call_id": call.id, "content": result}
Everything else in the file supports this loop: one function that makes the HTTP call, seven tools, a decorator that describes each tool to the model, a check that keeps file access inside your project, and a y/n prompt before anything risky.
One HTTP call
The entire network layer:
def chat(messages: list, tools: list) -> dict:
"""One blocking round-trip. Returns the assistant message dict."""
try:
r = requests.post(f"{BASE_URL}/chat/completions", timeout=120,
headers={"Authorization": f"Bearer {API_KEY}"},
json={"model": MODEL, "messages": messages,
"tools": tools, "tool_choice": "auto"})
r.raise_for_status()
return r.json()["choices"][0]["message"]
except requests.HTTPError:
raise RuntimeError(f"API {r.status_code}: {r.text[:500]}")
except requests.RequestException as e: # DNS, TLS, timeout, connection reset
raise RuntimeError(f"API unreachable: {e}")
tool_choice: "auto" lets the model choose between calling a tool and answering. Both failure paths become a RuntimeError with a printable message, because the only sensible reaction is “print it and keep the session alive”. A rate limit or a dropped connection shouldn’t cost you your conversation. One subtlety: requests doesn’t raise on a 4xx or 5xx status by itself, which is why raise_for_status() is there, and HTTPError is a subclass of RequestException, so those two except clauses must stay in this order.
The request is plain JSON:
{"model": "gpt-4o-mini",
"tools": [{"type": "function",
"function": {
"name": "read_file",
"description": "Read a file...",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]}}}],
"tool_choice": "auto",
"messages": [{"role": "system", "content": "You are a coding agent..."},
{"role": "user", "content": "run the tests, fix the bug"}]}
And a reply that wants a tool looks like this:
{"role": "assistant", "content": null,
"tool_calls": [{"id": "call_a1b2", "type": "function",
"function": {"name": "read_file",
"arguments": "{\"path\": \"fizzbuzz.py\"}"}}]}
Two details trip people up here.
First, arguments is a JSON string, not an object. The model writes it as text, so it can be malformed, and parsing it is your job. Some local servers send a dict instead; the agent accepts both.
Second, the id must come back. For every tool call the model sent, you have to send one message with role: "tool" and a matching tool_call_id. Miss one and the API rejects your next request outright.
Turning Python functions into tools
The model learns what a tool does from a description written in JSON Schema, a standard format for saying “this object has these fields, of these types, and these ones are required”. Writing that schema by hand next to every function is tedious, and the two drift apart. So a decorator builds it from the function signature:
TOOLS: dict = {} # name -> python function
SCHEMAS: list = [] # what the model sees, in registration order
JSON_TYPES = {str: "string", int: "integer", bool: "boolean"}
def tool(description: str, **params: str):
"""Register fn as a tool. The JSON type comes from each parameter's
annotation, a parameter with no default is required, and params holds
one line of help per parameter."""
def register(fn):
properties, required = {}, []
for name, p in inspect.signature(fn).parameters.items():
properties[name] = {"type": JSON_TYPES[p.annotation],
"description": params[name]}
if p.default is p.empty:
required.append(name)
SCHEMAS.append({"type": "function", "function": {
"name": fn.__name__, "description": description,
"parameters": {"type": "object", "properties": properties,
"required": required}}})
TOOLS[fn.__name__] = fn
return fn
return register
The type annotation gives the JSON type, a parameter without a default is required, and the decorator’s keyword arguments give one line of help per parameter. SCHEMAS is what gets sent to the model, and TOOLS[name] is how a call finds its function. Adding a tool means writing one decorated function.
If you forget a help string, or use an annotation the table doesn’t know, you get a KeyError when the file loads, before the model ever sees a broken schema.
The seven core tools
Every tool that touches a file first passes its path through safe_path(), which keeps the agent inside your project directory. More on that in Keeping it safe.
Looking around: read_file, list_files, grep
@tool("Read a file. Returns its contents with 1-based line numbers.",
path="file path relative to the workspace root",
start="first line to return (1-based, default 1)",
end="last line to return (default: end of file)")
def read_file(path: str, start: int = 1, end: int = None):
f = safe_path(path)
if not f.is_file():
return "error: not found"
lines = f.read_text(errors="replace").splitlines()
start = max(int(start), 1)
end = len(lines) if end is None else min(int(end), len(lines))
out = "\n".join(f"{i}: {lines[i - 1]}" for i in range(start, end + 1))
if end < len(lines):
out += f"\n...[{len(lines) - end} more lines; read them with start={end + 1}]"
return out or "(empty range)"
Line numbers let the model talk about the file (“the loop on line 4”), and you can check what it says against your editor. They’re for display only, though: edit_file matches text, not line numbers, so the model can’t damage a file by miscounting.
start and end handle files too big to read in one go. Without them, a 4,000-line file would be cut off mid-line and the model would fall back to sed -n, losing the numbering. Notice that the cut-off message says exactly what to call next (read them with start=...). That’s worth copying in any tool: when you have to cut something off, tell the model how to get the rest.
The two search tools share a file walker that skips directories which would otherwise flood the model with noise:
def walk(base: pathlib.Path):
"""Every file under base, skipping noise directories."""
for path in sorted(base.rglob("*")):
if path.is_file() and not SKIP_DIRS & set(path.relative_to(ROOT).parts):
yield path
@tool("List files recursively, optionally filtered by a glob on the filename.",
path="directory to list (default '.')",
pattern="glob matched against the file name, e.g. '*.py'")
def list_files(path: str = ".", pattern: str = "*"):
base = safe_path(path)
if not base.is_dir():
return "error: not a directory"
hits = [str(f.relative_to(ROOT)) for f in walk(base)
if fnmatch.fnmatch(f.name, pattern)]
out = "\n".join(hits[:500])
if len(hits) > 500:
out += f"\n...and {len(hits) - 500} more"
return out or "(no matches)"
@tool("Search file contents with a Python regular expression.",
regex="Python regex",
path="file or directory to search (default '.')")
def grep(regex: str, path: str = "."):
rx = re.compile(regex)
base = safe_path(path)
files = [base] if base.is_file() else list(walk(base))
hits = []
for f in files:
try:
text = f.read_text()
except (UnicodeDecodeError, OSError):
continue # binary or unreadable: skip
for i, line in enumerate(text.splitlines(), 1):
if rx.search(line):
hits.append(f"{f.relative_to(ROOT)}:{i}: {line.strip()}")
if len(hits) >= 200:
return "\n".join(hits) + "\n...(truncated at 200 matches)"
return "\n".join(hits) or "(no matches)"
Both tools cap their output, at 500 files and 200 matches. grep prints path:line: text, the format the model has seen countless times in its training data, so the system prompt doesn’t need to explain it.
Changing things: write_file, edit_file
@tool("Create or overwrite a file with exact content.",
path="file path relative to the workspace root",
content="full file content")
def write_file(path: str, content: str):
f = safe_path(path)
f.parent.mkdir(parents=True, exist_ok=True)
f.write_text(content)
return f"wrote {len(content.encode())} bytes to {path}"
Blunt: it overwrites. mkdir(parents=True) means “create src/utils/parse.py” takes one call instead of a trip through the shell. The byte count in the reply lets the model tell an empty write from a real one.
@tool("Replace an exact substring in a file. `old` must appear exactly once.",
path="file path relative to the workspace root",
old="exact text to replace, with enough context to be unique",
new="replacement text")
def edit_file(path: str, old: str, new: str):
f = safe_path(path)
if not f.is_file():
return "error: not found"
text = f.read_text()
n = text.count(old)
if n == 0:
return "error: old text not found"
if n > 1:
return f"error: old text occurs {n} times; include more context"
f.write_text(text.replace(old, new))
return "ok"
This is the most important tool in the file, and the “exactly once” rule is why.
Diffs look like the obvious format for edits, but to produce one the model has to get line offsets and surrounding context exactly right. There are many ways to be almost right, and a patch tool either rejects the whole thing or, worse, applies it in the wrong place. Find-and-replace on an exact snippet can fail in only two ways: the text isn’t there, or it’s there more than once. Both error messages tell the model what to do next. Models are good at copying a snippet they just read and bad at counting lines.
Refusing an ambiguous match, instead of replacing the first one, is the difference between a tool that silently edits the wrong return None and one that asks the model to try again.
Running things: bash
@tool("Run a shell command in the workspace root. 60s timeout.",
command="shell command")
def bash(command: str):
try:
p = subprocess.run(command, shell=True, cwd=ROOT, text=True,
capture_output=True, timeout=60)
except subprocess.TimeoutExpired:
return "error: timed out after 60s"
return f"{p.stdout}{p.stderr}\n[exit {p.returncode}]"
Nine lines that replace every tool you didn’t write: git, the test runner, the formatter, the compiler. Two choices here are deliberate. Standard output and error are joined rather than labelled, because the model wants the story of what happened, in order. And the exit code is always appended; without it, a command that fails quietly looks like a success, and the agent will cheerfully tell you your tests pass.
Asking you: ask_user
@tool("Ask the human a question and wait for their answer. Use only for a "
"decision the code cannot answer, never to ask permission.",
question="a single, specific question")
def ask_user(question: str):
if AUTO_APPROVE:
return "error: running unattended (--yes); use your best judgement"
print(f"\n[?] {question}")
try:
return input("your answer: ").strip() or "(no answer given)"
except (EOFError, KeyboardInterrupt):
return "(no answer given)"
This is the one thing bash can’t do: no shell command puts a human decision into the conversation. Without it, when the model hits a real fork (which of two config files did you mean? is this dead code or a public API?), it can only guess or stop. With it, the work continues after your answer.
Two details keep it from getting annoying. Under --yes it returns an error telling the model to decide for itself, since an unattended run has nobody to answer and would otherwise hang forever. And the description’s “never to ask permission” matters: without it, models use this tool to ask about things the y/n prompt already covers.
Extras: reaching the web
The seven tools above are the agent; with them it’s about 300 lines. Two optional tools add about 60 more and let the model look things up: web_search finds pages, and fetch_url reads one. Leave them out and nothing else changes.
@tool("Search the web with DuckDuckGo. Returns ranked title/url/snippet "
"triples; the snippet is a teaser, so fetch_url the ones that matter.",
query="search query, as you would type it into a search box",
max_results="how many results to return (default 5, max 25)")
def web_search(query: str, max_results: int = 5):
try:
from ddgs import DDGS # pip install ddgs
except ImportError:
return "error: web search needs the ddgs package (pip install ddgs)"
n = max(1, min(int(max_results), 25))
hits = []
for _ in (1, 2): # a rotated backend comes back empty at random
try:
hits = DDGS(timeout=30).text(query, max_results=n) or []
except Exception as e: # rate limits, captchas, DNS
if "no results" not in str(e).lower(): # ddgs raises on empty too
return f"error: search failed: {type(e).__name__}: {e}"
if hits:
break
out = [f"{i}. {h.get('title', '')}\n {h.get('href', '')}\n "
f"{' '.join((h.get('body') or '').split())}"
for i, h in enumerate(hits, 1)]
return "\n".join(out) or "(no results)"
ddgs is a small library that queries DuckDuckGo without an API key. It’s imported inside the function, so the core agent still needs nothing beyond requests. If ddgs isn’t installed, the model gets an error message telling it so, rather than the agent crashing at startup. The search runs at most twice because ddgs sometimes returns nothing for no reason.
@tool("Fetch a URL and return it as plain text, scripts and tags stripped. "
"Use it to actually read a page web_search pointed at.",
url="absolute http:// or https:// URL",
start="character offset to resume from (default 0)")
def fetch_url(url: str, start: int = 0):
if not url.startswith(("http://", "https://")):
return "error: url must start with http:// or https://"
try:
r = requests.get(url, headers={"User-Agent": UA}, timeout=30, stream=True)
r.raise_for_status()
kind = r.headers.get("Content-Type", "").split(";")[0].strip()
if not (kind.startswith("text/") or "json" in kind or "xml" in kind):
return f"error: {kind} is not text"
enc = r.encoding if "charset=" in r.headers.get("Content-Type", "") else "utf-8"
body = r.raw.read(4_000_000, decode_content=True).decode(enc, "replace")
except requests.RequestException as e: # DNS, TLS, 404, timeout
return f"error: fetch failed: {type(e).__name__}: {e}"
if kind.startswith("text/html"):
body = re.sub(r"(?is)<(script|style|noscript|svg|head|nav|footer|aside"
r"|form)\b.*?</\1>", " ", body)
main = re.search(r"(?is)<(main|article)\b.*?</\1>", body) # skip the chrome
body = main.group(0) if main else body
body = re.sub(r"(?i)<br\s*/?>|</(p|div|li|tr|h[1-6])>", "\n", body)
body = html.unescape(re.sub(r"(?s)<!--.*?-->|<[^>]+>", " ", body))
text = "\n".join(ln.strip() for ln in re.sub(r"[ \t]+", " ", body).splitlines()
if ln.strip())
start = max(int(start), 0)
out = text[start:start + PAGE_CHARS]
if len(text) > start + PAGE_CHARS:
out += (f"\n...[{len(text) - start - PAGE_CHARS} more chars; "
f"continue with start={start + PAGE_CHARS}]")
return out or "(no text on page)"
Raw HTML wastes the model’s attention on menus, scripts, and styling, so fetch_url throws those away. If the page has a <main> or <article> element, it keeps only that, then strips the remaining tags. It’s crude next to a real HTML parser, but it needs no extra dependency and handles documentation pages and blog posts well. Like read_file, it returns one 10,000-character page at a time and tells the model which start to use for the next one. It reads at most 4 MB, and it refuses images, PDFs, and other non-text responses instead of feeding the model binary data.
Keeping it safe
Two checks stand between the model and your machine.
The first keeps file tools inside your project:
def safe_path(p: str) -> pathlib.Path:
"""Resolve p under ROOT, refusing anything that escapes the workspace."""
path = (ROOT / p).resolve()
if path != ROOT and not path.is_relative_to(ROOT):
raise ValueError(f"path escapes workspace: {p}")
return path
The order is what matters: resolve first, then compare. .resolve() collapses .. and follows symlinks, so both ../../etc/passwd and a symlink pointing at /etc are caught by the same two lines. Checking the string before resolving would catch neither.
The second asks you before anything changes:
def confirm(name: str, args: dict) -> bool:
"""Ask the human before anything that writes to disk or runs a command."""
if AUTO_APPROVE:
return True
preview = args.get("command") or args.get("path", "?")
print(f"\n[{name}] {preview}")
body = args.get("content") or args.get("new")
if body:
for line in body.splitlines()[:10]:
print(f" | {line}")
try:
return input("allow? [y/N] ").strip().lower() in ("y", "yes")
except (EOFError, KeyboardInterrupt):
return False
Anything other than y or yes, including Ctrl-C, means no. The set of tools that need permission is one line, GATED = {"write_file", "edit_file", "bash"}, and it’s checked in one place, run_tool. The tools themselves never ask, which keeps them simple to test. --yes skips the prompt for unattended runs such as CI.
A “no” isn’t an error. The model receives user denied this action as that tool’s result and carries on. In practice it does the sensible thing: it explains what it wanted to change and hands you the patch to apply yourself.
The web tools aren’t in GATED, because they read rather than write. One caveat: a URL can carry data out, and a confused or manipulated model could put file contents in a query string. If you point this agent at code you can’t afford to leak, add web_search and fetch_url to GATED.
Errors are messages for the model
def run_tool(call: dict) -> str:
"""Execute one tool call from the model. Every failure comes back as
text, so the model can read it and try something else."""
name = call["function"]["name"]
raw = call["function"]["arguments"]
try:
args = raw if isinstance(raw, dict) else json.loads(raw or "{}")
print(f"\u2192 {name}({json.dumps(args)[:80]})")
if name not in TOOLS:
return f"error: unknown tool {name}"
if name in GATED and not confirm(name, args):
return "user denied this action"
out = TOOLS[name](**args)
except Exception as e:
out = f"error: {type(e).__name__}: {e}"
if len(out) > MAX_TOOL_OUTPUT:
extra = len(out) - MAX_TOOL_OUTPUT
out = out[:MAX_TOOL_OUTPUT] + f"\n...[truncated {extra} chars]"
return out
No exception from a tool ever reaches the loop. A made-up tool name, malformed arguments, a missing parameter, a wrong type, or a path outside the project all become a short line of text in the conversation, and the model reads it and tries something else. That’s what happened with tests/ in the opening demo. Crashes turn into conversation.
The length cap is the other half. The context window is the most text the model can see in one request, and a single cat of a minified JavaScript bundle could fill it. So every tool result is cut off at 20,000 characters.
The real loop
Here’s the pseudocode from earlier, as real Python:
def run_turn(messages: list) -> None:
"""Drive the model until it answers with plain text (or hits the cap)."""
for _ in range(MAX_STEPS):
msg = chat(messages, SCHEMAS)
messages.append(msg) # verbatim: tool_calls must round-trip
if msg.get("content"):
print(msg["content"])
if not msg.get("tool_calls"):
return
for call in msg["tool_calls"]:
messages.append({"role": "tool", "tool_call_id": call["id"],
"content": run_tool(call)})
messages.append({"role": "user",
"content": "Stop: step limit reached. Summarize what you did."})
final = chat(messages, [])
messages.append(final)
print(final.get("content", ""))
Seventeen lines. Three things in them are worth a closer look.
The model’s message is stored exactly as received. Providers add extra fields (refusal, annotations, their own ids), and some check that what you send back matches what they sent. Store the message as-is and you’re always right. Rebuild it from the fields you care about and you’re right until the day you aren’t.
MAX_STEPS keeps a stuck model from running forever. Models do get stuck, reading the same file again or retrying a failing command with small variations. Fifty rounds is plenty for real work. When the limit hits, the agent asks for a summary and makes one last call with no tools, so the model has to answer in plain text.
The message list only grows. Nothing is ever edited or removed, so each request is the previous one plus some new messages at the end. That makes prompt caching work. Providers save their processed version of a request’s beginning, and when the next request starts with exactly the same bytes, they skip that work and charge less for it. It’s also why SCHEMAS is built in a fixed order: the tool list goes out identical on every call. The final step-limit call is the one place the agent breaks the cache deliberately, by sending an empty tool list, which is fine for something that happens at most once per turn.
main() is the thin part. Pass a prompt on the command line and it runs one turn and exits. Pass nothing and you get an interactive session that reuses one messages list, so history carries over from turn to turn:
messages.append({"role": "user", "content": line})
try:
run_turn(messages)
except RuntimeError as e: # API error: report it, keep the session
print(e)
Run it yourself
| Variable | Default | Meaning |
|---|---|---|
AGENT_BASE_URL |
https://api.openai.com/v1 |
Any OpenAI-compatible root |
AGENT_API_KEY |
falls back to OPENAI_API_KEY |
Bearer token; empty is fine locally |
AGENT_MODEL |
gpt-4o-mini |
Model id as the endpoint names it |
The agent works on the current directory, so cd into your project first.
pip install requests # add ddgs if you want web_search
# OpenAI
export OPENAI_API_KEY=sk-...
python agent.py "add type hints to utils.py and run mypy"
# OpenRouter
export AGENT_BASE_URL=https://openrouter.ai/api/v1
export AGENT_API_KEY=sk-or-...
export AGENT_MODEL=anthropic/claude-sonnet-4
python agent.py
# Ollama, no key needed
AGENT_BASE_URL=http://localhost:11434/v1 AGENT_MODEL=qwen2.5-coder \
python agent.py "explain what this repo does"
Try the opening demo without --yes and answer n when edit_file asks permission. The agent stops, leaves the file alone, and prints the fix for you to apply by hand.
What’s missing
Five things a production agent has that this one doesn’t, and roughly what each costs:
- Streaming. Right now the answer arrives in one lump. Ask the API to stream and you get the reply in small fragments as it’s generated, including tool calls, which arrive in pieces and have to be reassembled. About 30 lines, and it’s the change you’ll notice most.
- Context compaction. Long sessions eventually overflow the context window. The usual fix: once the conversation passes a size threshold, ask the model to summarize the oldest half and replace it with the summary. About 25 lines. It throws away the prompt cache, by design.
- Parallel tool calls. Models often ask for several independent tools in one message, and this agent runs them one at a time. Running the read-only tools in a thread pool, while keeping writes in order, takes about 15 lines and speeds up exploration-heavy turns.
- Retries. A rate-limit or server error currently ends the turn with a printed message. Wrap
chat()in a retry loop that waits longer each time and respects the server’sRetry-Afterheader. About 15 lines. - A repo map. Serious agents start with a compact outline of the codebase, every file with its top-level functions and classes, so the model doesn’t have to discover the structure by searching. A regex pass over
defandclasslines gets most of the value in about 30 lines.
None of them changes the loop. Each one is more code around it.
Takeaways
- The loop is trivial; the tools are the product.
edit_file’s error messages did more for reliability than any prompt tweak. - Errors belong in the conversation, where the model can read them and recover, not in a stack trace.
- Only ever append to the message list. That one rule keeps tool-call ids matched, keeps prompt caching working, and lets you debug a whole session with
print(json.dumps(messages, indent=2)).
The commit history follows the same order as this post. The first commit is the skeleton plus read_file, the next five add one core tool each, and later commits add the permission prompt, line ranges for read_file, ask_user, and the web tools. Run git show on any of them to see one piece on its own.
About 300 lines. Read them, delete what you don’t need, and add the tool your workflow is missing.