A coding agent that works in a demo is easy; one you can use all day is a different program. If you’ve built an agent loop around a chat API, or wondered why tools like Claude Code run to thousands of lines, this post shows the five fixes that close most of that gap: streaming, retries, saved sessions, context compaction, and Ctrl-C that does the right thing. Each one is small, and the hard part of four of them turns out to be the same rule.
The full code is on GitHub: Fazel94/agent_500_lines (harness.py).
The starting point
A minimal coding agent is a loop. You send a chat API your conversation plus a list of tools (Python functions the model may ask you to run, such as read_file, edit_file, and bash). The model either answers in text or replies with tool calls. You run the tools, append their results to the conversation, and send it again, until the model stops asking for tools.
My version of that, agent.py, has seven core tools, a y/n prompt before anything that writes or runs commands, and about 300 lines in all. (I walk through it line by line in Build a coding agent in about 300 lines of Python, but you don’t need that post for this one.) It works: gpt-4o-mini, a model nobody calls good at code, used it to find and fix an off-by-one bug in six tool calls.
Use it for a day and the cracks show:
- Silence. The answer arrives in one lump after twenty seconds of nothing.
- Fragility. One rate-limit error from the API ends your turn.
- Amnesia. Close the terminal and the session is gone.
- A size limit. A long task fills the model’s context window and dies.
- A blunt Ctrl-C. It kills the whole program when you only meant to stop the command it was running.
None of these are model problems; a smarter model fixes none of them. They’re harness problems, and harness.py fixes all five in 145 extra lines. It’s still one file, and still needs only requests.
What is a harness?
The model is a pure function: (messages, tools) → message. It’s stateless. It has no memory, no hands, and no idea what happened thirty seconds ago. Everything that makes it feel like a colleague instead of a text predictor is code you write around it:
graph TD
U[human] --> P[presentation: streaming, traces, prompts]
P --> C[control loop: steps, MAX_STEPS, cancellation]
C --> X[context: transcript, compaction, persistence]
C --> G[policy: GATED, confirm, safe_path]
G --> D[dispatch: TOOLS registry, run_tool]
D --> T[capabilities: seven core tools + two web extras]
C --> N[transport: streaming, retries, error mapping]
N --> M[model API]
That’s the harness. The minimal agent has a few lines in most of these boxes. This post is about what happens when they have to survive daily use.
Throughout, transcript means the conversation list sent to the model on every request: the system prompt, your messages, the model’s replies, and one tool message answering each tool call.
Streaming
Twenty seconds of silence makes a tool feel dead. Streaming means asking the API for "stream": true. Instead of one JSON reply, the server then sends a series of small events (Server-Sent Events, or SSE: plain text lines starting with data:), each carrying a delta, the next few characters of the reply.
Text deltas are easy: print them and append them. Tool calls are the catch, because they also arrive in fragments that you have to reassemble yourself:
def merge_delta(msg: dict, delta: dict) -> None:
"""Fold one SSE delta into the assistant message being assembled."""
if delta.get("content"):
print(delta["content"], end="", flush=True)
msg["content"] += delta["content"]
for part in delta.get("tool_calls") or []:
slot = msg["calls"].setdefault(part.get("index", 0),
{"id": "", "type": "function",
"function": {"name": "", "arguments": ""}})
slot["id"] += part.get("id") or ""
fn = part.get("function") or {}
slot["function"]["name"] += fn.get("name") or ""
slot["function"]["arguments"] += fn.get("arguments") or ""
This function encodes four facts about the streaming format that no single document spells out:
- Tool calls are identified by
index, notid. Theidmay only appear in the first fragment, or arrive split across several. - The function
namecan be split across fragments. - The
argumentsstring definitely is. You receive{"pa, thenth": "fiz, thenzbuzz.py"}, and concatenate them. - One message can carry several calls, with their fragments interleaved. That’s why the accumulator is a dict keyed by index, and the final list is sorted:
calls = msg.pop("calls")
if calls:
msg["tool_calls"] = [calls[i] for i in sorted(calls)]
Appending with += to every field is the whole trick. slot["id"] += part.get("id") or "" looks redundant when the id arrives whole. It costs nothing, and it’s correct when the id arrives in pieces.
The read loop is nine lines. The only subtlety is that the final event, data: [DONE], isn’t JSON:
for raw in resp.iter_lines(): # server-sent events, one per line
line = raw.decode("utf-8", "replace").strip()
if not line.startswith("data:"):
continue
data = line[5:].strip()
if data == "[DONE]":
break
choices = json.loads(data).get("choices") or [{}]
merge_delta(msg, choices[0].get("delta") or {})
requests hands you a streamed response’s lines as they arrive through iter_lines(), so none of this needs async code. The harness also sends every request through one requests.Session, which keeps the connection open between calls. That saves a fresh TLS handshake (the encryption setup at the start of every HTTPS connection) on each round-trip, which adds up to a few seconds over a 40-step task.
Retries
The API fails in two different ways, and treating them alike is a mistake. Retry everything and a bad API key becomes four identical rejections and fifteen seconds of pointless waiting. So the error handling sorts failures first:
except requests.HTTPError:
detail = f"API {resp.status_code}: {resp.text[:200]}"
if resp.status_code == 429 or resp.status_code >= 500:
raise Retryable(detail, int(resp.headers.get("Retry-After") or 0) or None)
raise RuntimeError(detail)
except requests.RequestException as e: # DNS, TLS, timeout, connection reset
if msg and (msg["content"] or msg["calls"]):
raise RuntimeError(f"stream died mid-message: {e}")
raise Retryable(f"API unreachable: {e}")
Status 429 (“too many requests”) and the 5xx server errors mean “try again later”. A 400 (bad request) or 401 (bad key) means “this will never work”, and retrying just takes four times as long to be wrong.
Network errors are retryable only if nothing has streamed yet. This is the condition people miss. Once half of the model’s reply is on your screen, a silent retry would print the first half again, and you’d watch the agent stutter. It’s better to fail loudly than to misreport what the model said.
The backoff is what you’d expect: wait 1, 2, 4… seconds, but use the server’s Retry-After header when it sends one:
def chat(messages: list, tools: list, stream: bool = True) -> dict:
"""once(), with exponential backoff on the failures that deserve it."""
for attempt in range(MAX_RETRIES):
try:
return once(messages, tools, stream)
except Retryable as e:
delay = e.delay or min(2 ** attempt, 30)
print(f"[{e} | retry {attempt + 1}/{MAX_RETRIES} in {delay}s]")
time.sleep(delay)
raise RuntimeError(f"gave up after {MAX_RETRIES} attempts")
Printing each retry isn’t decoration. An agent that silently pauses for eight seconds looks broken; one that prints [API 429: {"error":"slow down"} | retry 1/4 in 1s] looks busy.
Sessions
The transcript is a list of dicts, so saving it is json.dumps. It’s embarrassing how much twenty lines buy:
def open_session(resume: bool):
"""Return (messages, path). --resume reopens the newest transcript."""
if resume:
path = newest_session()
if path:
messages = json.loads(path.read_text())
print(f"[resumed {path.name}: {len(messages)} messages]")
return messages, path
print("[no session to resume; starting a new one]")
SESSION_DIR.mkdir(parents=True, exist_ok=True)
return ([{"role": "system", "content": SYSTEM_PROMPT}],
SESSION_DIR / f"{time.strftime('%Y%m%d-%H%M%S')}.json")
The harness saves after every turn, and --resume picks up where you left off:
$ python harness.py --yes "create a file called zebra.txt containing the word stripes"
→ write_file({"path": "zebra.txt", "content": "stripes"})
I created a file called `zebra.txt` containing the word "stripes."
$ python harness.py --yes --resume "what file did you just create? answer from memory"
[resumed 20260920-111114.json: 5 messages]
I created a file called `zebra.txt` containing the word "stripes."
Two details are worth copying.
Sessions live in .agent/ inside your project, and .agent went into the list of directories the search tools skip in the same commit. Otherwise the agent’s grep starts matching its own memories, which is exactly as confusing as it sounds.
And the whole file is rewritten every turn instead of appended to. Compaction (next section) edits old messages in place, so an append-only log format like JSONL (one JSON object per line) would no longer match the real transcript.
Because the file is plain JSON, everything else is a one-liner. Inspect a session with jq, fork one with cp, or delete a bad turn in a text editor and resume. (If you delete by hand, keep each tool call with its result; the last section explains why.)
Compaction
A model can only read so much text per request; that limit is its context window. Every long session eventually exceeds it. The fix is to summarize the older half of the transcript and replace it with the summary.
That sounds like two lines of code, except the transcript isn’t a list of independent messages. A model reply containing tool calls, plus the tool messages answering it, form one unit. Cut between them and the API rejects the next request.
So the cut point isn’t “the middle”. It’s “the middle, then forward to a safe boundary”:
def cut_point(messages: list) -> int:
"""Halfway, then forward to a boundary that keeps tool results with their call."""
i = 1 + (len(messages) - 1) // 2
while i < len(messages) and messages[i].get("role") == "tool":
i += 1
return i
These few lines are the reason compaction works at all. The 1 + keeps the system prompt. The while loop steps past any tool results at the cut, so they stay with the reply that requested them.
def compact(messages: list) -> None:
"""Replace the older half of the transcript with a summary, in place."""
if len(json.dumps(messages)) < COMPACT_AT or len(messages) < 8:
return
cut = cut_point(messages)
old = messages[1:cut]
if not old:
return
print(f"[compacting {len(old)} messages]")
ask = [messages[0], *old, {"role": "user", "content":
"Summarize the work so far for your own future reference: files "
"read and changed, decisions made, what is still unfinished. Be "
"specific about names and line numbers. No preamble."}]
summary = chat(ask, [], stream=False).get("content") or "(summary failed)"
messages[1:cut] = [{"role": "user",
"content": f"Summary of earlier work:\n{summary}"}]
The size check counts characters, not tokens (the units models actually measure text in). Counting tokens exactly would need a tokenizer library, and len(json.dumps(...)) is within a factor of four, which is fine for a threshold you set conservatively.
The words “for your own future reference” matter. Ask for a summary and you get prose written for a human; ask for a handoff note and you get file names and line numbers. Here’s a real one, from a run with the threshold lowered to 1,500 characters:
Summary of earlier work:
- Read `f1.txt`: Lines 1-3 contain:
```
1: a
2: b
3: c
```
- Executed `echo one`.
- Executed `echo two`.
- Unfinished: Report the contents of `f1.txt` after the echo commands.
The agent finished the task after that, including re-reading the file it had been asked about.
Know the price, though. Providers cache the processed start of your transcript and charge less when the next request begins with the same bytes. Compaction rewrites that start, so the next call pays full price, and whatever was cut is gone for good. Treat compaction as what you do instead of failing, not something to do eagerly: if your model has a large context window, raise AGENT_COMPACT_AT to match.
Cancellation
In the minimal agent, Ctrl-C kills the program. That’s wrong twice over: you lose a session that was fine, and you usually only meant to stop this command, not the whole conversation.
Catching KeyboardInterrupt gets you halfway. The other half: the interrupt can land after the model asked for three tools but before you’ve answered all three. A transcript with an unanswered tool call is rejected by the API from then on. So cancelling has to repair what it interrupted:
def seal(messages: list) -> None:
"""Make an interrupted transcript valid: every tool_call needs a result."""
answered = {m.get("tool_call_id") for m in messages if m.get("role") == "tool"}
for m in list(messages):
for call in m.get("tool_calls") or []:
if call["id"] not in answered:
messages.append({"role": "tool", "tool_call_id": call["id"],
"content": "cancelled by user"})
messages.append({"role": "user", "content": "[interrupted by the user]"})
With that in place, a turn becomes something you can abandon:
def run_turn(messages: list) -> None:
"""One turn, interruptible. Ctrl-C ends the turn, not the session."""
try:
steps(messages)
except KeyboardInterrupt:
print("\n[cancelled]")
seal(messages)
The cancelled by user result also tells the model something: on the next turn it reads that and knows the command didn’t finish. Here it is interrupting a sleep 60:
> run the shell command: sleep 60
→ bash({"command": "sleep 60"})
^C
[cancelled]
> that got interrupted. in one sentence, what were you doing?
I was executing the command to pause the operation for 60 seconds.
The session survived, the API accepted the repaired transcript, and the model knew what had happened.
One rule, four features
Look at what streaming, retries, compaction, and cancellation have in common. Every hard part was the same problem in different clothes:
The transcript must be valid after every operation. Every tool call has exactly one
toolmessage answering it, and the two are never separated.
Streaming has to rebuild tool calls from fragments without losing one. Retries must not duplicate a half-sent message. Compaction may only cut at boundaries that keep calls with their results. Cancellation has to fill in the results it interrupted.
Break the rule and you don’t get a crash. You get an agent that works for six turns and then fails permanently with a 400 error, because the damage is in the history and every later request carries it.
That’s what a harness really is: code that keeps one data structure valid while a human, a model, and a filesystem all change it. Everything else is input and output.
Running and testing it
The harness reads the same settings as the minimal agent, plus one more, and adds a flag:
| Setting | Default | Effect |
|---|---|---|
AGENT_BASE_URL |
https://api.openai.com/v1 |
Any OpenAI-compatible API root (OpenRouter, vLLM, Ollama…) |
AGENT_API_KEY |
falls back to OPENAI_API_KEY |
API key; empty is fine for local servers |
AGENT_MODEL |
gpt-4o-mini |
Model name as the endpoint knows it |
AGENT_COMPACT_AT |
120000 |
Transcript size, in characters, that triggers summarization |
--resume / -r |
off | Reopen the newest transcript in .agent/ |
--yes / -y |
off | Skip the permission prompts; ask_user returns an error instead of waiting |
Several of these features are awkward to test against a real provider, because you can’t ask OpenAI for a 429 on demand. You don’t need to. The endpoint is just a URL, so flaky_server.py in the repo is a 32-line http.server that returns 429 with Retry-After: 1 on the first request and a canned SSE stream on the second. Point AGENT_BASE_URL at it and the real code path runs:
$ AGENT_BASE_URL=http://127.0.0.1:8931/v1 python harness.py --yes "hi"
[API 429: {"error":"slow down"} | retry 1/4 in 1s]
hello from the flaky server
A fake server like this lets you test the SSE parser against a stream you control byte for byte. The canned stream only carries text; add a tool call split across several events and you exercise merge_delta’s hard path, which is much easier than coaxing a real provider into splitting one at an awkward place. Point the harness at a closed port instead and you see the give-up path: four backoffs, then gave up after 4 attempts.
Compaction is testable the same way: set AGENT_COMPACT_AT=1500 and any real task triggers it within a few rounds. Cancellation needs a real terminal and a slow command: ask for sleep 60, press Ctrl-C, then ask the agent what it was doing. If it answers instead of failing with a 400, seal() did its job.
Still missing
What separates harness.py from something you’d ship to strangers:
- Concurrent tool calls. Models often ask for four tools at once; this runs them one after another.
- Permission scopes. The y/n prompt asks every time. Real harnesses remember “always allow
pytest” and match against patterns. - File-change detection. If you edit a file between the model reading it and writing it, the write silently overwrites your change. The fix is to store a hash of each file when it’s read and check it before writing.
- A real tokenizer. Counting characters is fine for a threshold and bad for a budget.
- Subagents. One conversation means one context window. Handing a search to a child agent with its own window is the biggest context saving available.
- A terminal UI. Rendered diffs, spinners, syntax highlighting, flicker-free redraws. This is where thousands of lines go, and it’s the part users judge you on.
Each is bounded and unglamorous. None needs a different architecture.
Takeaways
- The model is a commodity; the harness is not.
gpt-4o-miniproduced every transcript in this post. The difference between a frustrating agent and a good one was 145 lines of plumbing, not a better model. - Correctness lives in the transcript. Four of the five features here are one rule in different clothes. Write it down before you add anything that touches history.
- Users feel the unglamorous features. Nobody thanks you for a clever prompt. They notice that text streams, that Ctrl-C does the right thing, and that
--resumebrings back yesterday.
harness.py is 499 lines, including about 60 for two optional web tools. Read the diff against agent.py to see exactly what each fix cost.