Base agent
5. Bash tool
View this stage's code on GitHub →Read and Write cover files; Bash covers everything else — ls, rm, running a test suite, whatever the model decides to try. It's the one tool that converts "the model composed a shell command" into "it actually happened," which is as much power as it is risk: subprocess.run with shell=True executes literally what the model wrote.
Your task
- Add a
Bashtool: one requiredcommandstring. - Run it with your language's shell execution (
subprocess.run(..., shell=True)), capture stdout+stderr together, and append a note with the exit code if it failed.
The solution
Try the task above first. When you want to compare, this is exactly what changed since stage 4 (just diff 4 5 shows the same).
Show the solutionHide the solution +32 −10 lines
main.py+32 −10
⋯ """Stage 4: the agent loopStage 5: the Bash tool Up to now every run was one round trip: ask, maybe run one tool, exit.That breaks for multi-step tasks ("read a file and fix any bugs") becausethe model never gets to see a tool's result and react to it. The fix: keep `messages` around across turns, keep calling the API, andonly stop when the model answers with plain text (no more tool_calls).Each tool call's result is appended as its own `role: "tool"` message,tagged with that call's `tool_call_id` so the model knows which resultanswers which request.Shell access: subprocess.run captures stdout and stderr, and the combinedoutput goes back to the model as the tool's result (empty on a silentsuccess, like `rm file` with no output).""" import argparseimport jsonimport osimport subprocessimport sys from openai import OpenAI⋯ TOOLS = [ }, }, }, { "type": "function", "function": { "name": "Bash", "description": "Execute a shell command", "parameters": { "type": "object", "required": ["command"], "properties": { "command": { "type": "string", "description": "The command to execute", } }, }, }, },] ⋯ def Write(file_path, content): return f"Wrote to {file_path}" def Bash(command): completed = subprocess.run(command, shell=True, capture_output=True, text=True) output = completed.stdout + completed.stderr if completed.returncode != 0: output += f"\n(exit code {completed.returncode})" return output def execute_tool(name, arguments): if name == "Read": return Read(arguments["file_path"]) if name == "Write": return Write(arguments["file_path"], arguments["content"]) if name == "Bash": return Bash(arguments["command"]) raise RuntimeError(f"unknown tool: {name}")
Try it
Run this stage's own code:
just stage 5
cd .stages/05
just run "List the files in this directory with ls, then tell me how many there are."
Expect: A number matching the actual file count — check with ls yourself
.stages/05 is a git worktree: your checkout stays on main, so
just keeps working there. Compare with the previous stage with just diff 4 5,
or browse it on GitHub;
just clean-stages removes the worktrees.
Where this fits
The whole agent; this stage builds the highlighted part. Click any part to jump to its stage.