Base agent

5. Bash tool

View this stage's code on GitHub →

Read and Write cover files; Bash covers everything else — ls, rm, running a test suite, whatever the model decides to try. It's the one tool that converts "the model composed a shell command" into "it actually happened," which is as much power as it is risk: subprocess.run with shell=True executes literally what the model wrote.

Your task

  1. Add a Bash tool: one required command string.
  2. Run it with your language's shell execution (subprocess.run(..., shell=True)), capture stdout+stderr together, and append a note with the exit code if it failed.

The solution

Try the task above first. When you want to compare, this is exactly what changed since stage 4 (just diff 4 5 shows the same).

Show the solutionHide the solution +32 −10 lines
main.py+32 −10
⋯ """Stage 4: the agent loopStage 5: the Bash tool Up to now every run was one round trip: ask, maybe run one tool, exit.That breaks for multi-step tasks ("read a file and fix any bugs") becausethe model never gets to see a tool's result and react to it. The fix: keep `messages` around across turns, keep calling the API, andonly stop when the model answers with plain text (no more tool_calls).Each tool call's result is appended as its own `role: "tool"` message,tagged with that call's `tool_call_id` so the model knows which resultanswers which request.Shell access: subprocess.run captures stdout and stderr, and the combinedoutput goes back to the model as the tool's result (empty on a silentsuccess, like `rm file` with no output).""" import argparseimport jsonimport osimport subprocessimport sys from openai import OpenAI⋯ TOOLS = [            },        },    },    {        "type": "function",        "function": {            "name": "Bash",            "description": "Execute a shell command",            "parameters": {                "type": "object",                "required": ["command"],                "properties": {                    "command": {                        "type": "string",                        "description": "The command to execute",                    }                },            },        },    },]  ⋯ def Write(file_path, content):    return f"Wrote to {file_path}"  def Bash(command):    completed = subprocess.run(command, shell=True, capture_output=True, text=True)    output = completed.stdout + completed.stderr    if completed.returncode != 0:        output += f"\n(exit code {completed.returncode})"    return output  def execute_tool(name, arguments):    if name == "Read":        return Read(arguments["file_path"])    if name == "Write":        return Write(arguments["file_path"], arguments["content"])    if name == "Bash":        return Bash(arguments["command"])    raise RuntimeError(f"unknown tool: {name}")  

Try it

Run this stage's own code:

just stage 5
cd .stages/05
just run "List the files in this directory with ls, then tell me how many there are."

Expect: A number matching the actual file count — check with ls yourself

.stages/05 is a git worktree: your checkout stays on main, so just keeps working there. Compare with the previous stage with just diff 4 5, or browse it on GitHub; just clean-stages removes the worktrees.

Where this fits

The whole agent; this stage builds the highlighted part. Click any part to jump to its stage.

References

← 4. Agent loop 6. Extract tools module →