Skills
14. Subagents
View this stage's code on GitHub →context: fork runs a skill in a second, fully independent call of the same agent loop — its own empty conversation, seeded with only the skill's body, with just its final answer crossing back into the main conversation. Pulling run_agent_loop() out into its own function (agent.py) is what makes a "subagent" possible at all: it isn't a different kind of agent, just this same loop called again on a new messages list. The one real trap: a subagent must not get the Skill tool itself, or a forked skill's own folder header ("Skill: owl (located at ...)") reads to the model as an instruction to call Skill again — recursing forever instead of just answering.
Your task
- Add an optional
context: forkfrontmatter field. - Pull the agent loop into its own function so it can be called again on a brand new, empty messages list.
- When the
Skilltool resolves acontext: forkskill, seed that new list with only the skill's body (never the question that triggered it), run the loop on it, and return its final answer as the tool's result. - Don't give the subagent the Skill tool itself, or its own folder header reads as an instruction to call Skill again.
The solution
Try the task above first. When you want to compare, this is exactly what changed since stage 13 (just diff 13 14 shows the same).
Show the solutionHide the solution +149 −70 lines
.petite/skills/owl/SKILL.md+9 −0
⋯ ---name: owldescription: Use this skill when the user asks about the on-call rotation.context: fork--- Respond with exactly this sentence and nothing else: The owl is on call this week. Do not call any tools.
agent.py+93 −0
⋯ """Stage 14: subagents A skill with `context: fork` in its frontmatter runs in a subagent — asecond, independent run of this same loop, seeded with only the skill'sbody (not the question that triggered it), returning just its finalanswer to the main conversation. Pulling the loop out into run_agent_loop() is what makes that possible:main() calls it once for the top-level conversation, and the Skill tool(in tools.py) calls it again, on a brand new messages list, whenever itresolves a context: fork skill. Both calls share the same model/tools —a subagent isn't a different kind of agent, just a fresh conversation.""" import jsonimport osimport sys from openai import OpenAI API_KEY = os.getenv("OPENROUTER_API_KEY")BASE_URL = os.getenv("OPENROUTER_BASE_URL", "https://openrouter.ai/api/v1")MODEL = os.getenv("MODEL", "anthropic/claude-haiku-4.5")MAX_TURNS = 20 _client = None def _get_client(): global _client if _client is None: if not API_KEY: raise RuntimeError("OPENROUTER_API_KEY is not set") _client = OpenAI(api_key=API_KEY, base_url=BASE_URL) return _client def run_agent_loop(messages, tools, execute_tool, label="agent"): """Run the agent loop to completion over `messages` (mutated in place: assistant/tool messages are appended as the loop runs) and return the model's final text answer. `tools` is the OpenAI tools spec list; `execute_tool(name, arguments)` runs one tool call and returns its result string. `label` only affects the stderr log prefix, so a subagent's logs are easy to tell apart from the main conversation's. """ client = _get_client() for turn in range(1, MAX_TURNS + 1): print( f"[{label}] turn {turn}: calling model with {len(messages)} message(s)", file=sys.stderr, ) response = client.chat.completions.create( model=MODEL, messages=messages, tools=tools, ) if not response.choices: raise RuntimeError("no choices in response") message = response.choices[0].message messages.append(message.model_dump()) tool_calls = message.tool_calls if not tool_calls: return message.content print( f"[{label}] turn {turn}: {len(tool_calls)} tool call(s) requested", file=sys.stderr, ) for call in tool_calls: arguments = json.loads(call.function.arguments) print(f"[{label}] executing {call.function.name}({arguments})", file=sys.stderr) result = execute_tool(call.function.name, arguments) messages.append( { "role": "tool", "tool_call_id": call.id, "content": result, } ) raise RuntimeError(f"[{label}] exceeded {MAX_TURNS} turns without a final answer")
main.py+8 −63
⋯ """Stage 11: stacking multiple skills in one promptStage 14: subagents "/fox /walk 4127" now expands into one user message per skill — seeskills.py for how the expansion and argument-sharing rules work.The loop that used to live here moved to agent.py's run_agent_loop(), sothe Skill tool can call it again for a context: fork skill's own,separate conversation. main() is now just: build the initial messages,run the loop once, print the answer.""" import argparseimport jsonimport osimport sys from openai import OpenAI from agent import run_agent_loopfrom skills import build_skills_system_prompt, discover_skills, resolve_slash_commandfrom tools import TOOLS, execute_tool API_KEY = os.getenv("OPENROUTER_API_KEY")BASE_URL = os.getenv("OPENROUTER_BASE_URL", "https://openrouter.ai/api/v1")MODEL = os.getenv("MODEL", "anthropic/claude-haiku-4.5")MAX_TURNS = 20 def main(): parser = argparse.ArgumentParser(description="petite: a tiny AI coding assistant")⋯ def main(): ) args = parser.parse_args() if not API_KEY: raise RuntimeError("OPENROUTER_API_KEY is not set") client = OpenAI(api_key=API_KEY, base_url=BASE_URL) skill_bodies = resolve_slash_command(args.prompt) if skill_bodies is not None: print(⋯ def main(): print(f"[agent] advertising {len(skills)} skill(s)", file=sys.stderr) messages.insert(0, {"role": "system", "content": system_prompt}) for turn in range(1, MAX_TURNS + 1): print( f"[agent] turn {turn}: calling model with {len(messages)} message(s)", file=sys.stderr, ) response = client.chat.completions.create( model=MODEL, messages=messages, tools=TOOLS, ) if not response.choices: raise RuntimeError("no choices in response") message = response.choices[0].message messages.append(message.model_dump()) tool_calls = message.tool_calls if not tool_calls: print(message.content) return print( f"[agent] turn {turn}: {len(tool_calls)} tool call(s) requested", file=sys.stderr, ) for call in tool_calls: arguments = json.loads(call.function.arguments) print( f"[agent] executing {call.function.name}({arguments})", file=sys.stderr ) result = execute_tool(call.function.name, arguments) messages.append( { "role": "tool", "tool_call_id": call.id, "content": result, } ) raise RuntimeError(f"exceeded {MAX_TURNS} turns without a final answer") answer = run_agent_loop(messages, TOOLS, execute_tool, label="agent") print(answer) if __name__ == "__main__":
skills.py+11 −0
⋯ class SkillMeta(BaseModel): default=False, alias="disable-model-invocation" ) user_invocable: bool = Field(default=True, alias="user-invocable") context: str | None = None @field_validator("name") @classmethod⋯ def resolve_slash_command(prompt, skills_dir=SKILLS_DIR): break expanded_names.append(candidate) i += 1 if meta.context == "fork": # A forked skill runs in its own subagent, isolated from the # main conversation — it can't be combined with another # skill's body the way two normal invocations can. break if not expanded_names: return None⋯ def resolve_skill_invocation(name, args_text="", skills_dir=SKILLS_DIR): return _with_folder_header(name, substitute_arguments(body, args), skills_dir) def get_skill_meta(name, skills_dir=SKILLS_DIR): """Public lookup of one skill's validated frontmatter, or None.""" return _load_skill_meta(name, skills_dir, warn=False) def substitute_arguments(body, args): """Replace $ARGUMENTS, $ARGUMENTS[n] and $n placeholders in a skill body with the given positional arguments.
tools.py+28 −7
⋯ import subprocess from pydantic import BaseModel, Field, ValidationError from skills import resolve_skill_invocationfrom agent import run_agent_loopfrom skills import get_skill_meta, resolve_skill_invocation class ReadArgs(BaseModel):⋯ class SkillTool(Tool): def execute(self, name, args=""): try: result = resolve_skill_invocation(name, args) body = resolve_skill_invocation(name, args) except ValueError as e: return str(e) if result is None: if body is None: return f"Unknown skill: {name}" return result meta = get_skill_meta(name) if meta is not None and meta.context == "fork": # Run in a subagent: a brand new conversation that only ever # sees this body, not the question that triggered it. Only # its final answer crosses back into the main conversation. subagent_messages = [{"role": "user", "content": body}] answer = run_agent_loop( subagent_messages, SUBAGENT_TOOLS, execute_tool, label=f"subagent:{name}" ) return f"Skill {name} ran in a separate context and returned: {answer}" ALL_TOOLS = [ReadTool(), WriteTool(), BashTool(), SkillTool()] return body TOOLS = [tool.spec() for tool in ALL_TOOLS] TOOLS_BY_NAME = {tool.name: tool for tool in ALL_TOOLS}ALL_TOOLS = [ReadTool(), WriteTool(), BashTool()]SKILL_TOOL = SkillTool() # Subagents get Read/Write/Bash but never the Skill tool itself — without# this, a forked skill's folder header ("Skill: owl (located at ...)")# reads to a model as an instruction to call Skill again, recursing# forever instead of just following the body it was already given.SUBAGENT_TOOLS = [tool.spec() for tool in ALL_TOOLS] MAIN_TOOLS = ALL_TOOLS + [SKILL_TOOL]TOOLS = [tool.spec() for tool in MAIN_TOOLS] TOOLS_BY_NAME = {tool.name: tool for tool in MAIN_TOOLS} def execute_tool(name, arguments):
Try it
Run this stage's own code:
just stage 14
cd .stages/14
# bundled skill "owl": "...on-call rotation...", context: fork, body: "...The owl is on call this week."
just run "Who is on the on-call rotation right now?"
Expect: The final answer says the owl is on call this week — verify with logs that the subagent's own conversation never contained your question, only the skill body
.stages/14 is a git worktree: your checkout stays on main, so
just keeps working there. Compare with the previous stage with just diff 13 14,
or browse it on GitHub;
just clean-stages removes the worktrees.
Where this fits
The whole agent; this stage builds the highlighted part. Click any part to jump to its stage.