Two independent multi-agent reviews of the branch (high, then max effort) found
authority-confinement and robustness defects the per-step reviews could not
see. This commit addresses every confirmed finding. task_agent turned out to be
the surface that lagged its siblings on nearly every axis.
Risk gate (most severe):
- task_agent(skill=...) never enforced the high/critical-risk PRINCIPAL-load-
only gate that skills(load) / spawn_workstream / spawn_batch enforce, so a
model could route around it by delegating activation to a sub-agent. Enforce
it inline in _prepare_task on the row already fetched (no re-query, no drift
between get_skill_by_name and get_prompt_template_by_name).
- _high_risk_skill_denied now fails CLOSED on a storage fault: deny, never wave
the skill through. Denying (not returning "") also keeps spawn_batch's per-row
partial-success intact under a transient blip.
- (first round) extracted _high_risk_skill_denied onto spawn_workstream /
spawn_batch, closing the coordinator-side bypass.
Persona confinement (Principle 7 attenuation on the task_agent edge):
- A restrictive persona now attenuates the sub-agent's TOOLS, not just its
identity text — the tool lever is frozen into the item and filtered before
_run_agent.
- Honor ALL FOUR persona levers on the sub-agent, not two: a child persona's
mcp-off and memory-off levers now drop MCP tools (mcp__* + read_resource /
use_prompt) and the memory tool, matching a main session under the persona.
- Cap the sub-agent by the PARENT session's own persona grant too, so a
restricted principal cannot escalate authority by spawning.
- Add persona to the task_agent judge/audit func_args projection (policy +
audit parity with spawn).
- Persona-resolution failures defer to a clean tool error (try/except mirroring
_validate_child_persona) instead of an opaque "internal error".
Substitution / capability:
- substitute_args=False for capability contexts (defaults, task_agent) so a
literal $ARGUMENTS / $N in a body is preserved, not blanked; env vars still
resolve. The literal-$ARGUMENTS scan is deferred behind that guard (skipped on
every capability render).
- Drop the CLAUDE_SKILL_DIR alias (canonical TURNSTONE_SKILL_DIR only). That
name also lives in bash, where turnstone-as-a-node-inside-Claude-Code must not
shadow the host's value; claiming it in the prompt but deferring in bash
diverged the two surfaces (a review finding). turnstone now claims it in
neither surface. The CLAUDE_SESSION_ID / CLAUDE_EFFORT prompt aliases stay
(pure prompt values, no bash-namespace collision).
Skills-as-context:
- DEFAULT (always-on) skills stay in the identity system message — the standing
baseline, never a mid-session cache-bust; only a NAMED applied skill moves to
the user-role capability message. This shrinks the pending model-adherence
eval surface to the named-skill move alone.
Cleanups: consolidate a duplicated rationale comment; correct the now-stale
"task agents are not persona-filtered" note.
PRE-MERGE GATE unchanged: the §7 Q1 model-adherence eval (named-skill move,
this branch vs main) is not runnable in-tree and must clear before merge.
Step 3 of the skill/persona split: an applied skill (including default skills)
is CAPABILITY context, so its body no longer sits in the identity system
message. It rides its own message (user role) after the identity block, with a
short intro naming the active skill. The <available-skills> discovery catalog
stays in the system message.
Two consequences:
- The cached identity prefix (persona BASE + ENV + POLICIES + catalogs) stays
stable across skills(load): loading/clearing a skill changes only the
trailing capability message, not the identity block.
- The task_agent base (_agent_system_messages) is snapshotted BEFORE the skill
block, so a parent's applied skill no longer leaks into the sub-agent prefix
(the sub-agent supplies its own persona identity and skill via _exec_task).
PRE-MERGE GATE: the design gates this on a model-adherence eval (this branch vs
main) verifying the model follows a skill as well from a context message as it
did from the system message (design section 7 Q1; ASSUMED-neutral, UNVERIFIED).
That eval is not runnable in-tree and MUST clear before this branch merges.
Mechanical structure is pinned by TestSkillContextPlacement.
Deferred follow-up: sub-agent (task_agent) skill-resource materialization, so
${TURNSTONE_SKILL_DIR} stays literal on that path (unchanged since step 1).
Test helpers (_sys_content) now read the full prompt prefix (identity + skill
context) so placement-agnostic assertions keep working.
Skill-body placeholder substitution diverged by invocation context:
interactive load, default skills, and spawn-child ran the full
render + spec-substitute, while task_agent (_exec_task) ran
_render_template only -- so $ARGUMENTS and ${...} env placeholders
rendered literally on that one path.
Introduce _render_skill_body as the single render+substitute path and
route interactive load, defaults, and task_agent through it, so a skill
reading ${TURNSTONE_EFFORT} or $ARGUMENTS resolves identically wherever
it runs. A sub-agent has no invocation args, so bare $ARGUMENTS and the
positional $N / $ARGUMENTS[N] forms resolve to empty there -- matching
the defaults and spawn-child paths, not the old verbatim passthrough.
- Add ${TURNSTONE_*} as the canonical vendor-neutral spelling for the
env placeholders (SESSION_ID, EFFORT, SKILL_DIR); keep ${CLAUDE_*} as
a permanent back-compat alias so imported skills keep resolving.
- Bash env: export TURNSTONE_SKILL_DIR and SKILL_RESOURCES_DIR
unconditionally, but add CLAUDE_SKILL_DIR only when the host has not
set it, so turnstone does not shadow a real value when it runs as a
node inside Claude Code.
- Materialize skill resources before substituting the body, so
${TURNSTONE_SKILL_DIR} resolves to the concrete bundle path on the
interactive path.
Sub-agent resource materialization and moving identity to a first-class
persona are left to follow-ups; ${TURNSTONE_SKILL_DIR} stays literal on
the task_agent path for now (unchanged from prior behavior).