Compare commits

...

9 Commits

Author SHA1 Message Date
Patrick Buckley e86305c143 chore: bump version to 0.8.3 2026-03-17 17:06:00 -07:00
Patrick Buckley d0fc42195a chore: remove dead code and fix noisy JWT test warnings
Remove unused methods (ToolSearchManager.should_activate, get_all_tools),
dead attributes (_all_tools, _threshold), unused constant (DEFAULT_INTERVAL),
unused Scenario protocol class, and vestigial parameters (judge._evaluate_single
heuristic, SimEngine.simulate_llm_response turn_number). Lengthen JWT test
secrets to >= 32 bytes to suppress InsecureKeyLengthWarning from PyJWT.
2026-03-17 17:02:25 -07:00
Patrick Buckley 760321f7ee refactor: extract _resolve_capabilities and _without_tool helpers
Extract _resolve_capabilities() shared helper so _get_capabilities()
and _run_agent() use the same config-override logic instead of
duplicating inline. Add _without_tool() module-level helper to
deduplicate the tool-filtering listcomp.
2026-03-17 16:49:11 -07:00
Patrick Buckley 693e51f782 fix: address PR #119 review feedback
Add UI error notification and exc_info logging to run_one() exception
handler so tool failures are visible in the frontend. Apply config.toml
capability overrides when gating web_search in _run_agent(), matching
the pattern used by _get_capabilities().
2026-03-17 16:49:11 -07:00
Patrick Buckley ba07409724 fix: isolate parallel tool exceptions + gate web_search without backend
Two bugs: (1) an uncaught exception in one parallel tool call killed the
entire batch via pool.map(), losing all results including successful ones.
Wrap run_one() in try/except so failures return error strings instead of
propagating. (2) web_search was offered to local models even without a
Tavily API key — the model would attempt it, only to fail at execution
time. Filter web_search from _get_active_tools() and _run_agent() when
neither native support nor Tavily is available.

Closes https://github.com/turnstonelabs/turnstone/issues/117
2026-03-17 16:49:11 -07:00
Patrick Buckley c76a61841e fix: PR #118 round 2 — null-safe parser, docs, consistency
- Null-safe extraction for description, license, and compatibility in
  skill_parser.py — YAML bare keys (e.g. `description:`) no longer
  produce the literal string "None"
- Log warning on skill catalog storage failure instead of silent swallow
- Use `enabled == 1` in list_skills_by_activation for consistency with
  other prompt_templates queries in both storage backends
- Add parser tests for YAML null description, license, and compatibility
- Update governance.md: document runtime config editing on installed
  skills, two-column modal layout, SPDX license dropdown, origin badge
2026-03-17 16:09:06 -07:00
Patrick Buckley 341d2f604f fix: address PR #118 review feedback
- Regenerate OpenAPI snapshots (openapi-console.json) to include license
  and compatibility fields in SkillInfo/CreateSkillRequest/UpdateSkillRequest
- Omit version from create/update payloads when blank so server applies
  default "1.0.0" instead of storing empty string
- Push enabled_only + limit filters into list_skills_by_activation storage
  query (protocol, SQLite, PostgreSQL) instead of loading all rows and
  filtering in Python; session.py now passes enabled_only=True, limit=30
- License length cap ([:128]) was already applied in previous commit
2026-03-17 16:09:06 -07:00
Patrick Buckley 3f7f8495d6 feat: skills modal redesign + runtime config editing for installed skills
Redesigns the create/edit/view skill modal into a two-column spec manifest
layout (Identity/Manifest/Deployment | Skill Content) matching the Agent
Skills spec structure. Installed (readonly) skills can now have their runtime
config (model, temperature, token limits, enabled) edited independently of
the locked spec/content fields.

- Two-column spec layout with section headings (Identity, Manifest, Deployment,
  Skill Content); content textarea uses monospace font and fills the column
- h3 section headings for screen-reader nav; h3 UA stylesheet reset in CSS
- Runtime Config collapsible uses 3-column grid; license field is now a select
  of SPDX identifiers (MIT, Apache-2.0, GPL-3.0, AGPL-3.0, etc.)
- Origin badge (cyan) shows source URL for installed skills in view mode
- server.py: _SKILL_RUNTIME_CONFIG_FIELDS frozenset; readonly skills filter
  updates to config-only fields (spec fields silently dropped); audit action
  distinguishes skill.update.config from skill.update; license field capped
  at 128 chars in both create and update paths
- governance.js: spec fields disabled for readonly; config fields always
  editable; Save button shown for all skills (labeled "Save Config" when
  readonly); collapsible state reset between modal opens prevents state leak;
  esk-allowed-tools disabled state driven by auto_approve not readonly
- Tests: spec-only body on readonly skill → 400; config-only → 200 with
  spec fields unchanged; mixed body → config fields applied, spec dropped
2026-03-17 16:09:06 -07:00
Patrick Buckley dc464ac313 feat: Agent Skills standard compliance + frontend spec fields
Brings skills implementation into full compliance with agentskills.io:

Parser:
- Read `allowed-tools` (hyphenated, standard) only; stored as
  `allowed_tools` internally — no underscore fallback
- Reject consecutive hyphens in skill names
- Extract author/version from standard `metadata:` map with top-level
  fallback; null-safe (no "None" string for bare YAML keys)
- Truncate description at 1024 chars, compatibility at 500 chars (spec
  caps) with log warnings
- Lenient parsing mode (lenient=True) for cross-client import: sanitizes
  names, returns None on skip, malformed-YAML colon-value retry
- Type overloads: strict mode returns ParsedSkill, lenient returns
  ParsedSkill | None

Session:
- `<available-skills>` XML catalog in system messages for
  activation="search" skills (disabled ones filtered out, capped at 30)

Tool rename:
- `load_skill` tool → `skill` (JSON, session preparers/executors,
  approval labels, tests, docs)

Storage (migration 023):
- Add `license` and `compatibility` columns to prompt_templates
- skill_license / compatibility params on create_prompt_template across
  protocol, SQLite, PostgreSQL backends
- Add to SKILL_MUTABLE for update_prompt_template

API + server:
- SkillInfo, CreateSkillRequest, UpdateSkillRequest: license +
  compatibility fields
- Create/update/install endpoints extract and persist both fields
- Install endpoint maps parsed.license + parsed.compatibility from
  imported SKILL.md (previously discarded)
- _skill_to_response() includes both fields

Admin UI:
- Create + edit modals: version, license, compatibility fields
- Readonly (imported) skills: "edit" → "view" button, modal title
  "View Skill", all fields disabled, Save hidden, Cancel → "Close",
  collapsibles auto-expand, focus on Close button
- :disabled CSS for dark-theme modal inputs (bg-highlight, cursor
  not-allowed, dimmed text)
- Fix addEventListener stacking on auto-approve checkboxes → .onchange

SDK: license + compatibility on SkillInfo, CreateSkillRequest,
UpdateSkillRequest TypeScript interfaces

Docs: governance.md, judge.md, tools.md, README, diagram updated
2026-03-17 16:09:06 -07:00
43 changed files with 1658 additions and 354 deletions
+1 -1
View File
@@ -145,7 +145,7 @@ Turnstone includes a built-in governance layer for enterprise deployments — ma
- **RBAC** — 15 granular permissions, 3 built-in roles (admin / operator / viewer), custom roles, privilege escalation prevention
- **OIDC SSO** — single sign-on via any OpenID Connect provider (Okta, Azure AD, Google, Keycloak); Authorization Code Flow with PKCE, auto-provisioning, claim-based role mapping with demotion propagation; see [docs/oidc.md](docs/oidc.md)
- **Tool policies** — glob-pattern rules (`allow` / `deny` / `ask`) with priority ordering; automate approvals or lock down dangerous tools
- **Skills** — reusable behavioral profiles with system prompts, `{{variable}}` substitution, session config (model, temperature, token budget), install-time security scanning, version history, external discovery (skills.sh / GitHub), and runtime `load_skill` tool for model-driven skill activation
- **Skills** — reusable behavioral profiles with system prompts, `{{variable}}` substitution, session config (model, temperature, token budget), install-time security scanning, version history, external discovery (skills.sh / GitHub), and runtime `skill` tool for model-driven skill activation
- **Usage tracking** — per-request token and tool metrics, aggregation by day / model / user, automatic 90-day pruning
- **Audit logging** — append-only event trail for all admin mutations, IP-aware, 365-day retention
@@ -250,13 +250,11 @@ class "MCPClientManager" as MCPMgr {
' ToolSearchManager
class "ToolSearchManager" as ToolSearchMgr {
- _all_tools: list[dict]
- _always_on: list[dict]
- _deferred: list[dict]
- _expanded: dict[str, None]
- _index: BM25Index
--
+ should_activate() → bool
+ get_visible_tools() → list[dict]
+ get_deferred_tools() → list[dict]
+ get_expanded_names() → list[str]
@@ -33,7 +33,7 @@ package "Core Modules" as core #181825 {
}
package "Session Runtime" as runtime #181825 {
rectangle "load_skill tool\nsession.py" as loadtool
rectangle "skill tool\nsession.py" as loadtool
rectangle "set_skill()\nsession.py" as setskill
rectangle "_load_skills()\nsession.py" as loadskills
}
@@ -80,6 +80,7 @@ importui --> install : POST (github source)
' Annotations
note right of parser
YAML frontmatter -> ParsedSkill
allowed-tools (standard) -> allowed_tools (internal)
Anthropic + Hermes tag formats
Name validation (lowercase+hyphens)
end note
+20 -2
View File
@@ -70,7 +70,7 @@ etc.) since workstream templates were merged into the skills system in v0.8.0.
`{{node_id}}` (server node ID). Unrecognized placeholders are kept as-is.
- **Runtime switching**: `/template <name>` to switch, `/template clear` to revert
to defaults, `/template` to show current. Persisted across resume.
- **Model-driven loading**: The `load_skill` built-in tool lets the model
- **Model-driven loading**: The `skill` built-in tool lets the model
discover and activate skills mid-conversation. `search` action finds skills
by query (auto-approved); `load` action activates by name (requires user
approval since it changes session behavior). Main session only.
@@ -83,11 +83,15 @@ etc.) since workstream templates were merged into the skills system in v0.8.0.
precedence on name collision. MCP-synced content updates reset `is_default` to
prevent compromised servers from injecting defaults. Admin UI shows origin badge
and disables edit/delete for MCP-sourced skills.
- **Spec fields**: Skills support the full Agent Skills standard frontmatter:
`name`, `description`, `license`, `compatibility`, `metadata` (author, version),
`allowed-tools`. The `license` and `compatibility` fields are preserved on import
and editable in the admin UI. See https://agentskills.io/specification.
- **Security scanning**: Skills are automatically scanned at creation and update
time. The scanner evaluates four risk axes: content risk (command execution,
data exfiltration), supply chain risk (pipe-to-shell, transitive installs),
vulnerability risk (prompt injection, insecure credentials), and declared
capability risk (from `allowed_tools`). Results populate the `scan_status`
capability risk (from `allowed-tools` in SKILL.md). Results populate the `scan_status`
(safe/low/medium/high/critical) and `scan_report` (JSON breakdown) columns.
These fields are system-managed and cannot be overwritten via the admin API.
- **Discovery**: External skills can be discovered and installed from registries:
@@ -100,6 +104,20 @@ etc.) since workstream templates were merged into the skills system in v0.8.0.
Discovery view has search bar, result cards, and "Import from GitHub" modal.
- SDK: `discover_skills(q)` and `install_skill(source, skill_id=..., url=...)`
on both Python and TypeScript console clients.
- **Runtime config on installed skills**: Installed (readonly) skills can have
their runtime configuration edited — model, temperature, reasoning effort,
token budget, max tokens, agent max turns, auto-approve, allowed tools,
and enabled flag. The server restricts updates to these fields only via
`_SKILL_RUNTIME_CONFIG_FIELDS` filtering; spec/content fields (name,
description, tags, license, compatibility, content, activation) remain
immutable. The admin UI shows "Save Config" instead of "Save" for these
skills. Audit action: `skill.update.config`.
- **Admin UI**: Create/Edit skill modals use a two-column spec manifest layout
(left: Identity / Manifest / Deployment; right: Skill Content editor with
monospace font). Runtime Config is a collapsible 3-column grid below.
License uses an SPDX identifier dropdown (MIT, Apache-2.0, GPL-3.0, etc.).
Installed skills show a cyan origin badge with source URL, spec fields are
disabled, and all collapsible sections auto-expand in view mode.
### Usage Tracking
+1 -1
View File
@@ -299,7 +299,7 @@ four independent risk axes:
obfuscation, download-execute chains, executable URLs from untrusted domains
3. **Vulnerability risk** — prompt injection patterns, insecure credential
handling, third-party content exposure (indirect prompt injection surface)
4. **Declared capability risk** — parsed from the skill's `allowed_tools` field.
4. **Declared capability risk** — parsed from `allowed-tools` in the skill's SKILL.md.
`Bash(*)` (unrestricted shell) is high risk. `Bash(git:*)` is low.
Read-only tools are safe.
+5 -5
View File
@@ -492,7 +492,7 @@ data.get("mergedAt") is not None
---
### load_skill
### skill
Discover and activate skills at runtime during a conversation. The model can
search for available skills and load one by name, replacing the current active
@@ -543,7 +543,7 @@ pre-configure skills at workstream creation.
| `watch` | Monitor | No (create) | No | No | `command` |
| `read_resource`| MCP | No | Yes | Yes | `uri` |
| `use_prompt` | MCP | No | Yes | Yes | `name` |
| `load_skill` | Skills | No (load) | No | No | `name` |
| `skill` | Skills | No (load) | No | No | `name` |
| `tool_search`| Search | Yes | No | No | `query` |
---
@@ -591,9 +591,9 @@ CLI flags override the config file:
### How it works
1. **Threshold check**: At session startup, `ToolSearchManager.should_activate()`
counts total tools (built-in + MCP). If the count is below the threshold, tool
search stays off and all tools are sent to the model directly.
1. **Threshold check**: At session startup, if the total tool count (built-in + MCP)
is below the threshold, tool search stays off and all tools are sent to the model
directly.
2. **Partitioning**: When active, tools are split into two sets:
- **Always-on** -- the 17 built-in tools (members of `BUILTIN_TOOL_NAMES`).
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "turnstone"
version = "0.8.2"
version = "0.8.3"
description = "Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."
readme = "README.md"
license = "BUSL-1.1"
+45 -1
View File
@@ -2,7 +2,7 @@
"openapi": "3.1.0",
"info": {
"title": "turnstone Console API",
"version": "0.7.0",
"version": "0.8.2",
"description": "Cluster-wide visibility and control across all turnstone nodes."
},
"paths": {
@@ -6883,6 +6883,16 @@
"title": "Allowed Tools",
"type": "string"
},
"license": {
"default": "",
"title": "License",
"type": "string"
},
"compatibility": {
"default": "",
"title": "Compatibility",
"type": "string"
},
"scan_status": {
"default": "",
"title": "Scan Status",
@@ -7107,6 +7117,16 @@
"default": "[]",
"title": "Allowed Tools",
"type": "string"
},
"license": {
"default": "",
"title": "License",
"type": "string"
},
"compatibility": {
"default": "",
"title": "Compatibility",
"type": "string"
}
},
"required": [
@@ -7357,6 +7377,30 @@
],
"default": null,
"title": "Allowed Tools"
},
"license": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "License"
},
"compatibility": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Compatibility"
}
},
"title": "UpdateSkillRequest",
+1 -1
View File
@@ -2,7 +2,7 @@
"openapi": "3.1.0",
"info": {
"title": "turnstone Server API",
"version": "0.7.0",
"version": "0.8.2",
"description": "Single-node workstream management, chat interaction, and real-time streaming."
},
"paths": {
+6
View File
@@ -187,6 +187,8 @@ export interface SkillInfo {
notify_on_complete: string;
enabled: boolean;
allowed_tools: string;
license: string;
compatibility: string;
resource_count: number;
created: string;
updated: string;
@@ -214,6 +216,8 @@ export interface CreateSkillRequest {
notify_on_complete?: string;
enabled?: boolean;
allowed_tools?: string;
license?: string;
compatibility?: string;
}
export interface UpdateSkillRequest {
@@ -237,6 +241,8 @@ export interface UpdateSkillRequest {
notify_on_complete?: string;
enabled?: boolean;
allowed_tools?: string;
license?: string;
compatibility?: string;
}
export interface ListSkillsResponse {
+5 -5
View File
@@ -130,7 +130,7 @@ class TestParseScopes:
class TestJWT:
SECRET = "test-secret-key-for-jwt"
SECRET = "test-secret-key-for-jwt-min-32b!"
def test_round_trip(self):
scopes = frozenset({"read", "write"})
@@ -155,7 +155,7 @@ class TestJWT:
def test_invalid_signature(self):
token = create_jwt("user1", frozenset({"read"}), "db", self.SECRET)
assert validate_jwt(token, "wrong-secret") is None
assert validate_jwt(token, "wrong-secret-key-for-jwt-min-32b") is None
def test_malformed_token(self):
assert validate_jwt("not.a.jwt", self.SECRET) is None
@@ -217,7 +217,7 @@ class TestAuthenticateToken:
assert result.scopes == frozenset({"read", "write", "approve"})
def test_jwt_token(self):
secret = "test-secret"
secret = "test-secret-key-for-jwt-min-32b!"
jwt_tok = create_jwt("user1", frozenset({"read", "write"}), "db", secret)
cfg = AuthConfig(enabled=True)
result = _authenticate_token(jwt_tok, cfg, jwt_secret=secret)
@@ -304,7 +304,7 @@ class TestCheckRequestScopes:
assert result.has_scope("approve")
def test_jwt_with_scopes(self):
secret = "test"
secret = "test-secret-key-for-jwt-min-32b!"
jwt_tok = create_jwt("u1", frozenset({"read", "write"}), "db", secret)
cfg = AuthConfig(enabled=True)
allowed, status, msg, result = check_request(
@@ -319,7 +319,7 @@ class TestCheckRequestScopes:
assert result.user_id == "u1"
def test_jwt_insufficient_scope(self):
secret = "test"
secret = "test-secret-key-for-jwt-min-32b!"
jwt_tok = create_jwt("u1", frozenset({"read"}), "db", secret)
cfg = AuthConfig(enabled=True)
allowed, status, msg, _ = check_request(
-4
View File
@@ -182,7 +182,6 @@ class TestErrorHandling:
result = judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
MagicMock(),
)
assert result is None
@@ -215,7 +214,6 @@ class TestErrorHandling:
result = judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
MagicMock(),
)
assert result is None
@@ -259,7 +257,6 @@ class TestMultiTurnToolUse:
verdict = judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
MagicMock(),
)
assert verdict is not None
assert verdict.tier == "llm"
@@ -305,7 +302,6 @@ class TestMultiTurnToolUse:
judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
MagicMock(),
)
# Should have called create_completion exactly _JUDGE_MAX_TURNS times
assert provider.create_completion.call_count == 5
+143 -46
View File
@@ -1,4 +1,4 @@
"""Tests for the load_skill built-in tool."""
"""Tests for the skill built-in tool."""
from __future__ import annotations
@@ -9,25 +9,25 @@ from turnstone.core.tools import BUILTIN_TOOL_NAMES, PRIMARY_KEY_MAP
class TestToolRegistration:
"""Verify load_skill is registered correctly."""
"""Verify skill is registered correctly."""
def test_in_builtin_tool_names(self) -> None:
assert "load_skill" in BUILTIN_TOOL_NAMES
assert "skill" in BUILTIN_TOOL_NAMES
def test_not_agent_tool(self) -> None:
from turnstone.core.tools import AGENT_TOOLS
names = {t["function"]["name"] for t in AGENT_TOOLS}
assert "load_skill" not in names
assert "skill" not in names
def test_not_task_agent_tool(self) -> None:
from turnstone.core.tools import TASK_AGENT_TOOLS
names = {t["function"]["name"] for t in TASK_AGENT_TOOLS}
assert "load_skill" not in names
assert "skill" not in names
def test_has_primary_key(self) -> None:
assert PRIMARY_KEY_MAP.get("load_skill") == "name"
assert PRIMARY_KEY_MAP.get("skill") == "name"
# ---------------------------------------------------------------------------
@@ -82,12 +82,12 @@ def _make_session(skills: list[dict[str, Any]] | None = None):
class TestPrepareLoadSkill:
"""Test _prepare_load_skill validation and item dict shape."""
"""Test _prepare_skill validation and item dict shape."""
def test_load_valid(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "load", "name": "code-review"})
assert item["func_name"] == "load_skill"
item = session._prepare_skill("call-1", {"action": "load", "name": "code-review"})
assert item["func_name"] == "skill"
assert item["action"] == "load"
assert item["name"] == "code-review"
assert item["needs_approval"] is True
@@ -96,19 +96,19 @@ class TestPrepareLoadSkill:
def test_load_missing_name(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "load"})
item = session._prepare_skill("call-1", {"action": "load"})
assert "error" in item
assert "name" in item["error"].lower()
assert item["needs_approval"] is False
def test_load_empty_name(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "load", "name": ""})
item = session._prepare_skill("call-1", {"action": "load", "name": ""})
assert "error" in item
def test_search_with_query(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "code review"})
item = session._prepare_skill("call-1", {"action": "search", "query": "code review"})
assert item["action"] == "search"
assert item["query"] == "code review"
assert item["needs_approval"] is False
@@ -116,30 +116,30 @@ class TestPrepareLoadSkill:
def test_search_without_query(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search"})
item = session._prepare_skill("call-1", {"action": "search"})
assert item["action"] == "search"
assert item["query"] == ""
assert item["needs_approval"] is False
def test_invalid_action(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "delete"})
item = session._prepare_skill("call-1", {"action": "delete"})
assert "error" in item
assert "delete" in item["error"]
def test_empty_action(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": ""})
item = session._prepare_skill("call-1", {"action": ""})
assert "error" in item
def test_header_for_load(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "load", "name": "my-skill"})
item = session._prepare_skill("call-1", {"action": "load", "name": "my-skill"})
assert "my-skill" in item["header"]
def test_header_for_search(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "testing"})
item = session._prepare_skill("call-1", {"action": "search", "query": "testing"})
assert "testing" in item["header"]
@@ -149,7 +149,7 @@ class TestPrepareLoadSkill:
class TestExecLoadSkill:
"""Test _exec_load_skill execution logic."""
"""Test _exec_skill execution logic."""
def test_load_existing_skill(self) -> None:
skills = [
@@ -164,8 +164,8 @@ class TestExecLoadSkill:
session, _, fake_get = _make_session(skills)
with patch("turnstone.core.session.get_skill_by_name", side_effect=fake_get):
item = session._prepare_load_skill("call-1", {"action": "load", "name": "code-review"})
call_id, result = session._exec_load_skill(item)
item = session._prepare_skill("call-1", {"action": "load", "name": "code-review"})
call_id, result = session._exec_skill(item)
assert call_id == "call-1"
assert "code-review" in result
@@ -177,8 +177,8 @@ class TestExecLoadSkill:
session, _, fake_get = _make_session([])
with patch("turnstone.core.session.get_skill_by_name", side_effect=fake_get):
item = session._prepare_load_skill("call-1", {"action": "load", "name": "nope"})
call_id, result = session._exec_load_skill(item)
item = session._prepare_skill("call-1", {"action": "load", "name": "nope"})
call_id, result = session._exec_skill(item)
assert "not found" in result.lower()
assert session._set_skill_called == []
@@ -188,8 +188,8 @@ class TestExecLoadSkill:
session, _, fake_get = _make_session(skills)
with patch("turnstone.core.session.get_skill_by_name", side_effect=fake_get):
item = session._prepare_load_skill("call-1", {"action": "load", "name": "test"})
session._exec_load_skill(item)
item = session._prepare_skill("call-1", {"action": "load", "name": "test"})
session._exec_skill(item)
session.ui.on_tool_result.assert_called_once()
@@ -216,10 +216,10 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = skills
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "code"})
item = session._prepare_skill("call-1", {"action": "search", "query": "code"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "code-review" in result
# docs-writer shouldn't match "code" query
@@ -241,10 +241,10 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = skills
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search"})
item = session._prepare_skill("call-1", {"action": "search"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
# Should be limited to 10
assert result.count("skill-") == 10
@@ -254,10 +254,10 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = []
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "nonexistent"})
item = session._prepare_skill("call-1", {"action": "search", "query": "nonexistent"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "no skills found" in result.lower()
@@ -276,21 +276,21 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = skills
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "risky"})
item = session._prepare_skill("call-1", {"action": "search", "query": "risky"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "high" in result
def test_search_storage_failure_returns_empty(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "test"})
item = session._prepare_skill("call-1", {"action": "search", "query": "test"})
with patch(
"turnstone.core.storage._registry.get_storage", side_effect=RuntimeError("no storage")
):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "no skills found" in result.lower()
@@ -307,10 +307,8 @@ class TestExecLoadSkill:
session, _, fake_get = _make_session(skills)
with patch("turnstone.core.session.get_skill_by_name", side_effect=fake_get):
item = session._prepare_load_skill(
"call-1", {"action": "load", "name": "disabled-skill"}
)
call_id, result = session._exec_load_skill(item)
item = session._prepare_skill("call-1", {"action": "load", "name": "disabled-skill"})
call_id, result = session._exec_skill(item)
assert "not found" in result.lower()
assert session._set_skill_called == []
@@ -321,8 +319,8 @@ class TestExecLoadSkill:
session._skill_name = "active"
with patch("turnstone.core.session.get_skill_by_name", side_effect=fake_get):
item = session._prepare_load_skill("call-1", {"action": "load", "name": "active"})
call_id, result = session._exec_load_skill(item)
item = session._prepare_skill("call-1", {"action": "load", "name": "active"})
call_id, result = session._exec_skill(item)
assert "already active" in result.lower()
assert session._set_skill_called == []
@@ -352,10 +350,10 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = skills
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search"})
item = session._prepare_skill("call-1", {"action": "search"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "enabled-skill" in result
assert "disabled-skill" not in result
@@ -375,14 +373,113 @@ class TestExecLoadSkill:
mock_storage.list_prompt_templates.return_value = skills
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "search", "query": "code review"})
item = session._prepare_skill("call-1", {"action": "search", "query": "code review"})
with patch("turnstone.core.storage._registry.get_storage", return_value=mock_storage):
call_id, result = session._exec_load_skill(item)
call_id, result = session._exec_skill(item)
assert "code-review" in result
def test_preparer_load_has_approval_label(self) -> None:
session, _, _ = _make_session()
item = session._prepare_load_skill("call-1", {"action": "load", "name": "my-skill"})
assert item["approval_label"] == "load_skill__my-skill"
item = session._prepare_skill("call-1", {"action": "load", "name": "my-skill"})
assert item["approval_label"] == "skill__my-skill"
# ---------------------------------------------------------------------------
# Tests: Skill Catalog Disclosure (Agent Skills standard compliance)
# ---------------------------------------------------------------------------
class TestSkillCatalogDisclosure:
"""Verify <available-skills> catalog appears in system messages."""
def _build_session_with_system_messages(
self,
search_skills: list[dict[str, Any]] | None = None,
) -> Any:
"""Build a session and call _init_system_messages to get dev_parts."""
from turnstone.core.session import ChatSession
session = ChatSession.__new__(ChatSession)
ui = MagicMock()
session.ui = ui
session.model = "test-model"
session._ws_id = "ws-test"
session._node_id = "node-1"
session._skill_name = None
session._skill_content = None
session._skill_resources = {}
session._applied_skill_content = None
session.context_window = 128000
session.messages = []
session._config = {}
session.creative_mode = False
session.instructions = ""
session.system_messages = []
session._agent_system_messages = []
session.reasoning_effort = "medium"
session._pending_nudge = []
session._tool_search = None
session._mcp_client = None
session._notify_on_complete = "{}"
# Memory stubs
session._memory_config = MagicMock()
session._memory_config.fetch_limit = 0
session._user_id = ""
with (
patch(
"turnstone.core.session.list_skills_by_activation",
return_value=search_skills or [],
),
patch.object(session, "_get_visible_memories", return_value=[]),
):
session._init_system_messages()
return session
def test_catalog_present_with_search_skills(self) -> None:
skills = [
{"name": "pdf-processing", "description": "Extract PDF text and forms."},
{"name": "data-analysis", "description": "Analyze datasets."},
]
session = self._build_session_with_system_messages(search_skills=skills)
content = session.system_messages[0]["content"]
assert "<available-skills>" in content
assert "pdf-processing" in content
assert "data-analysis" in content
assert "</available-skills>" in content
def test_catalog_omitted_when_no_search_skills(self) -> None:
session = self._build_session_with_system_messages(search_skills=[])
content = session.system_messages[0]["content"]
assert "<available-skills>" not in content
def test_catalog_capped_at_30(self) -> None:
skills = [{"name": f"skill-{i:03d}", "description": f"Desc {i}"} for i in range(50)]
session = self._build_session_with_system_messages(search_skills=skills)
content = session.system_messages[0]["content"]
# Should include first 30, not all 50
assert "skill-029" in content
assert "skill-030" not in content
def test_catalog_escapes_html(self) -> None:
skills = [
{"name": "xss-test", "description": "Handle <script> & 'quotes'."},
]
session = self._build_session_with_system_messages(search_skills=skills)
content = session.system_messages[0]["content"]
assert "&lt;script&gt;" in content
assert "<script>" not in content.replace("<available-skills>", "").replace(
"</available-skills>", ""
).replace("<skill>", "").replace("</skill>", "").replace("<name>", "").replace(
"</name>", ""
).replace("<description>", "").replace("</description>", "")
def test_catalog_includes_hint(self) -> None:
skills = [{"name": "test", "description": "Test skill."}]
session = self._build_session_with_system_messages(search_skills=skills)
content = session.system_messages[0]["content"]
assert "/skill" in content
+3 -3
View File
@@ -131,7 +131,7 @@ def authorize_client(storage: SQLiteBackend, oidc_config: OIDCConfig) -> TestCli
)
app.state.oidc_config = oidc_config
app.state.auth_storage = storage
app.state.jwt_secret = "test-jwt-secret"
app.state.jwt_secret = "test-jwt-secret-key-padded-32b!!"
app.state.jwks_data = {"keys": []}
app.state.login_limiter = None
return TestClient(app, raise_server_exceptions=False)
@@ -468,7 +468,7 @@ class TestOIDCCallback:
)
app.state.oidc_config = _make_oidc_config()
app.state.auth_storage = backend
app.state.jwt_secret = "secret"
app.state.jwt_secret = "test-jwt-secret-key-padded-32b!!"
app.state.jwks_data = {"keys": []}
app.state.login_limiter = None
@@ -492,7 +492,7 @@ class TestOIDCCallback:
)
app.state.oidc_config = _make_oidc_config()
app.state.auth_storage = storage
app.state.jwt_secret = "secret"
app.state.jwt_secret = "test-jwt-secret-key-padded-32b!!"
app.state.jwks_data = {"keys": []}
limiter = LoginRateLimiter(max_attempts=1, window_seconds=300)
limiter.record("ip:testclient")
+147
View File
@@ -607,3 +607,150 @@ class TestPruneWorkstreams:
# Config rows should be cleaned up
assert load_workstream_config("orphan_cfg") == {}
assert load_workstream_config("stale_cfg") == {}
# ── Parallel tool exception isolation ────────────────────────────────
class TestParallelToolExceptionIsolation:
"""Bug #117: one tool raising should not kill the entire batch."""
def test_exception_in_one_tool_does_not_kill_batch(self, tmp_db, mock_openai_client):
from unittest.mock import patch
session = ChatSession(
client=mock_openai_client,
model="test-model",
ui=MagicMock(),
instructions=None,
temperature=0.5,
max_tokens=1000,
tool_timeout=10,
)
def succeed(item):
return item["call_id"], "ok"
def fail(item):
raise RuntimeError("boom")
items = [
{
"call_id": "c1",
"func_name": "bash",
"execute": succeed,
"needs_approval": False,
"header": "test",
"preview": "",
},
{
"call_id": "c2",
"func_name": "math",
"execute": fail,
"needs_approval": False,
"header": "test",
"preview": "",
},
]
tool_calls = [
{"id": "c1", "function": {"name": "bash", "arguments": "{}"}},
{"id": "c2", "function": {"name": "math", "arguments": "{}"}},
]
with (
patch.object(session, "_prepare_tool", side_effect=items),
patch.object(session, "_evaluate_intent"),
patch.object(session, "_emit_state"),
patch.object(session, "_init_system_messages"),
patch.object(session, "_check_cancelled"),
):
session.ui.approve_tools.return_value = (True, None)
results, _ = session._execute_tools(tool_calls)
assert results[0] == ("c1", "ok")
assert results[1][0] == "c2"
assert "Error executing math" in results[1][1]
assert "boom" in results[1][1]
# ── Web search tool gating ───────────────────────────────────────────
class TestWebSearchGating:
"""Bug #117: web_search should not be offered without a backend."""
def test_web_search_filtered_when_no_backend(self, tmp_db, mock_openai_client):
from unittest.mock import patch
from turnstone.core.providers._protocol import ModelCapabilities
session = ChatSession(
client=mock_openai_client,
model="local-model",
ui=MagicMock(),
instructions=None,
temperature=0.5,
max_tokens=1000,
tool_timeout=10,
)
caps = ModelCapabilities(supports_web_search=False)
with (
patch.object(session, "_get_capabilities", return_value=caps),
patch("turnstone.core.session.get_tavily_key", return_value=None),
):
tools = session._get_active_tools()
names = [t.get("function", {}).get("name") for t in tools]
assert "web_search" not in names
def test_web_search_kept_when_tavily_available(self, tmp_db, mock_openai_client):
from unittest.mock import patch
from turnstone.core.providers._protocol import ModelCapabilities
session = ChatSession(
client=mock_openai_client,
model="local-model",
ui=MagicMock(),
instructions=None,
temperature=0.5,
max_tokens=1000,
tool_timeout=10,
)
caps = ModelCapabilities(supports_web_search=False)
with (
patch.object(session, "_get_capabilities", return_value=caps),
patch("turnstone.core.session.get_tavily_key", return_value="tvly-test-key"),
):
tools = session._get_active_tools()
names = [t.get("function", {}).get("name") for t in tools]
assert "web_search" in names
def test_web_search_kept_when_native_support(self, tmp_db, mock_openai_client):
from unittest.mock import patch
from turnstone.core.providers._protocol import ModelCapabilities
session = ChatSession(
client=mock_openai_client,
model="gpt-5-search-api",
ui=MagicMock(),
instructions=None,
temperature=0.5,
max_tokens=1000,
tool_timeout=10,
)
caps = ModelCapabilities(supports_web_search=True)
with (
patch.object(session, "_get_capabilities", return_value=caps),
patch("turnstone.core.session.get_tavily_key", return_value=None),
):
tools = session._get_active_tools()
names = [t.get("function", {}).get("name") for t in tools]
assert "web_search" in names
+3 -3
View File
@@ -74,7 +74,7 @@ class TestSimEngine:
def test_llm_response_returns_content(self, engine):
async def _test():
content, tool_calls = await engine.simulate_llm_response(True, 1)
content, tool_calls = await engine.simulate_llm_response(True)
assert isinstance(content, str)
assert len(content) > 0
assert isinstance(tool_calls, list)
@@ -85,8 +85,8 @@ class TestSimEngine:
async def _test():
e1 = SimEngine(fast_config, rng=random.Random(123))
e2 = SimEngine(fast_config, rng=random.Random(123))
c1, t1 = await e1.simulate_llm_response(True, 1)
c2, t2 = await e2.simulate_llm_response(True, 1)
c1, t1 = await e1.simulate_llm_response(True)
c2, t2 = await e2.simulate_llm_response(True)
assert c1 == c2
assert len(t1) == len(t2)
+316 -6
View File
@@ -18,7 +18,7 @@ description: Automated code review skill
author: Test Author
version: 2.0.0
tags: [python, review, quality]
allowed_tools: [read_file, list_directory]
allowed-tools: [read_file, list_directory]
license: MIT
compatibility: ">=0.7"
---
@@ -187,13 +187,13 @@ Content.
class TestAllowedTools:
"""Verify allowed_tools parsing."""
"""Verify allowed-tools parsing (Agent Skills standard hyphenated field)."""
def test_list_format(self) -> None:
raw = """\
---
name: tools-list
allowed_tools: [bash, read_file]
allowed-tools: [bash, read_file]
---
Content.
@@ -201,11 +201,12 @@ Content.
result = parse_skill_md(raw)
assert result.allowed_tools == ["bash", "read_file"]
def test_csv_format(self) -> None:
def test_space_delimited_format(self) -> None:
"""Standard format per Agent Skills spec."""
raw = """\
---
name: tools-csv
allowed_tools: "bash, read_file, write_file"
name: tools-space
allowed-tools: "bash read_file write_file"
---
Content.
@@ -219,6 +220,19 @@ Content.
name: no-tools
---
Content.
"""
result = parse_skill_md(raw)
assert result.allowed_tools == []
def test_underscore_key_not_read(self) -> None:
"""allowed_tools (underscore) is not a SKILL.md field — ignored by parser."""
raw = """\
---
name: legacy-key
allowed_tools: [bash, read_file]
---
Content.
"""
result = parse_skill_md(raw)
@@ -247,3 +261,299 @@ class TestValidateSkillName:
assert validate_skill_name("HAS-UPPER") is not None
assert validate_skill_name("has space") is not None
assert validate_skill_name("-leading-hyphen") is not None
def test_consecutive_hyphens_rejected(self) -> None:
"""Agent Skills spec: consecutive hyphens not allowed."""
assert validate_skill_name("foo--bar") is not None
assert "consecutive hyphens" in (validate_skill_name("a--b") or "")
# Single hyphens are fine
assert validate_skill_name("foo-bar") is None
# -- Agent Skills Standard Compliance Tests -----------------------------------
class TestStandardAllowedTools:
"""Agent Skills spec: 'allowed-tools' (hyphenated), space-delimited."""
def test_list_format(self) -> None:
raw = """\
---
name: standard-tools
allowed-tools: ["Bash(git:*)", "Bash(jq:*)", "Read"]
---
Content.
"""
result = parse_skill_md(raw)
assert result.allowed_tools == ["Bash(git:*)", "Bash(jq:*)", "Read"]
def test_space_delimited(self) -> None:
"""Standard format: space-delimited string."""
raw = """\
---
name: space-tools
allowed-tools: "Bash(git:*) Bash(jq:*) Read"
---
Content.
"""
result = parse_skill_md(raw)
assert result.allowed_tools == ["Bash(git:*)", "Bash(jq:*)", "Read"]
def test_mixed_space_comma_delimiters(self) -> None:
raw = """\
---
name: mixed-delim
allowed-tools: "Read, Write Bash"
---
Content.
"""
result = parse_skill_md(raw)
assert result.allowed_tools == ["Read", "Write", "Bash"]
class TestStandardMetadataNesting:
"""Standard puts author/version under metadata map."""
def test_metadata_author(self) -> None:
raw = """\
---
name: nested-author
description: Test skill
metadata:
author: example-org
version: "2.0"
---
Content.
"""
result = parse_skill_md(raw)
assert result.author == "example-org"
assert result.version == "2.0"
def test_top_level_takes_precedence(self) -> None:
raw = """\
---
name: precedence
description: Test skill
author: top-level
version: 1.0.0
metadata:
author: nested
version: "2.0"
---
Content.
"""
result = parse_skill_md(raw)
assert result.author == "top-level"
assert result.version == "1.0.0"
def test_metadata_version_only(self) -> None:
raw = """\
---
name: version-only
description: Test
metadata:
version: "3.5.1"
---
Content.
"""
result = parse_skill_md(raw)
assert result.version == "3.5.1"
assert result.author == ""
def test_null_author_uses_default(self) -> None:
"""YAML null/bare key must not produce the string 'None'."""
raw = """\
---
name: null-author
description: Test
author:
---
Content.
"""
result = parse_skill_md(raw)
assert result.author == ""
assert result.version == "1.0.0"
def test_null_version_uses_default(self) -> None:
raw = """\
---
name: null-version
description: Test
version:
---
Content.
"""
result = parse_skill_md(raw)
assert result.version == "1.0.0"
def test_null_description_falls_back_to_body(self) -> None:
"""YAML null description must not produce 'None' string."""
raw = """\
---
name: null-desc
description:
---
First paragraph here.
"""
result = parse_skill_md(raw)
assert result.description == "First paragraph here."
assert "None" not in result.description
def test_null_license_and_compatibility(self) -> None:
"""YAML null license/compatibility must not produce 'None' string."""
raw = """\
---
name: null-fields
description: Test
license:
compatibility:
---
Content.
"""
result = parse_skill_md(raw)
assert result.license == ""
assert result.compatibility == ""
class TestStandardFieldLengths:
"""Spec caps: description <= 1024, compatibility <= 500."""
def test_description_truncated_at_1024(self) -> None:
long_desc = "x" * 1200
raw = f"""\
---
name: long-desc
description: "{long_desc}"
---
Content.
"""
result = parse_skill_md(raw)
assert len(result.description) == 1024
def test_compatibility_truncated_at_500(self) -> None:
long_compat = "y" * 600
raw = f"""\
---
name: long-compat
description: Short
compatibility: "{long_compat}"
---
Content.
"""
result = parse_skill_md(raw)
assert len(result.compatibility) == 500
def test_short_fields_unchanged(self) -> None:
raw = """\
---
name: short
description: Brief
compatibility: Requires git
---
Content.
"""
result = parse_skill_md(raw)
assert result.description == "Brief"
assert result.compatibility == "Requires git"
class TestLenientMode:
"""Lenient parsing for cross-client skill ingestion."""
def test_invalid_name_sanitized(self) -> None:
raw = """\
---
name: Invalid_Name!
description: A test skill
---
Content.
"""
result = parse_skill_md(raw, lenient=True)
assert result is not None
assert result.name == "invalidname"
def test_unsalvageable_name_returns_none(self) -> None:
raw = """\
---
name: "!!!"
description: A test skill
---
Content.
"""
assert parse_skill_md(raw, lenient=True) is None
def test_missing_description_returns_none(self) -> None:
raw = """\
---
name: no-desc
---
"""
assert parse_skill_md(raw, lenient=True) is None
def test_broken_yaml_returns_none(self) -> None:
raw = """\
---
name: [broken: yaml: {{{
---
Content.
"""
assert parse_skill_md(raw, lenient=True) is None
def test_malformed_yaml_colon_in_description_recovers(self) -> None:
"""Standard recommends retrying unquoted colon values."""
raw = """\
---
name: colon-desc
description: Use this skill when: the user asks about PDFs
---
Content.
"""
result = parse_skill_md(raw, lenient=True)
# The frontmatter library may parse this fine, but if not,
# the retry mechanism should recover.
assert result is not None
assert result.name == "colon-desc"
assert "PDF" in result.description
def test_strict_mode_still_raises(self) -> None:
"""Default strict mode unchanged."""
raw = """\
---
name: Invalid_Name!
description: A test skill
---
Content.
"""
with pytest.raises(ValueError):
parse_skill_md(raw)
def test_consecutive_hyphens_lenient(self) -> None:
raw = """\
---
name: foo--bar
description: A test skill
---
Content.
"""
result = parse_skill_md(raw, lenient=True)
assert result is not None
assert "--" not in result.name
+103 -4
View File
@@ -802,8 +802,8 @@ class TestSkillAPI:
)
assert resp.status_code == 404
def test_update_skill_readonly_rejected(self, api_client, api_storage):
"""Updating a readonly (MCP-sourced) skill returns 403."""
def test_update_skill_readonly_spec_fields_rejected(self, api_client, api_storage):
"""Updating spec fields on a readonly skill returns 400 (filtered to nothing)."""
_create_template(
api_storage,
"s1",
@@ -815,9 +815,65 @@ class TestSkillAPI:
)
resp = api_client.put(
"/v1/api/admin/skills/s1",
json={"description": "hacked"},
json={"description": "hacked", "content": "evil"},
)
assert resp.status_code == 403
assert resp.status_code == 400
assert "runtime config" in resp.json()["error"].lower()
def test_update_skill_readonly_runtime_config_allowed(self, api_client, api_storage):
"""Runtime config fields can be updated on a readonly (installed) skill."""
_create_template(
api_storage,
"s1",
"installed-skill",
"external content",
origin="source",
readonly=True,
)
resp = api_client.put(
"/v1/api/admin/skills/s1",
json={"model": "gpt-5", "temperature": 0.5, "enabled": False},
)
assert resp.status_code == 200
data = resp.json()
assert data["model"] == "gpt-5"
assert data["temperature"] == 0.5
assert data["enabled"] is False
# Spec fields must remain unchanged
assert data["content"] == "external content"
def test_update_skill_readonly_mixed_body_filters_spec(self, api_client, api_storage):
"""When JS sends all fields for a readonly skill, spec fields are silently dropped."""
_create_template(
api_storage,
"s1",
"installed",
"original content",
origin="source",
readonly=True,
)
resp = api_client.put(
"/v1/api/admin/skills/s1",
# Simulate what the browser form submits: every field present
json={
"name": "hacked",
"content": "evil content",
"description": "tampered",
"model": "gpt-5",
"enabled": False,
"token_budget": 50000,
},
)
assert resp.status_code == 200
data = resp.json()
# Config fields updated
assert data["model"] == "gpt-5"
assert data["enabled"] is False
assert data["token_budget"] == 50000
# Spec fields unchanged
assert data["name"] == "installed"
assert data["content"] == "original content"
assert data["description"] == ""
def test_update_skill_recomputes_token_estimate(self, api_client, api_storage):
"""Updating content recomputes token_estimate."""
@@ -1156,6 +1212,49 @@ class TestSkillSessionConfigApplication:
parsed_tools = json.loads(tpl["allowed_tools"])
assert parsed_tools == ["bash", "read_file", "write_file"]
def test_license_compatibility_roundtrip(self, db):
"""Agent Skills spec fields license and compatibility round-trip."""
db.create_prompt_template(
template_id="spec1",
name="spec-fields-skill",
category="general",
content="Spec test.",
skill_license="Apache-2.0",
compatibility="Requires git, docker, jq",
)
tpl = db.get_skill_by_name("spec-fields-skill")
assert tpl is not None
assert tpl["license"] == "Apache-2.0"
assert tpl["compatibility"] == "Requires git, docker, jq"
def test_license_compatibility_default_empty(self, db):
"""license and compatibility default to empty string."""
db.create_prompt_template(
template_id="spec2",
name="no-spec-fields",
category="general",
content="No spec fields.",
)
tpl = db.get_skill_by_name("no-spec-fields")
assert tpl is not None
assert tpl["license"] == ""
assert tpl["compatibility"] == ""
def test_update_license_compatibility(self, db):
"""license and compatibility can be updated."""
db.create_prompt_template(
template_id="spec3",
name="updatable-spec",
category="general",
content="Test.",
)
db.update_prompt_template("spec3", license="MIT")
db.update_prompt_template("spec3", compatibility="Python 3.11+")
tpl = db.get_prompt_template("spec3")
assert tpl is not None
assert tpl["license"] == "MIT"
assert tpl["compatibility"] == "Python 3.11+"
# ---------------------------------------------------------------------------
# 7. Migration behavior tests
-11
View File
@@ -128,17 +128,9 @@ class TestToolSearchManager:
return ToolSearchManager(
all_tools,
always_on_names={"bash", "read_file", "edit_file"},
threshold=5,
max_results=3,
)
def test_should_activate_above_threshold(self, manager):
assert manager.should_activate()
def test_should_not_activate_below_threshold(self, builtin_tools):
mgr = ToolSearchManager(builtin_tools, always_on_names={"bash", "read_file", "edit_file"})
assert not mgr.should_activate()
def test_visible_tools_initially_builtin_only(self, manager):
visible = manager.get_visible_tools()
names = {_tool_name(t) for t in visible}
@@ -201,9 +193,6 @@ class TestToolSearchManager:
names = {_tool_name(t) for t in deferred}
assert "mcp__github__create_issue" not in names
def test_get_all_tools_returns_everything(self, manager, builtin_tools, mcp_tools):
assert len(manager.get_all_tools()) == len(builtin_tools) + len(mcp_tools)
def test_search_tool_definition_format(self, manager):
defn = manager.get_search_tool_definition()
assert defn["type"] == "function"
+1 -1
View File
@@ -112,7 +112,7 @@ class TestToolsMetadata:
"watch": "command",
"read_resource": "uri",
"use_prompt": "name",
"load_skill": "name",
"skill": "name",
}
assert expected == PRIMARY_KEY_MAP
+1 -1
View File
@@ -1,3 +1,3 @@
"""turnstone - Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."""
__version__ = "0.8.2"
__version__ = "0.8.3"
+6
View File
@@ -307,6 +307,8 @@ class SkillInfo(BaseModel):
notify_on_complete: str = "{}"
enabled: bool = True
allowed_tools: str = "[]"
license: str = ""
compatibility: str = ""
scan_status: str = ""
scan_report: str = "{}"
scan_version: str = ""
@@ -337,6 +339,8 @@ class CreateSkillRequest(BaseModel):
notify_on_complete: str = "{}"
enabled: bool = True
allowed_tools: str = "[]"
license: str = ""
compatibility: str = ""
class UpdateSkillRequest(BaseModel):
@@ -360,6 +364,8 @@ class UpdateSkillRequest(BaseModel):
notify_on_complete: str | None = None
enabled: bool | None = None
allowed_tools: str | None = None
license: str | None = None
compatibility: str | None = None
class ListSkillsResponse(BaseModel):
+39 -3
View File
@@ -2134,6 +2134,24 @@ async def admin_delete_policy(request: Request) -> JSONResponse:
_VALID_ACTIVATIONS = {"named", "default", "search"}
# Fields that may be updated on installed (readonly) skills.
# These are local runtime configuration — not part of the SKILL.md spec —
# so they don't compromise the fidelity of an externally-sourced skill.
_SKILL_RUNTIME_CONFIG_FIELDS = frozenset(
{
"model",
"temperature",
"reasoning_effort",
"max_tokens",
"token_budget",
"agent_max_turns",
"auto_approve",
"allowed_tools",
"enabled",
"notify_on_complete",
}
)
def _parse_skill_session_config(body: dict[str, Any]) -> tuple[dict[str, Any], JSONResponse | None]:
"""Parse and validate session config fields from a skill request body.
@@ -2286,6 +2304,8 @@ def _skill_to_response(r: dict[str, Any], resource_count: int = 0) -> dict[str,
"notify_on_complete": r.get("notify_on_complete", "{}"),
"enabled": r.get("enabled", True),
"allowed_tools": r.get("allowed_tools", "[]"),
"license": r.get("license", ""),
"compatibility": r.get("compatibility", ""),
"scan_status": r.get("scan_status", ""),
"scan_report": r.get("scan_report", "{}"),
"scan_version": r.get("scan_version", ""),
@@ -2370,6 +2390,8 @@ async def admin_create_skill(request: Request) -> JSONResponse:
org_id = str(body.get("org_id", "")).strip()[:64]
author = str(body.get("author", "")).strip()[:256]
version = str(body.get("version", "1.0.0")).strip()[:64]
license_val = str(body.get("license", "")).strip()[:128]
compatibility = str(body.get("compatibility", "")).strip()[:500]
raw_tags = body.get("tags", [])
if isinstance(raw_tags, list):
@@ -2418,6 +2440,8 @@ async def admin_create_skill(request: Request) -> JSONResponse:
tags=tags_str,
version=version,
author=author,
skill_license=license_val,
compatibility=compatibility,
activation=activation,
token_estimate=token_estimate,
**session_fields,
@@ -2456,8 +2480,7 @@ async def admin_update_skill(request: Request) -> JSONResponse:
existing = storage.get_prompt_template(skill_id)
if existing is None:
return JSONResponse({"error": "Skill not found"}, status_code=404)
if existing.get("readonly"):
return JSONResponse({"error": "MCP-sourced skills are read-only"}, status_code=403)
is_readonly = bool(existing.get("readonly"))
body = await read_json_or_400(request)
if isinstance(body, JSONResponse):
@@ -2497,6 +2520,10 @@ async def admin_update_skill(request: Request) -> JSONResponse:
updates["author"] = str(body["author"]).strip()[:256]
if "version" in body:
updates["version"] = str(body["version"]).strip()[:64]
if "license" in body:
updates["license"] = str(body["license"]).strip()[:128]
if "compatibility" in body:
updates["compatibility"] = str(body["compatibility"]).strip()[:500]
if "tags" in body:
raw_tags = body["tags"]
if isinstance(raw_tags, list):
@@ -2509,6 +2536,13 @@ async def admin_update_skill(request: Request) -> JSONResponse:
tag_str = "[]"
updates["tags"] = tag_str
# Installed (readonly) skills: restrict updates to runtime config only.
# Spec/content fields are locked to preserve external-source fidelity.
if is_readonly:
updates = {k: v for k, v in updates.items() if k in _SKILL_RUNTIME_CONFIG_FIELDS}
if not updates:
return JSONResponse({"error": "No runtime config fields to update"}, status_code=400)
# Snapshot current state for version history before applying update
existing_versions = storage.list_skill_versions(skill_id)
version_int = len(existing_versions) + 1
@@ -2526,7 +2560,7 @@ async def admin_update_skill(request: Request) -> JSONResponse:
record_audit(
storage,
audit_uid,
"skill.update",
"skill.update.config" if is_readonly else "skill.update",
"skill",
skill_id,
updates,
@@ -3203,6 +3237,8 @@ async def admin_skill_install(request: Request) -> JSONResponse:
source_url=pkg_source_url,
version=parsed.version,
author=parsed.author,
skill_license=parsed.license,
compatibility=parsed.compatibility,
activation="named",
token_estimate=token_estimate,
allowed_tools=allowed_tools_str,
+154 -59
View File
@@ -768,7 +768,7 @@ function _renderGovSkills(items) {
t.resource_count +
" res</span>";
}
var editDisabled = t.readonly ? " disabled" : "";
var editLabel = t.readonly ? "view" : "edit";
var deleteDisabled = "";
html +=
'<div class="admin-row" role="listitem">' +
@@ -794,9 +794,9 @@ function _renderGovSkills(items) {
'<span class="admin-col admin-col-actions">' +
'<button class="admin-btn-action" data-edit-tmpl="' +
escapeHtml(t.template_id) +
'"' +
editDisabled +
">edit</button>" +
'">' +
editLabel +
"</button>" +
'<button class="admin-btn-danger" data-delete-tmpl="' +
escapeHtml(t.template_id) +
'" data-tmpl-name="' +
@@ -870,6 +870,9 @@ function showCreateTemplateModal() {
document.getElementById("skill-description").value = "";
document.getElementById("skill-tags").value = "";
document.getElementById("skill-author").value = "";
document.getElementById("skill-version").value = "";
document.getElementById("skill-license").value = "";
document.getElementById("skill-compatibility").value = "";
document.getElementById("skill-activation").value = "named";
document.getElementById("ctm-content").value = "";
document.getElementById("ctm-variables").textContent = "(none)";
@@ -888,11 +891,9 @@ function showCreateTemplateModal() {
document.getElementById("csk-allowed-tools").value = "";
document.getElementById("csk-allowed-tools").disabled = false;
document.getElementById("csk-enabled").checked = true;
document
.getElementById("csk-auto-approve")
.addEventListener("change", function () {
document.getElementById("csk-allowed-tools").disabled = this.checked;
});
document.getElementById("csk-auto-approve").onchange = function () {
document.getElementById("csk-allowed-tools").disabled = this.checked;
};
document.getElementById("create-template-error").style.display = "none";
// Clear resource list
_pendingResources = [];
@@ -949,31 +950,38 @@ function submitCreateTemplate() {
.filter(Boolean)
: [];
document.getElementById("ctm-submit").disabled = true;
var csVersion = (document.getElementById("skill-version").value || "").trim();
var createBody = {
name: name,
category: document.getElementById("ctm-category").value,
description: (
document.getElementById("skill-description").value || ""
).trim(),
tags: JSON.stringify(tagsArray),
author: (document.getElementById("skill-author").value || "").trim(),
license: (document.getElementById("skill-license").value || "").trim(),
compatibility: (
document.getElementById("skill-compatibility").value || ""
).trim(),
activation: document.getElementById("skill-activation").value,
content: content,
variables: JSON.stringify(varList),
is_default: document.getElementById("ctm-default").checked,
model: document.getElementById("csk-model").value.trim(),
auto_approve: document.getElementById("csk-auto-approve").checked,
temperature: csTemp ? parseFloat(csTemp) : null,
reasoning_effort: document.getElementById("csk-reasoning-effort").value,
max_tokens: csMaxTok ? parseInt(csMaxTok, 10) : null,
token_budget: csBudget ? parseInt(csBudget, 10) : 0,
agent_max_turns: csMaxTurns ? parseInt(csMaxTurns, 10) : null,
allowed_tools: JSON.stringify(csAllowedArr),
enabled: document.getElementById("csk-enabled").checked,
};
if (csVersion) createBody.version = csVersion;
authFetch("/v1/api/admin/skills", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
name: name,
category: document.getElementById("ctm-category").value,
description: (
document.getElementById("skill-description").value || ""
).trim(),
tags: JSON.stringify(tagsArray),
author: (document.getElementById("skill-author").value || "").trim(),
activation: document.getElementById("skill-activation").value,
content: content,
variables: JSON.stringify(varList),
is_default: document.getElementById("ctm-default").checked,
model: document.getElementById("csk-model").value.trim(),
auto_approve: document.getElementById("csk-auto-approve").checked,
temperature: csTemp ? parseFloat(csTemp) : null,
reasoning_effort: document.getElementById("csk-reasoning-effort").value,
max_tokens: csMaxTok ? parseInt(csMaxTok, 10) : null,
token_budget: csBudget ? parseInt(csBudget, 10) : 0,
agent_max_turns: csMaxTurns ? parseInt(csMaxTurns, 10) : null,
allowed_tools: JSON.stringify(csAllowedArr),
enabled: document.getElementById("csk-enabled").checked,
}),
body: JSON.stringify(createBody),
})
.then(function (r) {
if (!r.ok)
@@ -1052,6 +1060,9 @@ function showEditTemplateModal(tmplId) {
}
document.getElementById("etm-tags").value = tagsDisplay;
document.getElementById("etm-author").value = tmpl.author || "";
document.getElementById("etm-version").value = tmpl.version || "";
document.getElementById("etm-license").value = tmpl.license || "";
document.getElementById("etm-compatibility").value = tmpl.compatibility || "";
document.getElementById("etm-activation").value = tmpl.activation || "named";
document.getElementById("etm-content").value = tmpl.content;
_updateVarsDisplay("etm-content", "etm-variables");
@@ -1086,11 +1097,9 @@ function showEditTemplateModal(tmplId) {
document.getElementById("esk-allowed-tools").disabled =
tmpl.auto_approve || false;
document.getElementById("esk-enabled").checked = tmpl.enabled !== false;
document
.getElementById("esk-auto-approve")
.addEventListener("change", function () {
document.getElementById("esk-allowed-tools").disabled = this.checked;
});
document.getElementById("esk-auto-approve").onchange = function () {
document.getElementById("esk-allowed-tools").disabled = this.checked;
};
document.getElementById("edit-template-error").style.display = "none";
// Scan report section
var scanSection = document.getElementById("etm-scan-section");
@@ -1180,12 +1189,91 @@ function showEditTemplateModal(tmplId) {
});
};
}
// Reset collapsible state before applying readonly rules (prevents state leak
// when switching between readonly and editable skills in the same session)
var allDetails = document.querySelectorAll(
"#edit-template-box .admin-details",
);
for (var d = 0; d < allDetails.length; d++) allDetails[d].open = false;
// --- Readonly mode for imported skills ---
var isReadonly = tmpl.readonly || false;
var editTitle = document.getElementById("edit-template-title");
if (editTitle)
editTitle.textContent = isReadonly ? "View Skill" : "Edit Skill";
// Origin badge — show provenance for installed skills
var originBadge = document.getElementById("etm-origin-badge");
if (originBadge) {
if (isReadonly && tmpl.source_url) {
originBadge.textContent = "Installed from \u00a0" + tmpl.source_url;
originBadge.style.display = "inline-flex";
} else if (isReadonly && tmpl.origin && tmpl.origin !== "manual") {
originBadge.textContent = "Installed skill";
originBadge.style.display = "inline-flex";
} else {
originBadge.style.display = "none";
}
}
var submitBtn = document.getElementById("etm-submit");
if (submitBtn) {
submitBtn.style.display = "";
submitBtn.textContent = isReadonly ? "Save Config" : "Save";
}
// Spec/content fields: locked for installed skills (preserve source fidelity)
[
"etm-name",
"etm-category",
"etm-description",
"etm-tags",
"etm-author",
"etm-version",
"etm-license",
"etm-compatibility",
"etm-activation",
"etm-content",
"etm-default",
].forEach(function (id) {
var el = document.getElementById(id);
if (el) el.disabled = isReadonly;
});
// Runtime config fields: always editable (local settings, not part of SKILL.md spec)
[
"esk-model",
"esk-temperature",
"esk-reasoning-effort",
"esk-max-tokens",
"esk-token-budget",
"esk-agent-max-turns",
"esk-auto-approve",
"esk-enabled",
].forEach(function (id) {
var el = document.getElementById(id);
if (el) el.disabled = false;
});
// esk-allowed-tools follows auto_approve state, not readonly state
var allowedToolsEl = document.getElementById("esk-allowed-tools");
if (allowedToolsEl) allowedToolsEl.disabled = tmpl.auto_approve || false;
var cancelBtn = document.querySelector("#edit-template-box .modal-cancel");
if (cancelBtn) cancelBtn.textContent = isReadonly ? "Close" : "Cancel";
// Auto-expand Runtime Config collapsible for installed skills so config is visible
if (isReadonly) {
var details = document.querySelectorAll(
"#edit-template-box .admin-details",
);
for (var d = 0; d < details.length; d++) details[d].open = true;
}
// --- Skill Resources ---
var resSection = document.getElementById("etm-resources-section");
if (resSection) {
_loadSkillResources(tmplId, tmpl.readonly || false);
_loadSkillResources(tmplId, isReadonly);
}
_etmTrapHandler = _installTrap("edit-template-overlay", "edit-template-box");
// Focus management
if (isReadonly) {
if (cancelBtn) cancelBtn.focus();
} else {
document.getElementById("etm-name").focus();
}
}
function hideEditTemplateModal() {
@@ -1466,31 +1554,38 @@ function submitEditTemplate() {
.filter(Boolean)
: [];
document.getElementById("etm-submit").disabled = true;
var esVersion = (document.getElementById("etm-version").value || "").trim();
var updateBody = {
name: document.getElementById("etm-name").value.trim(),
category: document.getElementById("etm-category").value,
description: (
document.getElementById("etm-description").value || ""
).trim(),
tags: JSON.stringify(tagsArray),
author: (document.getElementById("etm-author").value || "").trim(),
license: (document.getElementById("etm-license").value || "").trim(),
compatibility: (
document.getElementById("etm-compatibility").value || ""
).trim(),
activation: document.getElementById("etm-activation").value,
content: content,
variables: JSON.stringify(varList),
is_default: document.getElementById("etm-default").checked,
model: document.getElementById("esk-model").value.trim(),
auto_approve: document.getElementById("esk-auto-approve").checked,
temperature: esTemp ? parseFloat(esTemp) : null,
reasoning_effort: document.getElementById("esk-reasoning-effort").value,
max_tokens: esMaxTok ? parseInt(esMaxTok, 10) : null,
token_budget: esBudget ? parseInt(esBudget, 10) : 0,
agent_max_turns: esMaxTurns ? parseInt(esMaxTurns, 10) : null,
allowed_tools: JSON.stringify(esAllowedArr),
enabled: document.getElementById("esk-enabled").checked,
};
if (esVersion) updateBody.version = esVersion;
authFetch("/v1/api/admin/skills/" + id, {
method: "PUT",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
name: document.getElementById("etm-name").value.trim(),
category: document.getElementById("etm-category").value,
description: (
document.getElementById("etm-description").value || ""
).trim(),
tags: JSON.stringify(tagsArray),
author: (document.getElementById("etm-author").value || "").trim(),
activation: document.getElementById("etm-activation").value,
content: content,
variables: JSON.stringify(varList),
is_default: document.getElementById("etm-default").checked,
model: document.getElementById("esk-model").value.trim(),
auto_approve: document.getElementById("esk-auto-approve").checked,
temperature: esTemp ? parseFloat(esTemp) : null,
reasoning_effort: document.getElementById("esk-reasoning-effort").value,
max_tokens: esMaxTok ? parseInt(esMaxTok, 10) : null,
token_budget: esBudget ? parseInt(esBudget, 10) : 0,
agent_max_turns: esMaxTurns ? parseInt(esMaxTurns, 10) : null,
allowed_tools: JSON.stringify(esAllowedArr),
enabled: document.getElementById("esk-enabled").checked,
}),
body: JSON.stringify(updateBody),
})
.then(function (r) {
if (!r.ok)
+170 -91
View File
@@ -848,61 +848,100 @@ window.TURNSTONE_KB_SHORTCUTS = [
<!-- Create Skill Modal -->
<div id="create-template-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="create-template-title">
<div id="create-template-box" class="admin-modal admin-modal-wide">
<div id="create-template-box" class="admin-modal admin-modal-wide admin-modal-skill">
<h2 id="create-template-title">Create Skill</h2>
<div id="create-template-error" role="alert" aria-live="assertive"></div>
<label for="ctm-name">Name</label>
<input id="ctm-name" type="text" placeholder="e.g. Code Review Agent" autocomplete="off">
<label for="ctm-category">Category</label>
<select id="ctm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="skill-description">Description</label>
<textarea id="skill-description" rows="2" placeholder="Brief description for discovery"></textarea>
<label for="skill-tags">Tags</label>
<input id="skill-tags" type="text" placeholder="Comma-separated tags">
<label for="skill-author">Author</label>
<input id="skill-author" type="text" placeholder="Author name">
<label for="skill-activation">Activation</label>
<select id="skill-activation">
<option value="named">Named</option>
<option value="default">Default (auto-apply)</option>
<option value="search">Search (BM25 discoverable)</option>
</select>
<label for="ctm-content">Content <span class="label-hint">system message text, use {{model}}, {{ws_id}}, {{node_id}} for placeholders</span></label>
<textarea id="ctm-content" rows="6" placeholder="You are a code reviewer using {{model}}..."></textarea>
<label>Variables <span class="label-hint">auto-detected from content &mdash; available: model, ws_id, node_id</span></label>
<div id="ctm-variables" class="label-hint" style="padding:4px 0;min-height:1.2em"></div>
<label class="admin-checkbox"><input id="ctm-default" type="checkbox"> Set as default for new workstreams</label>
<div class="skill-spec-body">
<div class="skill-spec-col skill-spec-col-meta">
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Identity</h3>
<label for="ctm-name">Name</label>
<input id="ctm-name" type="text" placeholder="e.g. code-review" autocomplete="off">
<label for="ctm-category">Category</label>
<select id="ctm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="skill-description">Description</label>
<textarea id="skill-description" rows="2" placeholder="Brief description for skill discovery"></textarea>
</div>
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Manifest</h3>
<label for="skill-tags">Tags</label>
<input id="skill-tags" type="text" placeholder="python, review, quality">
<label for="skill-author">Author</label>
<input id="skill-author" type="text" placeholder="Author name">
<label for="skill-version">Version</label>
<input id="skill-version" type="text" placeholder="1.0.0">
<label for="skill-license">License</label>
<select id="skill-license">
<option value="">— not specified —</option>
<option value="MIT">MIT</option>
<option value="Apache-2.0">Apache-2.0</option>
<option value="GPL-2.0">GPL-2.0</option>
<option value="GPL-3.0">GPL-3.0</option>
<option value="LGPL-2.1">LGPL-2.1</option>
<option value="LGPL-3.0">LGPL-3.0</option>
<option value="AGPL-3.0">AGPL-3.0</option>
<option value="BSD-2-Clause">BSD-2-Clause</option>
<option value="BSD-3-Clause">BSD-3-Clause</option>
<option value="ISC">ISC</option>
<option value="MPL-2.0">MPL-2.0</option>
<option value="Unlicense">Unlicense</option>
<option value="Proprietary">Proprietary</option>
</select>
<label for="skill-compatibility">Compatibility <span class="label-hint">environment requirements, max 500 chars</span></label>
<input id="skill-compatibility" type="text" placeholder="Requires git, docker, etc." maxlength="500">
</div>
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Deployment</h3>
<label for="skill-activation">Activation <span class="label-hint">how models discover this skill</span></label>
<select id="skill-activation">
<option value="named">Named — explicit /skill invocation</option>
<option value="default">Default — auto-applied to every session</option>
<option value="search">Search — BM25 discoverable</option>
</select>
<label class="admin-checkbox"><input id="ctm-default" type="checkbox"> Apply to new workstreams by default</label>
</div>
</div>
<div class="skill-spec-col skill-spec-col-content">
<div class="skill-spec-section skill-spec-section-content">
<h3 class="skill-spec-heading">Skill Content <span class="label-hint">system message &mdash; {{model}}, {{ws_id}}, {{node_id}}</span></h3>
<textarea id="ctm-content" class="skill-content-area" placeholder="You are a code reviewer using {{model}}..."></textarea>
<div class="skill-vars-row">
<span class="skill-vars-label">Variables</span>
<div id="ctm-variables" class="skill-vars-display label-hint"></div>
</div>
</div>
</div>
</div>
<details class="admin-details">
<summary>Session Config <span class="label-hint">optional &mdash; applied when skill is selected for a workstream</span></summary>
<label for="csk-model">Model</label>
<input id="csk-model" type="text" placeholder="Default model">
<label for="csk-temperature">Temperature</label>
<input id="csk-temperature" type="number" step="0.1" min="0" max="2" placeholder="System default">
<label for="csk-reasoning-effort">Reasoning Effort</label>
<select id="csk-reasoning-effort">
<option value="">System default</option>
<option value="low">Low</option>
<option value="medium">Medium</option>
<option value="high">High</option>
</select>
<label for="csk-max-tokens">Max Tokens</label>
<input id="csk-max-tokens" type="number" min="1" placeholder="System default">
<label for="csk-token-budget">Token Budget</label>
<input id="csk-token-budget" type="number" min="0" placeholder="0 = unlimited">
<label for="csk-agent-max-turns">Agent Max Turns</label>
<input id="csk-agent-max-turns" type="number" min="1" placeholder="System default">
<summary>Runtime Config <span class="label-hint">model, temperature, token limits</span></summary>
<div class="skill-config-grid">
<div><label for="csk-model">Model</label><input id="csk-model" type="text" placeholder="Default model"></div>
<div><label for="csk-temperature">Temperature</label><input id="csk-temperature" type="number" step="0.1" min="0" max="2" placeholder="System default"></div>
<div>
<label for="csk-reasoning-effort">Reasoning Effort</label>
<select id="csk-reasoning-effort">
<option value="">System default</option>
<option value="low">Low</option>
<option value="medium">Medium</option>
<option value="high">High</option>
</select>
</div>
<div><label for="csk-max-tokens">Max Tokens</label><input id="csk-max-tokens" type="number" min="1" placeholder="System default"></div>
<div><label for="csk-token-budget">Token Budget</label><input id="csk-token-budget" type="number" min="0" placeholder="0 = unlimited"></div>
<div><label for="csk-agent-max-turns">Agent Max Turns</label><input id="csk-agent-max-turns" type="number" min="1" placeholder="System default"></div>
</div>
<label class="admin-checkbox"><input id="csk-auto-approve" type="checkbox"> Auto-approve all tools</label>
<label for="csk-allowed-tools">Allowed Tools <span class="label-hint">comma-separated tool names for auto-approve</span></label>
<input id="csk-allowed-tools" type="text" placeholder="bash, read_file, write_file">
<label class="admin-checkbox"><input id="csk-enabled" type="checkbox" checked> Enabled</label>
</details>
<details class="admin-details">
<summary>Resources <span class="label-hint">optional bundled files (scripts, references, assets)</span></summary>
<summary>Resources <span class="label-hint">bundled files (scripts, references, assets)</span></summary>
<div id="ctm-resources-list" role="list" aria-live="polite" aria-label="Pending resources"></div>
<div style="margin-top:8px;display:flex;flex-direction:column;gap:6px">
<label for="ctm-res-path">Path</label>
@@ -923,55 +962,95 @@ window.TURNSTONE_KB_SHORTCUTS = [
<!-- Edit Skill Modal -->
<div id="edit-template-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="edit-template-title">
<div id="edit-template-box" class="admin-modal admin-modal-wide">
<div id="edit-template-box" class="admin-modal admin-modal-wide admin-modal-skill">
<h2 id="edit-template-title">Edit Skill</h2>
<div id="etm-origin-badge" class="skill-origin-badge" style="display:none"></div>
<div id="edit-template-error" role="alert" aria-live="assertive"></div>
<input id="etm-id" type="hidden">
<label for="etm-name">Name</label>
<input id="etm-name" type="text" autocomplete="off">
<label for="etm-category">Category</label>
<select id="etm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="etm-description">Description</label>
<textarea id="etm-description" rows="2" placeholder="Brief description for discovery"></textarea>
<label for="etm-tags">Tags</label>
<input id="etm-tags" type="text" placeholder="Comma-separated tags">
<label for="etm-author">Author</label>
<input id="etm-author" type="text" placeholder="Author name">
<label for="etm-activation">Activation</label>
<select id="etm-activation">
<option value="named">Named</option>
<option value="default">Default (auto-apply)</option>
<option value="search">Search (BM25 discoverable)</option>
</select>
<label for="etm-content">Content</label>
<textarea id="etm-content" rows="6"></textarea>
<label>Variables <span class="label-hint">auto-detected from content &mdash; available: model, ws_id, node_id</span></label>
<div id="etm-variables" class="label-hint" style="padding:4px 0;min-height:1.2em"></div>
<label class="admin-checkbox"><input id="etm-default" type="checkbox"> Set as default</label>
<div class="skill-spec-body">
<div class="skill-spec-col skill-spec-col-meta">
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Identity</h3>
<label for="etm-name">Name</label>
<input id="etm-name" type="text" autocomplete="off">
<label for="etm-category">Category</label>
<select id="etm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="etm-description">Description</label>
<textarea id="etm-description" rows="2" placeholder="Brief description for skill discovery"></textarea>
</div>
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Manifest</h3>
<label for="etm-tags">Tags</label>
<input id="etm-tags" type="text" placeholder="python, review, quality">
<label for="etm-author">Author</label>
<input id="etm-author" type="text" placeholder="Author name">
<label for="etm-version">Version</label>
<input id="etm-version" type="text" placeholder="1.0.0">
<label for="etm-license">License</label>
<select id="etm-license">
<option value="">— not specified —</option>
<option value="MIT">MIT</option>
<option value="Apache-2.0">Apache-2.0</option>
<option value="GPL-2.0">GPL-2.0</option>
<option value="GPL-3.0">GPL-3.0</option>
<option value="LGPL-2.1">LGPL-2.1</option>
<option value="LGPL-3.0">LGPL-3.0</option>
<option value="AGPL-3.0">AGPL-3.0</option>
<option value="BSD-2-Clause">BSD-2-Clause</option>
<option value="BSD-3-Clause">BSD-3-Clause</option>
<option value="ISC">ISC</option>
<option value="MPL-2.0">MPL-2.0</option>
<option value="Unlicense">Unlicense</option>
<option value="Proprietary">Proprietary</option>
</select>
<label for="etm-compatibility">Compatibility <span class="label-hint">environment requirements, max 500 chars</span></label>
<input id="etm-compatibility" type="text" placeholder="Requires git, docker, etc." maxlength="500">
</div>
<div class="skill-spec-section">
<h3 class="skill-spec-heading">Deployment</h3>
<label for="etm-activation">Activation <span class="label-hint">how models discover this skill</span></label>
<select id="etm-activation">
<option value="named">Named — explicit /skill invocation</option>
<option value="default">Default — auto-applied to every session</option>
<option value="search">Search — BM25 discoverable</option>
</select>
<label class="admin-checkbox"><input id="etm-default" type="checkbox"> Apply to new workstreams by default</label>
</div>
</div>
<div class="skill-spec-col skill-spec-col-content">
<div class="skill-spec-section skill-spec-section-content">
<h3 class="skill-spec-heading">Skill Content <span class="label-hint">{{model}}, {{ws_id}}, {{node_id}}</span></h3>
<textarea id="etm-content" class="skill-content-area"></textarea>
<div class="skill-vars-row">
<span class="skill-vars-label">Variables</span>
<div id="etm-variables" class="skill-vars-display label-hint"></div>
</div>
</div>
</div>
</div>
<details class="admin-details">
<summary>Session Config <span class="label-hint">applied when skill is selected for a workstream</span></summary>
<label for="esk-model">Model</label>
<input id="esk-model" type="text" placeholder="Default model">
<label for="esk-temperature">Temperature</label>
<input id="esk-temperature" type="number" step="0.1" min="0" max="2" placeholder="System default">
<label for="esk-reasoning-effort">Reasoning Effort</label>
<select id="esk-reasoning-effort">
<option value="">System default</option>
<option value="low">Low</option>
<option value="medium">Medium</option>
<option value="high">High</option>
</select>
<label for="esk-max-tokens">Max Tokens</label>
<input id="esk-max-tokens" type="number" min="1" placeholder="System default">
<label for="esk-token-budget">Token Budget</label>
<input id="esk-token-budget" type="number" min="0" placeholder="0 = unlimited">
<label for="esk-agent-max-turns">Agent Max Turns</label>
<input id="esk-agent-max-turns" type="number" min="1" placeholder="System default">
<summary>Runtime Config <span class="label-hint">model, temperature, token limits</span></summary>
<div class="skill-config-grid">
<div><label for="esk-model">Model</label><input id="esk-model" type="text" placeholder="Default model"></div>
<div><label for="esk-temperature">Temperature</label><input id="esk-temperature" type="number" step="0.1" min="0" max="2" placeholder="System default"></div>
<div>
<label for="esk-reasoning-effort">Reasoning Effort</label>
<select id="esk-reasoning-effort">
<option value="">System default</option>
<option value="low">Low</option>
<option value="medium">Medium</option>
<option value="high">High</option>
</select>
</div>
<div><label for="esk-max-tokens">Max Tokens</label><input id="esk-max-tokens" type="number" min="1" placeholder="System default"></div>
<div><label for="esk-token-budget">Token Budget</label><input id="esk-token-budget" type="number" min="0" placeholder="0 = unlimited"></div>
<div><label for="esk-agent-max-turns">Agent Max Turns</label><input id="esk-agent-max-turns" type="number" min="1" placeholder="System default"></div>
</div>
<label class="admin-checkbox"><input id="esk-auto-approve" type="checkbox"> Auto-approve all tools</label>
<label for="esk-allowed-tools">Allowed Tools <span class="label-hint">comma-separated tool names for auto-approve</span></label>
<input id="esk-allowed-tools" type="text" placeholder="bash, read_file, write_file">
+140
View File
@@ -1166,6 +1166,14 @@
outline: none;
box-shadow: 0 0 0 3px var(--accent-dim);
}
.admin-modal input:disabled, .admin-modal select:disabled, .admin-modal textarea:disabled {
opacity: 0.55;
cursor: not-allowed;
background: var(--bg-highlight);
border-color: var(--border);
color: var(--fg-dim);
}
.admin-modal label.admin-checkbox input:disabled { opacity: 0.4; }
.admin-modal input::placeholder, .admin-modal textarea::placeholder { color: var(--fg-dim); opacity: 0.6; }
.admin-modal textarea { resize: vertical; min-height: 40px; }
.admin-modal [role="alert"] { color: var(--red); font-size: 12px; margin-bottom: 8px; display: none; }
@@ -1199,6 +1207,138 @@
.admin-details summary .label-hint { font-weight: 400; }
.admin-details label:first-of-type { margin-top: 4px; }
/* ==========================================================================
Skill Spec Modal two-column manifest layout
Left: Identity / Manifest / Deployment | Right: Skill Content
========================================================================== */
.admin-modal-skill { padding: 28px 28px 24px; }
.skill-spec-body {
display: grid;
grid-template-columns: 1fr 1.55fr;
gap: 0;
margin-bottom: 12px;
}
.skill-spec-col-meta {
border-right: 1px solid var(--border);
padding-right: 22px;
}
.skill-spec-col-content {
padding-left: 22px;
display: flex;
flex-direction: column;
}
.skill-spec-section { margin-bottom: 14px; }
.skill-spec-section:last-child { margin-bottom: 0; }
/* h3 used for screen-reader heading structure; reset UA defaults */
h3.skill-spec-heading { font-size: inherit; margin-block: 0; }
.skill-spec-heading {
font-family: var(--font-display);
font-size: 10px;
font-weight: 700;
text-transform: uppercase;
letter-spacing: 0.1em;
color: var(--accent);
padding-bottom: 5px;
margin: 14px 0 6px;
border-bottom: 1px solid var(--accent-dim);
}
.skill-spec-section:first-child .skill-spec-heading { margin-top: 0; }
.skill-spec-heading .label-hint {
text-transform: none;
letter-spacing: 0;
font-weight: 400;
font-size: 10px;
opacity: 1;
}
.skill-spec-section-content {
flex: 1;
display: flex;
flex-direction: column;
}
.skill-content-area {
flex: 1;
min-height: 220px;
font-family: var(--font-mono) !important;
font-size: 11.5px !important;
line-height: 1.65 !important;
}
.skill-vars-row {
display: flex;
align-items: center;
gap: 8px;
margin-top: 8px;
min-height: 18px;
}
.skill-vars-label {
font-family: var(--font-display);
font-size: 10px;
font-weight: 700;
text-transform: uppercase;
letter-spacing: 0.1em;
color: var(--fg-dim);
white-space: nowrap;
flex-shrink: 0;
}
.skill-vars-display { font-size: 11px; }
.skill-config-grid {
display: grid;
grid-template-columns: 1fr 1fr 1fr;
gap: 8px 14px;
margin: 4px 0 2px;
}
.skill-config-grid > div { min-width: 0; }
/* Origin badge — shown for remotely installed (readonly) skills */
.skill-origin-badge {
display: inline-flex;
align-items: center;
gap: 6px;
font-family: var(--font-display);
font-size: 9px;
font-weight: 700;
text-transform: uppercase;
letter-spacing: 0.1em;
color: var(--cyan);
background: rgba(103, 232, 249, 0.07);
border: 1px solid rgba(103, 232, 249, 0.18);
border-radius: var(--radius-sm);
padding: 5px 10px;
margin-bottom: 14px;
max-width: 100%;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.skill-origin-badge::before {
content: "\2193";
font-size: 11px;
flex-shrink: 0;
}
@media (max-width: 700px) {
.skill-spec-body { grid-template-columns: 1fr; }
.skill-spec-col-meta {
border-right: none;
padding-right: 0;
border-bottom: 1px solid var(--border);
padding-bottom: 16px;
margin-bottom: 16px;
}
.skill-spec-col-content { padding-left: 0; }
.skill-config-grid { grid-template-columns: 1fr 1fr; }
}
.modal-buttons { display: flex; gap: 10px; margin-top: 20px; }
.modal-cancel {
flex: 1;
+1 -2
View File
@@ -975,7 +975,7 @@ class IntentJudge:
"""Daemon thread: run LLM judge for each item and invoke callback."""
for item, h_verdict in zip(items, heuristic_verdicts, strict=True):
try:
llm_verdict = self._evaluate_single(item, messages, h_verdict)
llm_verdict = self._evaluate_single(item, messages)
# Arbitrate: only callback when LLM upgrades the heuristic
if llm_verdict and llm_verdict.confidence > h_verdict.confidence:
callback(llm_verdict)
@@ -990,7 +990,6 @@ class IntentJudge:
self,
item: dict[str, Any],
messages: list[dict[str, Any]],
heuristic: IntentVerdict,
) -> IntentVerdict | None:
"""Run LLM judge for a single tool call. Returns verdict or None."""
start = time.monotonic()
+9 -2
View File
@@ -179,10 +179,17 @@ def list_default_skills(org_id: str = "") -> list[dict[str, Any]]:
return []
def list_skills_by_activation(activation: str) -> list[dict[str, Any]]:
def list_skills_by_activation(
activation: str,
*,
enabled_only: bool = False,
limit: int = 0,
) -> list[dict[str, Any]]:
"""Return skills filtered by activation value, ordered by name."""
try:
return get_storage().list_skills_by_activation(activation)
return get_storage().list_skills_by_activation(
activation, enabled_only=enabled_only, limit=limit
)
except Exception:
return []
+100 -40
View File
@@ -39,6 +39,7 @@ from turnstone.core.memory import (
get_skill_by_name,
get_workstream_display_name,
list_default_skills,
list_skills_by_activation,
list_structured_memories,
list_workstreams_with_history,
load_messages,
@@ -128,6 +129,11 @@ _MAX_SKILL_CONTENT: int = 32768
_TEMPLATE_VAR_RE = re.compile(r"\{\{(\w+)\}\}")
def _without_tool(tools: list[dict[str, Any]], name: str) -> list[dict[str, Any]]:
"""Return *tools* with the named tool removed."""
return [t for t in tools if t.get("function", {}).get("name") != name]
def _render_template(content: str, context: dict[str, str]) -> str:
"""Replace ``{{variable}}`` placeholders in a single pass.
@@ -341,7 +347,6 @@ class ChatSession:
self._tool_search = ToolSearchManager(
self._tools,
always_on_names=set(BUILTIN_TOOL_NAMES),
threshold=tool_search_threshold,
max_results=tool_search_max_results,
)
# Skill: explicit name overrides is_default skills
@@ -360,11 +365,16 @@ class ChatSession:
def model_alias(self) -> str | None:
return self._model_alias
def _get_capabilities(self) -> ModelCapabilities:
def _resolve_capabilities(
self,
provider: LLMProvider,
model: str,
alias: str | None = None,
) -> ModelCapabilities:
"""Get model capabilities, applying config.toml overrides if present."""
caps = self._provider.get_capabilities(self.model)
if self._registry and self._model_alias:
cfg: ModelConfig = self._registry.get_config(self._model_alias)
caps = provider.get_capabilities(model)
if self._registry and alias:
cfg: ModelConfig = self._registry.get_config(alias)
if cfg.capabilities:
fields = {f.name for f in dataclasses.fields(type(caps))}
overrides = {k: v for k, v in cfg.capabilities.items() if k in fields}
@@ -372,6 +382,10 @@ class ChatSession:
caps = dataclasses.replace(caps, **overrides)
return caps
def _get_capabilities(self) -> ModelCapabilities:
"""Get capabilities for the current model."""
return self._resolve_capabilities(self._provider, self.model, self._model_alias)
def _save_config(self) -> None:
"""Persist LLM-affecting config so resumed workstreams behave identically."""
save_workstream_config(
@@ -512,7 +526,6 @@ class ChatSession:
self._tool_search = ToolSearchManager(
self._tools,
always_on_names=set(BUILTIN_TOOL_NAMES),
threshold=self._tool_search_threshold,
max_results=self._tool_search_max_results,
)
# Restore previously expanded tools that still exist
@@ -856,6 +869,28 @@ class ChatSession:
)
lines.append("</skill-resources>")
dev_parts.append("\n".join(lines))
# Skill catalog: disclose search-activated skills so the model
# knows they exist (Agent Skills standard progressive disclosure).
try:
search_skills = list_skills_by_activation("search", enabled_only=True, limit=30)
except Exception:
log.warning("session.skill_catalog_failed", exc_info=True)
search_skills = []
if search_skills:
catalog_lines = ["<available-skills>"]
for sk in search_skills[:30]:
sk_name = _html_escape(sk.get("name", ""))
sk_desc = _html_escape(sk.get("description", "")[:200])
catalog_lines.append(
f" <skill><name>{sk_name}</name><description>{sk_desc}</description></skill>"
)
catalog_lines.append("</available-skills>")
catalog_lines.append(
"Additional skills are available. When a task matches a skill "
"description, ask the user to activate it with `/skill <name>`, "
"or use `/skill search <query>` to find relevant skills."
)
dev_parts.append("\n".join(catalog_lines))
if self.instructions:
dev_parts.append("")
dev_parts.append(self.instructions)
@@ -915,19 +950,29 @@ class ChatSession:
- Client-side fallback: send visible tools + synthetic tool_search.
Without tool search: return self._tools unchanged.
Web search gating: ``web_search`` is removed when the model has
no native search support and no Tavily API key is configured.
"""
if self.creative_mode:
return None
if not self._tool_search:
return self._tools
# Check if provider supports native tool search
caps = self._get_capabilities()
if caps.supports_tool_search:
# Provider handles defer_loading — send all tools
return self._tools
# Client-side fallback: visible tools + search tool
visible = self._tool_search.get_visible_tools()
return visible + [self._tool_search.get_search_tool_definition()]
if not self._tool_search:
tools = self._tools
else:
if caps.supports_tool_search:
# Provider handles defer_loading — send all tools
tools = self._tools
else:
# Client-side fallback: visible tools + search tool
visible = self._tool_search.get_visible_tools()
tools = visible + [self._tool_search.get_search_tool_definition()]
# Gate web_search: only include when a backend exists
if not caps.supports_web_search and not get_tavily_key():
tools = _without_tool(tools, "web_search")
return tools
def _get_deferred_names(self) -> frozenset[str] | None:
"""Return names of deferred tools for native provider search, or None."""
@@ -1936,7 +1981,7 @@ class ChatSession:
it["func_args"] = {"url": it.get("url", ""), "question": it.get("question", "")}
elif name == "web_search":
it["func_args"] = {"query": it.get("query", ""), "topic": it.get("topic", "")}
elif name == "load_skill":
elif name == "skill":
it["func_args"] = {"action": it.get("action", ""), "name": it.get("name", "")}
elif name == "watch":
it["func_args"] = {
@@ -2057,8 +2102,17 @@ class ChatSession:
return item["call_id"], item["error"]
if item.get("denied"):
return item["call_id"], item.get("denial_msg", "Denied by user")
result: tuple[str, str | list[dict[str, Any]]] = item["execute"](item)
return result
try:
result: tuple[str, str | list[dict[str, Any]]] = item["execute"](item)
return result
except (KeyboardInterrupt, GenerationCancelled):
raise
except Exception as e:
func = item.get("func_name", "unknown")
msg = f"Error executing {func}: {e}"
log.warning("tool_exec.failed", tool=func, error=str(e), exc_info=True)
self.ui.on_error(msg)
return item["call_id"], msg
if len(items) == 1:
results = [run_one(items[0])]
@@ -2203,7 +2257,7 @@ class ChatSession:
"watch": self._prepare_watch,
"read_resource": self._prepare_read_resource,
"use_prompt": self._prepare_use_prompt,
"load_skill": self._prepare_load_skill,
"skill": self._prepare_skill,
}
preparer = preparers.get(func_name)
if not preparer:
@@ -3027,10 +3081,10 @@ class ChatSession:
"limit": max(1, min(limit, 50)),
}
# -- load_skill prepare/execute --------------------------------------------
# -- skill prepare/execute -------------------------------------------------
def _prepare_load_skill(self, call_id: str, args: dict[str, Any]) -> dict[str, Any]:
"""Prepare a load_skill action (load or search)."""
def _prepare_skill(self, call_id: str, args: dict[str, Any]) -> dict[str, Any]:
"""Prepare a skill action (load or search)."""
action = (args.get("action") or "").strip().lower()
if action == "load":
@@ -3038,20 +3092,20 @@ class ChatSession:
if not name:
return {
"call_id": call_id,
"func_name": "load_skill",
"header": "\u2717 load_skill: name is required",
"func_name": "skill",
"header": "\u2717 skill: name is required",
"preview": "",
"needs_approval": False,
"error": "Error: 'name' is required for load action",
}
return {
"call_id": call_id,
"func_name": "load_skill",
"header": f"\u2699 load_skill: {name}",
"func_name": "skill",
"header": f"\u2699 skill: {name}",
"preview": "",
"needs_approval": True,
"approval_label": f"load_skill__{name}",
"execute": self._exec_load_skill,
"approval_label": f"skill__{name}",
"execute": self._exec_skill,
"action": "load",
"name": name,
}
@@ -3060,26 +3114,26 @@ class ChatSession:
query = (args.get("query") or "").strip()
return {
"call_id": call_id,
"func_name": "load_skill",
"func_name": "skill",
"header": f"\u2699 skill search{': ' + query[:80] if query else ''}",
"preview": "",
"needs_approval": False,
"execute": self._exec_load_skill,
"execute": self._exec_skill,
"action": "search",
"query": query,
}
return {
"call_id": call_id,
"func_name": "load_skill",
"header": "\u2717 load_skill: invalid action",
"func_name": "skill",
"header": "\u2717 skill: invalid action",
"preview": "",
"needs_approval": False,
"error": f"Error: action must be 'load' or 'search', got '{action}'",
}
def _exec_load_skill(self, item: dict[str, Any]) -> tuple[str, str]:
"""Execute a load_skill action."""
def _exec_skill(self, item: dict[str, Any]) -> tuple[str, str]:
"""Execute a skill action."""
call_id = item["call_id"]
action = item["action"]
@@ -3088,12 +3142,12 @@ class ChatSession:
skill_data = get_skill_by_name(name)
if not skill_data or not skill_data.get("enabled", True):
msg = f"Error: skill '{name}' not found"
self.ui.on_tool_result(call_id, "load_skill", msg)
self.ui.on_tool_result(call_id, "skill", msg)
return call_id, msg
if self._skill_name == name:
msg = f"Skill '{name}' is already active"
self.ui.on_tool_result(call_id, "load_skill", msg)
self.ui.on_tool_result(call_id, "skill", msg)
return call_id, msg
self.set_skill(name)
@@ -3106,7 +3160,7 @@ class ChatSession:
if scan:
parts.append(f"Security tier: {scan}")
msg = "\n".join(parts)
self.ui.on_tool_result(call_id, "load_skill", msg)
self.ui.on_tool_result(call_id, "skill", msg)
return call_id, msg
# action == "search"
@@ -3116,7 +3170,7 @@ class ChatSession:
rows = get_storage().list_prompt_templates(limit=50)
except Exception:
log.warning("load_skill.search_storage_error", exc_info=True)
log.warning("skill.search_storage_error", exc_info=True)
rows = []
# Filter out disabled skills
@@ -3160,7 +3214,7 @@ class ChatSession:
if not rows:
msg = "No skills found" + (f" matching '{query}'" if query else "")
self.ui.on_tool_result(call_id, "load_skill", msg)
self.ui.on_tool_result(call_id, "skill", msg)
return call_id, msg
lines = [f"Found {len(rows)} skill(s):", ""]
@@ -3182,7 +3236,7 @@ class ChatSession:
lines.append(line)
msg = "\n".join(lines)
self.ui.on_tool_result(call_id, "load_skill", msg)
self.ui.on_tool_result(call_id, "skill", msg)
return call_id, msg
# -- MCP tool prepare/execute ----------------------------------------------
@@ -3671,6 +3725,12 @@ class ChatSession:
agent_client, agent_model, _ = self._registry.resolve(self._registry.agent_model)
agent_provider = self._registry.get_provider(self._registry.agent_model)
# Gate web_search: remove when no backend exists for the agent model
agent_alias = self._registry.agent_model if self._registry else None
agent_caps = self._resolve_capabilities(agent_provider, agent_model, agent_alias)
if not agent_caps.supports_web_search and not get_tavily_key():
tools = _without_tool(tools, "web_search")
# Build extra params for agent calls
agent_extra: dict[str, Any] | None = None
if agent_provider.provider_name == "openai":
+148 -23
View File
@@ -2,19 +2,40 @@
Pure functions, no I/O. Accepts raw SKILL.md text and returns a
:class:`ParsedSkill` dataclass.
Compliant with the Agent Skills specification (https://agentskills.io/specification).
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from typing import Any
from typing import Any, Literal, overload
import frontmatter
# Name validation: lowercase letters, digits, hyphens, max 64 chars
from turnstone.core.log import get_logger
log = get_logger(__name__)
# Name validation: lowercase letters, digits, hyphens, max 64 chars.
# Note: consecutive hyphens checked separately (not expressible in a
# single character-class regex without a lookahead).
_NAME_RE = re.compile(r"^[a-z0-9][a-z0-9\-]{0,62}[a-z0-9]$|^[a-z0-9]$")
# Split allowed-tools on whitespace or commas (standard uses spaces,
# legacy turnstone format uses commas). Tool expressions must not
# contain internal whitespace (e.g. "Bash(git:*)" not "Bash(git: *)").
_LIST_SPLIT_RE = re.compile(r"[\s,]+")
# Malformed YAML recovery: match a bare ``description:`` line whose
# value contains an unquoted colon (the most common cross-client issue).
_BARE_DESC_RE = re.compile(r"^(description:\s*)(.+)$", re.MULTILINE)
# Field length caps from the Agent Skills specification.
_MAX_DESCRIPTION_LEN = 1024
_MAX_COMPATIBILITY_LEN = 500
@dataclass(frozen=True)
class ParsedSkill:
@@ -55,13 +76,39 @@ def _extract_tags(meta: dict[str, Any]) -> list[str]:
return []
def _extract_list(meta: dict[str, Any], key: str) -> list[str]:
"""Extract a list of strings from frontmatter, with fallback."""
val = meta.get(key)
if isinstance(val, list):
return [str(v) for v in val if v]
if isinstance(val, str) and val:
return [v.strip() for v in val.split(",") if v.strip()]
def _extract_str(meta: dict[str, Any], key: str, default: str = "") -> str:
"""Extract a string field, checking top-level then ``metadata.*`` fallback.
Handles YAML ``null`` / bare keys gracefully (returns *default*
rather than the string ``"None"``).
"""
raw = meta.get(key)
val = str(raw).strip() if raw is not None else ""
if val:
return val
# Standard puts author/version under metadata map
nested = meta.get("metadata")
if isinstance(nested, dict):
raw = nested.get(key)
val = str(raw).strip() if raw is not None else ""
if val:
return val
return default
def _extract_list(meta: dict[str, Any], *keys: str) -> list[str]:
"""Extract a list of strings from frontmatter.
Tries each *key* in order (first match wins). String values are
split on whitespace or commas to handle both the Agent Skills
standard (space-delimited) and legacy comma-delimited formats.
"""
for key in keys:
val = meta.get(key)
if isinstance(val, list):
return [str(v) for v in val if v]
if isinstance(val, str) and val:
return [v for v in _LIST_SPLIT_RE.split(val) if v]
return []
@@ -71,22 +118,66 @@ def validate_skill_name(name: str) -> str | None:
return "name is required"
if len(name) > 64:
return f"name exceeds 64 characters ({len(name)})"
if "--" in name:
return "name must not contain consecutive hyphens"
if not _NAME_RE.match(name):
return "name must be lowercase alphanumeric with hyphens (e.g. 'code-review')"
return None
def parse_skill_md(raw: str) -> ParsedSkill:
"""Parse SKILL.md (YAML frontmatter + markdown body).
def _try_parse_frontmatter(raw: str) -> frontmatter.Post:
"""Parse YAML frontmatter with a single malformed-YAML retry.
Handles missing or malformed frontmatter gracefully returns a
``ParsedSkill`` with defaults for any missing fields.
Raises ``ValueError`` if ``name`` is missing or invalid.
The most common cross-client issue is unquoted description values
containing colons (e.g. ``description: Use when: the user asks``).
On initial failure, wrap the description value in quotes and retry.
"""
try:
post = frontmatter.loads(raw)
return frontmatter.loads(raw)
except Exception:
pass # fall through to retry
# Retry: quote the description line
def _quote_desc(m: re.Match[str]) -> str:
prefix = m.group(1)
value = m.group(2).strip()
escaped = value.replace('"', '\\"')
return f'{prefix}"{escaped}"'
fixed = _BARE_DESC_RE.sub(_quote_desc, raw)
if fixed != raw:
try:
return frontmatter.loads(fixed)
except Exception:
pass
raise ValueError("Failed to parse SKILL.md YAML frontmatter")
@overload
def parse_skill_md(raw: str, *, lenient: Literal[False] = ...) -> ParsedSkill: ...
@overload
def parse_skill_md(raw: str, *, lenient: Literal[True]) -> ParsedSkill | None: ...
def parse_skill_md(raw: str, *, lenient: bool = False) -> ParsedSkill | None:
"""Parse SKILL.md (YAML frontmatter + markdown body).
When *lenient* is ``False`` (default strict mode), raises
``ValueError`` on missing/invalid name or unparseable YAML.
When *lenient* is ``True`` (for external import / cross-client
ingestion), logs warnings and returns ``None`` for unskippable
failures instead of raising.
"""
try:
post = _try_parse_frontmatter(raw)
except Exception as exc:
if lenient:
log.warning("skill_parser.yaml_failed", error=str(exc))
return None
raise ValueError(f"Failed to parse SKILL.md frontmatter: {exc}") from exc
meta: dict[str, Any] = dict(post.metadata)
@@ -96,10 +187,20 @@ def parse_skill_md(raw: str) -> ParsedSkill:
name = str(meta.get("name", "")).strip().lower()
name_err = validate_skill_name(name)
if name_err:
raise ValueError(name_err)
if lenient:
log.warning("skill_parser.name_invalid", name=name, error=name_err)
# Try to salvage: strip invalid chars, truncate
sanitized = re.sub(r"[^a-z0-9-]", "", name).strip("-")
sanitized = re.sub(r"-{2,}", "-", sanitized)[:64].strip("-")
if not sanitized or validate_skill_name(sanitized):
return None
name = sanitized
else:
raise ValueError(name_err)
# Description — frontmatter or first paragraph of body
description = str(meta.get("description", "")).strip()
raw_desc = meta.get("description")
description = str(raw_desc).strip() if raw_desc is not None else ""
if not description and body:
first_line = body.split("\n")[0].strip()
# Skip markdown headings
@@ -107,15 +208,39 @@ def parse_skill_md(raw: str) -> ParsedSkill:
first_line = first_line.lstrip("# ").strip()
description = first_line[:256]
if not description and lenient:
log.warning("skill_parser.no_description", name=name)
return None
# Spec caps
if len(description) > _MAX_DESCRIPTION_LEN:
log.warning(
"skill_parser.description_truncated",
name=name,
length=len(description),
)
description = description[:_MAX_DESCRIPTION_LEN]
raw_compat = meta.get("compatibility")
compatibility = str(raw_compat).strip() if raw_compat is not None else ""
if len(compatibility) > _MAX_COMPATIBILITY_LEN:
log.warning(
"skill_parser.compatibility_truncated",
name=name,
length=len(compatibility),
)
compatibility = compatibility[:_MAX_COMPATIBILITY_LEN]
return ParsedSkill(
name=name,
description=description,
content=body,
tags=_extract_tags(meta),
author=str(meta.get("author", "")).strip(),
version=str(meta.get("version", "1.0.0")).strip(),
allowed_tools=_extract_list(meta, "allowed_tools"),
license=str(meta.get("license", "")).strip(),
compatibility=str(meta.get("compatibility", "")).strip(),
author=_extract_str(meta, "author"),
version=_extract_str(meta, "version", default="1.0.0"),
# Standard uses "allowed-tools" (hyphenated); stored internally as allowed_tools
allowed_tools=_extract_list(meta, "allowed-tools"),
license=_extract_str(meta, "license"),
compatibility=compatibility,
raw_frontmatter=meta,
)
+18 -3
View File
@@ -1508,6 +1508,8 @@ class PostgreSQLBackend:
notify_on_complete: str = "{}",
enabled: bool = True,
allowed_tools: str = "[]",
skill_license: str = "",
compatibility: str = "",
) -> None:
# Sync is_default from activation when activation is explicitly set
if activation == "default":
@@ -1540,6 +1542,8 @@ class PostgreSQLBackend:
"activation": activation,
"token_estimate": token_estimate,
"allowed_tools": allowed_tools,
"license": skill_license,
"compatibility": compatibility,
"scan_status": scan_status,
"scan_report": scan_report,
"scan_version": scan_version,
@@ -1679,13 +1683,24 @@ class PostgreSQLBackend:
conn.commit()
return result.rowcount > 0
def list_skills_by_activation(self, activation: str) -> list[dict[str, Any]]:
def list_skills_by_activation(
self,
activation: str,
*,
enabled_only: bool = False,
limit: int = 0,
) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
q = (
sa.select(prompt_templates)
.where(prompt_templates.c.activation == activation)
.order_by(prompt_templates.c.name)
).fetchall()
)
if enabled_only:
q = q.where(prompt_templates.c.enabled == 1)
if limit > 0:
q = q.limit(limit)
rows = conn.execute(q).fetchall()
return [
_row_to_dict(r, "is_default", "readonly", "auto_approve", "enabled") for r in rows
]
+9 -1
View File
@@ -579,6 +579,8 @@ class StorageBackend(Protocol):
notify_on_complete: str = "{}",
enabled: bool = True,
allowed_tools: str = "[]",
skill_license: str = "",
compatibility: str = "",
) -> None:
"""Create a prompt template (skill)."""
...
@@ -617,7 +619,13 @@ class StorageBackend(Protocol):
"""Count prompt templates, optionally filtered by org_id."""
...
def list_skills_by_activation(self, activation: str) -> list[dict[str, Any]]:
def list_skills_by_activation(
self,
activation: str,
*,
enabled_only: bool = False,
limit: int = 0,
) -> list[dict[str, Any]]:
"""Return prompt templates filtered by activation value, ordered by name."""
...
+2
View File
@@ -311,6 +311,8 @@ prompt_templates = sa.Table(
sa.Column("activation", sa.Text, nullable=False, server_default="named"),
sa.Column("token_estimate", sa.Integer, nullable=False, server_default="0"),
sa.Column("allowed_tools", sa.Text, nullable=False, server_default="[]"), # JSON array
sa.Column("license", sa.Text, nullable=False, server_default=""),
sa.Column("compatibility", sa.Text, nullable=False, server_default=""),
sa.Column("scan_status", sa.Text, nullable=False, server_default=""),
sa.Column("scan_report", sa.Text, nullable=False, server_default="{}"), # JSON
sa.Column("installed_at", sa.Text, nullable=False, server_default=""),
+18 -3
View File
@@ -1532,6 +1532,8 @@ class SQLiteBackend:
notify_on_complete: str = "{}",
enabled: bool = True,
allowed_tools: str = "[]",
skill_license: str = "",
compatibility: str = "",
) -> None:
# Sync is_default from activation when activation is explicitly set
if activation == "default":
@@ -1564,6 +1566,8 @@ class SQLiteBackend:
"activation": activation,
"token_estimate": token_estimate,
"allowed_tools": allowed_tools,
"license": skill_license,
"compatibility": compatibility,
"scan_status": scan_status,
"scan_report": scan_report,
"scan_version": scan_version,
@@ -1703,13 +1707,24 @@ class SQLiteBackend:
conn.commit()
return result.rowcount > 0
def list_skills_by_activation(self, activation: str) -> list[dict[str, Any]]:
def list_skills_by_activation(
self,
activation: str,
*,
enabled_only: bool = False,
limit: int = 0,
) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
q = (
sa.select(prompt_templates)
.where(prompt_templates.c.activation == activation)
.order_by(prompt_templates.c.name)
).fetchall()
)
if enabled_only:
q = q.where(prompt_templates.c.enabled == 1)
if limit > 0:
q = q.limit(limit)
rows = conn.execute(q).fetchall()
return [
_row_to_dict(r, "is_default", "readonly", "auto_approve", "enabled") for r in rows
]
+2
View File
@@ -54,6 +54,8 @@ SKILL_MUTABLE = frozenset(
"notify_on_complete",
"enabled",
"allowed_tools",
"license",
"compatibility",
"scan_version",
"scan_status",
"scan_report",
@@ -0,0 +1,34 @@
"""Add license and compatibility columns to prompt_templates.
Agent Skills standard (agentskills.io) defines license and compatibility
as optional SKILL.md frontmatter fields. These were parsed but discarded
prior to this migration.
Revision ID: 023
Revises: 022
Create Date: 2026-03-17
"""
import sqlalchemy as sa
from alembic import op
revision = "023"
down_revision = "022"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column(
"prompt_templates",
sa.Column("license", sa.Text, nullable=False, server_default=""),
)
op.add_column(
"prompt_templates",
sa.Column("compatibility", sa.Text, nullable=False, server_default=""),
)
def downgrade() -> None:
op.drop_column("prompt_templates", "compatibility")
op.drop_column("prompt_templates", "license")
-11
View File
@@ -64,15 +64,12 @@ class ToolSearchManager:
all_tools: list[dict[str, Any]],
always_on_names: set[str],
*,
threshold: int = 20,
max_results: int = 5,
) -> None:
self._all_tools = all_tools
self._always_on: list[dict[str, Any]] = []
self._deferred: list[dict[str, Any]] = []
self._deferred_by_name: dict[str, dict[str, Any]] = {}
self._expanded: dict[str, None] = {} # ordered set (preserves discovery order)
self._threshold = threshold
self._max_results = max_results
for tool in all_tools:
@@ -90,10 +87,6 @@ class ToolSearchManager:
# Pre-compute server summary for the search tool description
self._server_hint = _mcp_server_summary(self._deferred)
def should_activate(self) -> bool:
"""Return True if tool search should be active (enough tools)."""
return len(self._all_tools) > self._threshold
def get_visible_tools(self) -> list[dict[str, Any]]:
"""Return always-on tools + any expanded (discovered) tools."""
result = list(self._always_on)
@@ -107,10 +100,6 @@ class ToolSearchManager:
"""Return tools that are currently deferred (not yet discovered)."""
return [t for t in self._deferred if _tool_name(t) not in self._expanded]
def get_all_tools(self) -> list[dict[str, Any]]:
"""Return the full tool list (for native provider modes)."""
return list(self._all_tools)
def search(self, query: str) -> list[dict[str, Any]]:
"""Search deferred tools by query, return top-k matches.
-1
View File
@@ -32,7 +32,6 @@ MAX_WATCHES_PER_WS = 5
MIN_INTERVAL = 10 # seconds
MAX_INTERVAL = 86_400 # 24 hours
DEFAULT_MAX_POLLS = 100
DEFAULT_INTERVAL = 300 # 5 minutes
MAX_OUTPUT_SIZE = 65_536 # truncate stored/dispatched output at 64 KB
# Safe builtins exposed to condition expressions.
+1 -3
View File
@@ -66,9 +66,7 @@ class SimEngine:
self._config = config
self._rng = rng or random.Random(config.seed)
async def simulate_llm_response(
self, first_round: bool, turn_number: int
) -> tuple[str, list[dict[str, Any]]]:
async def simulate_llm_response(self, first_round: bool) -> tuple[str, list[dict[str, Any]]]:
"""Simulate an LLM response.
Returns ``(content_text, tool_calls)`` where *tool_calls* may be
-1
View File
@@ -75,7 +75,6 @@ class SimWorkstream:
self._set_state("thinking", correlation_id)
content, tool_calls = await self._engine.simulate_llm_response(
rounds == 0,
self._turn_count,
)
await self._stream_content(content, correlation_id)
+1 -10
View File
@@ -5,7 +5,7 @@ from __future__ import annotations
import asyncio
import logging
import time
from typing import TYPE_CHECKING, Any, Protocol
from typing import TYPE_CHECKING, Any
from turnstone.mq.broker import RedisBroker
from turnstone.mq.protocol import SendMessage
@@ -18,15 +18,6 @@ if TYPE_CHECKING:
log = logging.getLogger("turnstone.sim.scenario")
class Scenario(Protocol):
async def run(
self,
cluster: SimCluster,
config: SimConfig,
metrics: MetricsCollector,
) -> None: ...
class SteadyStateScenario:
"""Inject messages at a constant rate for the configured duration."""
@@ -1,5 +1,5 @@
{
"name": "load_skill",
"name": "skill",
"description": "Load or search for skills. Actions: 'load' activates a skill by name (replaces current skill), 'search' finds available skills by query.",
"parameters": {
"type": "object",
Generated
+1 -1
View File
@@ -2168,7 +2168,7 @@ wheels = [
[[package]]
name = "turnstone"
version = "0.8.2"
version = "0.8.3"
source = { editable = "." }
dependencies = [
{ name = "alembic" },