Three issues caught by Copilot on the initial PR:
1. SKILL.md size cap was measured in code points, not UTF-8 bytes.
`len(str)` is a *lower* bound on encoded byte length — multi-byte
chars (emoji, CJK) inflate up to 4×, so a 100k-emoji SKILL.md
(400KB encoded) would slip past the 256KB cap. Switch to
`len(contents.encode("utf-8"))` and surface lone-surrogate failures
as SkillSourceError instead of dropping them silently. New
regression test feeds emoji content.
2. _skills_sh_source_url did not normalize the skill_id, so a sloppy
id from `/api/search` (whitespace, surrounding slashes) would pass
`_split_skills_sh_id`'s charset check (which strips first) and
produce a malformed persisted source_url that broke the
discover-UI dedup contract. Strip the id inside the helper, and
reconstruct the canonical id from validated parts in
download_skill's listing so downstream callers never see the raw
input.
3. The catch-all `except Exception:` around create_prompt_template
relabeled every storage failure (DB connection, disk full,
permission errors) as "conflict", masking operational issues.
Translate IntegrityError → StorageConflictError at the storage
shim (matching the pattern already used for OIDC user
provisioning) in both sqlite and postgres backends, then catch
StorageConflictError specifically in the install handler. Real
conflicts → "conflict" + warning; other exceptions → new
"internal error" reason + log.exception.
Tests: +3 (oversized multibyte SKILL.md, source_url normalization,
storage-layer conflict translation). 226 passing.
The skills.sh install path was failing with 404s because their public
API surface changed: /api/skills/{id} is gone, replaced by
/api/skill/[owner]/[repo]/[skill] (auth-walled) and
/api/download/[owner]/[repo]/[skill] (unauthenticated, returns the
SKILL.md + bundled resources inline as JSON). The error was not
surfacing in logs because admin_skill_install had a silent
`except Exception:` around create_prompt_template that relabeled every
storage failure as "conflict" with no log entry.
- Replace SkillsShClient.resolve_github_url with download_skill that
hits /api/download/{owner}/{repo}/{skill} and returns a SkillPackage
directly. No GitHub round-trip; no rate-limit surface.
- Add _split_skills_sh_id with strict per-segment charset validation
([A-Za-z0-9._-]+) so URL-hostile content can't produce a malformed
request or divergent persisted source_url.
- Use len(contents) instead of len(contents.encode("utf-8",
errors="ignore")) for the SKILL.md size cap — errors='ignore' was
silently dropping invalid units, making the cap bypassable.
- Extract _accept_resource(rel_path, byte_size) gate predicate; share
it between download_skill and the GitHub _find_resource_files helper.
- Have search() derive a deterministic source_url from the skill id
when /api/search omits one (which it currently always does), so the
discover-UI "already installed" check matches what download_skill
persists.
- Add structured logging across admin_skill_install and
admin_skill_discover: a shared _log_install_failure helper for the
four except branches (was four near-duplicate log calls with one
drift), plus per-resource failure tallying — partial-resource
installs now surface failed_resources in the response and audit
record instead of silently committing the skill row with missing
assets.
Tests: 7 new — empty/non-list files, oversized SKILL.md, resource
cap, non-text extension filtering, plus _split_skills_sh_id charset
rejection (whitespace, query chars). Verified end-to-end against
live skills.sh with tavily-search.
* feat: skill discovery — search and install skills from external sources
Add discovery UI and API for finding and installing skills from
skills.sh registries and GitHub repositories with one-click install,
SKILL.md frontmatter parsing, and security scan integration.
Core modules:
- skill_parser.py: ParsedSkill dataclass, parse_skill_md() with YAML
frontmatter support (Anthropic + Hermes tag formats), name validation
- skill_sources.py: SkillsShClient (async search + resolve),
fetch_skill_from_github (SKILL.md + bundled resource fetching with
256KB cap, text extension filter, GitHub API tree traversal)
API:
- GET /v1/api/admin/skills/discover — search with installed annotation
and scan_status for installed skills
- POST /v1/api/admin/skills/install — fetch, parse, duplicate check,
create with origin="source" readonly=true, store resources, audit
Also fixes pre-existing bug where _skill_to_response omitted scan_status,
scan_report, scan_version fields — scan tier badges in the installed
skills table were silently empty despite data existing in storage.
Admin UI: pill toggle (Installed/Discover), discovery cards with scan
tier badges, GitHub import modal with proper focus trap/Escape/backdrop,
scoped selectors preventing MCP↔Skills cross-tab state corruption.
SDK: discover_skills() + install_skill() on Python (async+sync) and
TypeScript console clients.
48 new tests across 3 test files. All 2632 tests pass.
* fix: address copilot review — 404 vs 502, O(n) lookups, branch fallback
- SkillNotFoundError subclass: install returns 404 when SKILL.md is
missing, 502 only for connectivity/upstream errors
- get_skill_by_source_url() + list_installed_skill_urls(): indexed
storage lookups replace O(n) full-table scans with content blobs
- Default branch fallback: tries main then master when URL doesn't
specify a branch
- Path normalization: strip trailing slash once, remove redundant
candidate
- SDK install_skill() returns typed SkillInfo with response_model
- Tree size guard: skip resource tree if response >2MB