Peter Steinberger
97b13b1735
fix(llama-cpp): preserve managed local embedding batch capacity ( #128721 )
...
Preserve the contributor fix from #127288 and its original Git author. Thanks @cantoblanco for the original issue and pull request.
Co-authored-by: Alex <alex@example.com >
2026-08-24 05:46:59 -07:00
Peter Steinberger
01e8887959
refactor(providers): return prepared dynamic models directly ( #126574 )
2026-08-21 11:40:52 -07:00
Peter Steinberger
4118f31d89
fix(llama-cpp): make endpoint auth transitions reproducible ( #126498 )
2026-08-19 18:37:28 -07:00
Peter Steinberger
0135046830
refactor(llama-cpp): use one provider for managed and existing servers ( #126434 )
...
* refactor(llama-cpp): unify server ownership modes
* test(llama-cpp): preserve shared discovery limits
* fix(plugin-sdk): retain provider auth removal export
2026-08-19 13:57:33 -07:00
Onur Solmaz
c2de3206d4
feat(llama-cpp): support external llama-server
...
* feat(llama-cpp): add external server provider
* feat(llama-cpp): document external server setup
* refactor(llama-cpp): harden external provider boundaries
* fix(llama-cpp): support external structured output
* fix(llama-cpp): isolate replacement endpoint credentials
* test(llama-cpp): register external live shard
* fix(llama-cpp): preserve explicit endpoint authorization
* fix(llama-cpp): clear disabled inline credentials
* fix(llama-cpp): preserve external local service configs
* test(llama-cpp): cover retained external configs
* test(llama-cpp): cover authorization precedence
2026-08-19 17:32:00 +03:00
Jacqueline Henriksen
72c227c7cc
fix(llama): support embedding-only managed servers ( #125383 )
...
* fix(llama): support embedding-only managed servers
* docs(llama): describe embedding-only setup
* fix(llama): remove unused model reference export
* fix(llama-cpp): keep guided setup chat-capable
---------
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com >
2026-08-17 23:29:43 -07:00
Peter Steinberger
6aa27d6ecd
refactor: retire August compat windows (embedding API, pi aliases, target parser, spawning hook, setup exports, WhatsApp inbound aliases) ( #124416 )
...
* refactor(plugin-sdk): retire embedded Pi aliases
* refactor(channels): retire explicit target compatibility
* refactor(plugins): retire subagent spawning hook
* refactor(plugin-sdk): retire shipped channel setup exports
* refactor(whatsapp): retire inbound callback aliases
Proof: focused build and WhatsApp E2E green; broad WhatsApp suite 188/189 files green. extensions/whatsapp/src/monitor-inbox.policy.test.ts flakes only in the parallel batch and passes isolated (10/10).
* refactor(plugin-sdk): retire memory embedding registrar
Migrate every bundled provider and manifest to registerEmbeddingProvider and contracts.embeddingProviders. Preserve memory-specific batching, local-service acquisition, index identity, and auto-selection through the canonical generic registry adapter, then remove the parallel registrar, registry, diagnostics, contracts, tests, and docs.
* chore(plugin-sdk): tighten retired surface budgets
Pin the post-retirement public SDK surface to 144 entrypoints, 4,312 exports, 2,564 callable exports, and 1,133 deprecated exports; agent-harness-runtime now permits exactly nine deprecated exports.
2026-08-15 22:43:47 -07:00
Jason (Json)
5f1bbed42d
fix(doctor): report missing managed local embedding setup ( #123575 )
...
* fix(gateway): expose startup blockers before cutover
* fix(gateway): include session blockers in preflight
* fix(gateway): keep preflight finding type private
* fix(gateway): preflight startup auth blockers
* fix(gateway): complete startup preflight readiness
* fix(llama-cpp): keep preflight remediation private
* fix(gateway): keep preflight passive and activation-aware
* fix(gateway): apply startup guard in preflight
* fix(gateway): align auth mode preflight
* fix(gateway): keep preflight state reads isolated
Share the read-only inspection snapshot scope across duplicated runtime chunks so blocked gateway preflight remains non-mutating when bundled provider artifacts read canonical state.
* fix(gateway): keep startup preflight passive
* fix(gateway): ignore inactive embedding owner shadows
* fix(gateway): preserve startup preflight parity
* fix(llama-cpp): keep cache inspection types private
* fix(gateway): close startup preflight parity gaps
* fix(gateway): handle uninitialized memory databases
* test(gateway): observe shell fallback portably
* fix(llama-cpp): normalize embedding model paths
* refactor(gateway): drop broad startup preflight surface
* fix(doctor): report missing managed local embedding setup
* style(memory): simplify setup enablement check
* fix(memory): keep diagnostic result type private
* fix(memory): inspect local setup with remote secret refs
* fix(memory): keep doctor index inspection immutable
* fix(memory): make readiness inspection owner-aware
* fix(doctor): mirror memory slot allowlist policy
* test(doctor): use canonical memory slot id
* fix(doctor): normalize memory provider ids
* fix(doctor): resolve external embedding readiness owner
* fix(plugins): keep embedding inspection result internal
* fix(doctor): isolate plugin state during lint
* fix(doctor): route lint metadata through snapshot
* fix(cli): keep doctor lint startup source-only
* fix(cli): keep doctor lint compile-cache free
* fix(doctor): keep local embedding readiness opt-in
* test(doctor): preserve plugin artifact roots during lint
* fix(doctor): refresh memory readiness registration
* test(doctor): type nullable provider policy mock
* fix(doctor): scope lint state snapshot to provider check
* fix(doctor): isolate selected plugin state checks
* test(doctor): restore only scoped environment
* fix(doctor): defer readiness state inspection
* fix(doctor): keep deferred config reads isolated
* fix(doctor): keep plugin state mode internal
* fix(config): preserve default plugin validation
2026-08-15 20:15:06 -06:00
Jason (Json)
a02ed2cfda
fix(plugin-sdk): installed llama providers fail to load after upgrade ( #124041 )
...
* fix(plugin-sdk): let installed llama providers load after upgrade
* test(tooling): include new memory race importer
2026-08-14 23:57:24 -06:00
Peter Steinberger
1348387076
refactor(plugins): replace node-llama-cpp with managed llama-server ( #123105 )
...
Move llama.cpp chat and local embeddings onto a verified externally managed llama-server runtime. Remove the in-process native runtime, forked embedding workers, and node-llama-cpp dependency while preserving guided setup, local GGUF models, tool-capable agent runs, diagnostics, and operator docs.
2026-08-13 16:58:20 -07:00
Vincent Koc
76ba214b98
fix(llama-cpp): dispose runtime on plugin stop
2026-08-04 16:10:51 +08:00
Peter Steinberger
1cc374b2f0
test(llama-cpp): consolidate provider fixtures ( #118426 )
2026-08-02 20:59:50 -07:00
Peter Steinberger
b521e0a6dc
fix(llama-cpp): stop hijacking explicitly configured HTTP provider routes ( #118293 )
...
* fix(llama-cpp): preserve configured HTTP routes
* chore: drop release-owned CHANGELOG edit from PR branch
2026-08-02 17:06:58 -07:00
Vincent Koc
bce957fe61
fix(llama-cpp): clarify local model setup
2026-08-02 18:04:30 +08:00
Peter Steinberger
e4d1b7e0d3
refactor(plugins): move plugin contributions into the registry bundle ( #117372 )
...
* refactor(plugins): move plugin contributions into the registry bundle
* fix(plugins): guard embedding owner union and drop unused test-util imports
* fix(plugins): break registry facade import cycles
* fix(plugins): remove obsolete registry snapshot seams
* test(gateway): preserve plugin runtime mock exports
* test(auto-reply): install builder registry in diagnostics fixture
* test(plugins): activate registry-backed capability fixtures
2026-08-01 06:57:01 -07:00
Peter Steinberger
98a066c90e
fix(agents): project llama.cpp-safe tool schemas ( #115598 )
...
Co-authored-by: Jithin Mohandas <mohandasjithin@gmail.com >
2026-07-29 01:33:32 -04:00
Peter Steinberger
658b601ee5
feat(llama-cpp): in-process local GGUF text inference provider ( #109444 )
...
* feat(llama-cpp): add in-process text inference
* test(llama-cpp): narrow setup provider fixture
* fix(llama-cpp): trim public surface and refresh docs map
* fix(llama-cpp): import Context type in inference test
2026-07-16 18:53:55 -07:00
Peter Steinberger
2eb9c7ebd7
refactor(extensions): privatize small plugin internals ( #107774 )
...
* refactor(llama-cpp): privatize embedding internals
* refactor(parallel): privatize MCP response helpers
* refactor(logbook): privatize analysis parsers
* refactor(crabbox): narrow worker provider exports
* refactor(acpx): privatize process reaper internals
* chore(deadcode): refresh unused-export baseline
2026-07-14 14:15:44 -07:00
Peter Steinberger
81941f2d68
test: enable noUncheckedIndexedAccess for extension tests ( #105343 )
2026-07-12 12:48:22 +01:00
Vincent Koc
85a96409f1
feat(memory): surface llama.cpp diagnostics
2026-07-11 16:40:14 +08:00
Vincent Koc
d3f7f7d1fc
chore(deadcode): remove unused test-only helpers
2026-06-22 15:48:43 +08:00
liuhao1024
94e6255666
feat(memory): apply outputDimensionality truncation to local GGUF embeddings ( fixes #58765 ) ( #93758 )
...
* feat(memory): apply outputDimensionality truncation to local GGUF embeddings
The outputDimensionality config field was passed through to the local
embedding provider but never applied. Local GGUF models (e.g.
Qwen3-Embedding-0.6B) always returned their full dimension vector.
Apply slice(0, N) after normalization so MRL-capable models can benefit
from dimension truncation — matching the behavior already supported by
Gemini embedding-2 and OpenAI providers.
Fixes #58765
* fix(memory): preserve local embedding dimensions through worker
---------
Co-authored-by: Vincent Koc <25068+vincentkoc@users.noreply.github.com >
2026-06-17 05:05:49 +08:00
mushuiyu_xydt
44e6caff54
fix(memory): accept local default model path migration ( #92954 )
...
* fix(memory): accept local default model path migration
Treat the official local default embedding model's hf URI and downloaded GGUF path identities as equivalent so upgraded local memory indexes do not pause solely on path-format changes.
* fix(memory): satisfy local identity lint
Avoid filtered array tail access in the local model filename helper while preserving the same compatibility behavior.
* fix(memory): preserve local embedding identity aliases
---------
Co-authored-by: Vincent Koc <25068+vincentkoc@users.noreply.github.com >
2026-06-15 09:29:42 +08:00
Onur Solmaz
3137110167
fix(memory): move local llama.cpp runtime to provider plugin
...
* fix(memory): move local llama.cpp runtime to provider plugin
* chore: ignore llama cpp dynamic dependency
* test: remove invalid local provider alias fixture
* chore: refresh llama cpp shrinkwrap
* chore: drop stale memory embedding defaults facade
2026-06-09 14:30:35 +08:00