mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
e3af600a90
* feat(deploy): vllm-litellm example — 3-model co-resident shape + HF loader Update the unified-memory inference example to the validated GB10 Spark shape: qwen3.6-27B-FP8 (reasoning) + gemma-4-12B-it (perception) + Qwen3-Reranker-4B, all co-resident on one GPU behind LiteLLM, loaded by HF id into a mounted HF_HOME cache. - qwen: MTP spec-decode + runai_streamer (weight load ~166s->1s) + full 256K at util 0.50 (default KV) - gemma on the OpenAI lane (audio), reranker direct on :8002/rerank - sequential startup + page-cache-drop guidance; runai_streamer kept on the big model only (its buffers break small models' KV budgets) - README: HF-id loader, DGX Spark (validated) + AMD Strix Halo (ROCm) setup, tuning notes, troubleshooting - wheel-check ALLOW entries for the example files (supersedes #687) * docs(deploy): clarify AMD edits are compose literals (Copilot review) In the Strix Halo guidance, --max-model-len and --load-format runai_streamer are hard-coded in docker-compose.yml's vllm-qwen command, not .env vars — say where to edit them.