Skip to content
#

llama-server

Here are 74 public repositories matching this topic...

A standalone, all-on-by-default token-optimization + memory layer for LLM requests: minifies tools/system/messages (code-fence aware), dedups tool results, distills old turns, prunes to a budget, type-compresses output; hot/warm/cold memory manager with atomic writes. Async proxy for Anthropic, OpenAI & Ollama + MCP server (stdio) + CLI. MIT.

  • Updated Sep 10, 2026
  • Python
dsh-local-llm-controller

为DSH接入本地大模型能力:在「设置→插件」页一键启停本地 llama.cpp 大模型(双槽x双模态x双预设),卡片内配置、一条命令安装、自动注册,装完即用| Enable local large model capabilities for DSH: One-click start/stop for local llama.cpp models (dual‑slot × dual‑modal × dual‑preset) right in Settings → Plugins; configure within the card, install with a single command, automatically register, and ready to use

  • Updated Sep 12, 2026
  • JavaScript

Add this topic to your repo

To associate your repository with the llama-server topic, visit your repo's landing page and select "manage topics."

Learn more