Auto-discovers models from OpenAI-compatible endpoints and populates opencode's model list. Works with llama.cpp, Ollama, LM Studio, and anything else that uses the /v1/models API.
opencode doesn't auto-discover models from local endpoints. You have to manually define each model in your config. This plugin fixes that by hitting your endpoint's /models endpoint and loading whatever's available.
- Discovers models from your OpenAI-compatible endpoint, respecting any you've already configured manually
- Beautifies ugly model IDs into readable names (
qwen2-5-7b-instruct->Qwen 2.5 7B Instruct) - Detects context and output limits from standard fields (
context_length,max_model_len,max_completion_tokens, etc.) - Falls back to
min(context/4, 32000)for output limits when undetectable - Supports auth via
apiKeyin provider options - Maps llama-router's API addon report: reasoning effort variants, the advertised default effort, a thinking-off variant,
supported_parameters, and fullpricing - Works with any
@ai-sdk/openai-compatibleprovider in your config
Add to your opencode.json:
{
"plugin": [
"HarutoHiroki/OpenCode-AutoDiscovery#main"
]
}{
"plugin": [
"/path/to/OpenCode-AutoDiscovery"
]
}No config needed. Just make sure your OpenAI-compatible provider is defined in your opencode.json:
{
"provider": {
"my-local-llm": {
"npm": "@ai-sdk/openai-compatible",
"options": {
"baseURL": "http://localhost:8080/v1"
}
}
}
}For endpoints that require auth, add apiKey:
{
"provider": {
"my-gateway": {
"npm": "@ai-sdk/openai-compatible",
"options": {
"baseURL": "https://gateway.example.com/v1",
"apiKey": "your-api-key"
}
}
}
}The plugin will automatically discover models for any provider using @ai-sdk/openai-compatible. If you've already manually configured some models on a provider, only the remaining ones are discovered.
On opencode startup, the plugin:
- Scans your providers for any using
@ai-sdk/openai-compatible - Hits each provider's
/v1/modelsendpoint (with Bearer auth ifapiKeyis set) - Reads context/output limits from standard response fields (
context_length,max_model_len,max_completion_tokens,max_output_tokens,meta.n_ctx) - Skips models you've already configured manually
- Beautifies model IDs into human-readable names
- Registers discovered models with opencode
llama-router rewrites its /v1/models report through an API addon (src/api_addon.py) into an OpenRouter-style shape. The plugin maps that report onto opencode's model config:
reasoning_effort.levels-> one variant per level, each sendingreasoning_effort: <level>on the wirereasoning_effort.default-> the model's defaultreasoningEffort, so "no variant" matches the router's advertised defaultreasoning_effort.disable-> anonevariant that turns thinking off:nonesendsreasoning_effort: "none",lowestsends the lowest level,qwensendschat_template_kwargs: { "enable_thinking": false }supported_parameters->tool_callfrom the presence oftoolspricing->cost, includingcache_readandcache_writearchitecture.input_modalities/output_modalities-> input/output modalitiescontext_length/max_output_tokens->limit.context/limit.output
Models with a reasoning surface also get reasoning: true so opencode advertises the capability.
The plugin has built-in knowledge of common model prefixes and tags:
qwen2-5-7b-instruct->Qwen 2.5 7B Instructgemma-3-27b-it->Gemma 3 27B Instructglm-4-9b-chat->GLM 4 9B Chatkimi-k2-0905-preview->Kimi K2 0905 Preview
Unknown parts are capitalized and joined with spaces.
MIT, optionally credit this if you implement the code in your codebase, would be appreciated.