OpenAI-compatible Servers
Connect vLLM, llama.cpp, LM Studio, SGLang, or any other server with the OpenAI /v1 API. Add it as the OpenAI-compatible provider. Archestra uses its Chat Completions and Embeddings endpoints.
Connecting a Server
Start your model server before adding it. Archestra must be able to reach its URL; localhost inside a container refers to that container.
- Go to Model Providers → Add API Key and select OpenAI-compatible.
- Set Base URL to the server's OpenAI API root, for example
http://model-server:8000/v1. - Enter an API key if your server requires one. Otherwise leave it blank.
- Click Test & Create.
The key appears in the provider table, and the models returned by your server appear under Models. A server that lists several models needs only one key entry. Add separate entries for servers on different URLs.
Proxy clients use https://<archestra-host>/v1/vllm. Through the Model Router, use model IDs prefixed with vllm:.
Deployment Configuration
Set ARCHESTRA_VLLM_BASE_URL to configure a default endpoint. A Base URL set on a provider key takes precedence. ARCHESTRA_CHAT_VLLM_API_KEY supplies its default credential when needed.
The API key variable alone does not create a provider key at startup: the Base URL variable must also be set. If connection testing fails, check that the backend can reach the server and that /models is available under the configured API root.