vllm-chat
Interactive chat client for a running vLLM OpenAI-compatible server
TLDR
SYNOPSIS
vllm chat [--url url] [--model-name name] [--api-key key] [--system-prompt text] [-q|--quick message] [--stats]
DESCRIPTION
vllm chat is the vllm subcommand that talks to an already running OpenAI-compatible HTTP server (vllm serve). It uses the official OpenAI Python client: chat.completions.create with stream=True. Reasoning deltas print inside <think> ... </think> when the server sends them.It does not load weights. Start vllm serve first, then vllm chat. vllm complete is the matching non-chat completions client.
PARAMETERS
--url url
OpenAI-compatible base URL. Default http://localhost:8000/v1 (the vllm serve default).--model-name name
Model id sent in chat completions. Default: first id from GET /v1/models.--api-key key
Bearer token for the client. Overrides OPENAI_API_KEY. If neither is set, the client uses EMPTY. Must match vllm serve --api-key / VLLM_API_KEY when the server requires a key.--system-prompt text
Optional system message inserted at the start of the conversation (for models that honor one).-q, --quick message
Send one user message, print the streamed reply, exit. Without -q, reads lines from the terminal (> prompt) until EOF.--stats
After each reply, print time to first token (ms) and tokens per second. Enables stream_options.include_usage on the request.
INSTALL
CAVEATS
Subcommand of vllm. Requires a reachable /v1 server; connection errors mean serve is not listening on --url. --api-key on this client only authenticates the OpenAI routes the CLI calls. On serve, --api-key still does not cover every HTTP path (see vllm). Interactive mode keeps the full transcript in memory for follow-ups. Default URL assumes port 8000.
SEE ALSO
vllm(1), ollama(1), llama.cpp(1), huggingface-cli(1)
