01
Connect your client
Use the API key supplied by Mixvoip. Keep the real key on your server or in a private secret manager. Never expose it in browser code, a repository or application logs.
Base URL
llm.mixvoip.ai/v1
Authentication
Bearer token
Transport
HTTPS + JSON
export MIXVOIP_API_KEY="YOUR_API_KEY"
export MIXVOIP_BASE_URL="https://llm.mixvoip.ai/v1"
export MIXVOIP_MODEL="Qwen/Qwen3.8-27B-FP8"02
Choose an available model
List the models exposed by your API key before integrating. The supplied key described in the September 2026 manual exposes three model types.
curl --fail-with-body "$MIXVOIP_BASE_URL/models" \
-H "Authorization: Bearer $MIXVOIP_API_KEY"| Model ID | Capability | Use with chat endpoint |
|---|---|---|
| Qwen/Qwen3.8-27B-FP8 | Chat and reasoning | Yes |
| qwen3-embedding-4b | Embeddings | No |
| whisper-large-v3 | Transcription | No |
03
Send your first prompt
Call the OpenAI-compatible chat completions endpoint. Read the completed answer from choices[0].message.content.
curl --fail-with-body "$MIXVOIP_BASE_URL/chat/completions" \
-H "Authorization: Bearer $MIXVOIP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.8-27B-FP8",
"messages": [{
"role": "user",
"content": "Explain DNS in two sentences."
}]
}'04
Python quickstart
Install the official OpenAI Python SDK and provide the Mixvoip base URL. The SDK then sends requests to the Mixvoip service rather than the default OpenAI endpoint.
python -m pip install openaiimport os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MIXVOIP_API_KEY"],
base_url=os.environ["MIXVOIP_BASE_URL"],
timeout=120.0,
max_retries=2,
)
response = client.chat.completions.create(
model=os.environ["MIXVOIP_MODEL"],
messages=[{
"role": "user",
"content": "Explain DNS in two sentences.",
}],
)
print(response.choices[0].message.content or "")Timeouts and retries
The 120-second timeout is an example client setting, not a service guarantee. Adjust it to the workload. SDK retries can increase total latency and can create another billable request.
05
Stream the answer
Streaming is verified for the chat model above. Set stream=True and append the content delta from each returned chunk.
stream = client.chat.completions.create(
model=os.environ["MIXVOIP_MODEL"],
messages=[{"role": "user", "content": "What is DNS?"}],
stream=True,
)
try:
for chunk in stream:
if chunk.choices:
text = chunk.choices[0].delta.content or ""
print(text, end="", flush=True)
finally:
stream.close()
print()Conversation history: include earlier user and assistant messages in the messages array on each request. Keep the full history within the selected model's context limit.
06
Reasoning controls and limits
Start with a basic request. Add optional model-specific controls only after confirming that the selected model supports them.
Accepted for Qwen/Qwen3.8-27B-FP8. Acceptance alone does not confirm a behavioural change.
The model catalogue reports this limit. It was not stress-tested in the supplied verification.
Test requests used max_tokens=256 successfully. Non-streaming responses returned prompt, completion and total token usage. Omit sampling controls unless needed and supported.
07
Handle common failures
Treat authentication, request validation, limits and temporary service failures differently. Do not retry every error automatically.
400 / 422Invalid requestCheck the model ID, message structure and optional parameters. Remove unsupported controls before retrying.
401 / 403Authentication or accessCheck the key and Bearer header. Confirm that the key is active and permitted to use the selected model.
404Route or model not foundUse the SDK base URL ending in /v1 without /chat/completions. Check the exact model ID against the model list.
429Rate or quota limitRespect Retry-After when present. Back off for rate limits and contact Mixvoip if quota is exhausted.
5xx / timeoutTemporary failureUse bounded retries with exponential backoff and jitter. A retry can create another billable request.
08
Before you deploy
Move from a successful test request to a production integration with explicit security, reliability and observability choices.
Need an API key?
Request access and describe your intended workload.
