Back to SMOX.AI

Developer documentation · September 2026

Mixvoip
Reasoning API

Connect an OpenAI-compatible client to locally hosted chat, embedding and transcription models through one authenticated API.

API connection
Base URL
https://llm.mixvoip.ai/v1
Authentication
Bearer API key
Content type
application/json
Download developer manual

01

Connect your client

Use the API key supplied by Mixvoip. Keep the real key on your server or in a private secret manager. Never expose it in browser code, a repository or application logs.

Base URL

llm.mixvoip.ai/v1

Authentication

Bearer token

Transport

HTTPS + JSON

Bash / Zsh
export MIXVOIP_API_KEY="YOUR_API_KEY"
export MIXVOIP_BASE_URL="https://llm.mixvoip.ai/v1"
export MIXVOIP_MODEL="Qwen/Qwen3.8-27B-FP8"

02

Choose an available model

List the models exposed by your API key before integrating. The supplied key described in the September 2026 manual exposes three model types.

List models
curl --fail-with-body "$MIXVOIP_BASE_URL/models" \
  -H "Authorization: Bearer $MIXVOIP_API_KEY"
Model IDCapabilityUse with chat endpoint
Qwen/Qwen3.8-27B-FP8Chat and reasoningYes
qwen3-embedding-4bEmbeddingsNo
whisper-large-v3TranscriptionNo

03

Send your first prompt

Call the OpenAI-compatible chat completions endpoint. Read the completed answer from choices[0].message.content.

cURL
curl --fail-with-body "$MIXVOIP_BASE_URL/chat/completions" \
  -H "Authorization: Bearer $MIXVOIP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-27B-FP8",
    "messages": [{
      "role": "user",
      "content": "Explain DNS in two sentences."
    }]
  }'

04

Python quickstart

Install the official OpenAI Python SDK and provide the Mixvoip base URL. The SDK then sends requests to the Mixvoip service rather than the default OpenAI endpoint.

Install SDK
python -m pip install openai
Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["MIXVOIP_API_KEY"],
    base_url=os.environ["MIXVOIP_BASE_URL"],
    timeout=120.0,
    max_retries=2,
)

response = client.chat.completions.create(
    model=os.environ["MIXVOIP_MODEL"],
    messages=[{
        "role": "user",
        "content": "Explain DNS in two sentences.",
    }],
)

print(response.choices[0].message.content or "")

Timeouts and retries

The 120-second timeout is an example client setting, not a service guarantee. Adjust it to the workload. SDK retries can increase total latency and can create another billable request.

05

Stream the answer

Streaming is verified for the chat model above. Set stream=True and append the content delta from each returned chunk.

Python streaming
stream = client.chat.completions.create(
    model=os.environ["MIXVOIP_MODEL"],
    messages=[{"role": "user", "content": "What is DNS?"}],
    stream=True,
)

try:
    for chunk in stream:
        if chunk.choices:
            text = chunk.choices[0].delta.content or ""
            print(text, end="", flush=True)
finally:
    stream.close()
print()

Conversation history: include earlier user and assistant messages in the messages array on each request. Keep the full history within the selected model's context limit.

06

Reasoning controls and limits

Start with a basic request. Add optional model-specific controls only after confirming that the selected model supports them.

Verified optional fieldreasoning_effort="medium"

Accepted for Qwen/Qwen3.8-27B-FP8. Acceptance alone does not confirm a behavioural change.

Reported context limit131,072 input tokens

The model catalogue reports this limit. It was not stress-tested in the supplied verification.

Test requests used max_tokens=256 successfully. Non-streaming responses returned prompt, completion and total token usage. Omit sampling controls unless needed and supported.

07

Handle common failures

Treat authentication, request validation, limits and temporary service failures differently. Do not retry every error automatically.

400 / 422Invalid request

Check the model ID, message structure and optional parameters. Remove unsupported controls before retrying.

401 / 403Authentication or access

Check the key and Bearer header. Confirm that the key is active and permitted to use the selected model.

404Route or model not found

Use the SDK base URL ending in /v1 without /chat/completions. Check the exact model ID against the model list.

429Rate or quota limit

Respect Retry-After when present. Back off for rate limits and contact Mixvoip if quota is exhausted.

5xx / timeoutTemporary failure

Use bounded retries with exponential backoff and jitter. A retry can create another billable request.

08

Before you deploy

Move from a successful test request to a production integration with explicit security, reliability and observability choices.

Keep API keys on the server and out of browser code.
Do not commit keys to repositories or print them in logs.
Record status, model, latency and request ID when available.
Inspect finish_reason for truncated responses.
Track returned token usage and apply quota controls.
Confirm current model limits and pricing with Mixvoip.

Need an API key?

Request access and describe your intended workload.

Request API access