MIXERLEAD

LLM API Reference

The MixerLead API follows the OpenAI Chat Completions and Embeddings formats. Base URL: https://api.mixerlead.com/v1.

Quickstart

from openai import OpenAI

client = OpenAI(base_url="https://api.mixerlead.com/v1", api_key="ml_live_YOUR_KEY")
reply = client.chat.completions.create(
    model="mixerlead-ai/meta/llama-3.3-70b-instruct-fp8-fast",
    messages=[{"role": "user", "content": "Hello!"}],
)

Choose a Model

  • Chat (chatbots, writing, summaries): Gemma 4 26B, Llama 3.1 8B, Mistral Small 3.1, GLM-4.7 Flash.
  • Reasoning (maths, logic, planning): GPT-OSS 120B / 20B, Qwen3 30B, DeepSeek R1. Set max_tokens to 1000+.
  • Coding: Qwen2.5 Coder 32B, Kimi K2.6.
  • Vision (images as base64): Llama 4 Scout, Mistral Small 3.1.
  • Tool calling / agents: Llama 3.3 70B, GPT-OSS 120B.
  • Fast & low-cost: Granite 4.0 Micro, Llama 3.2 1B / 3B.
  • Embeddings (search, RAG): BGE-M3, Qwen3 Embedding, BGE Large EN.

Full list with credit multipliers: models page.

Authentication

Send your key as Authorization: Bearer ml_live_… or in the X-API-Key header.

Chat Completions

POST /v1/chat/completions with messages, max_tokens, temperature, top_p, stop, seed, response_format, tools and stream.

Streaming

Set stream: true for Server-Sent Events ending in data: [DONE]. Add stream_options: {"include_usage": true} for a final usage chunk.

Tool Calling

Models that support tools accept OpenAI-style tools and return tool_calls.

Embeddings

POST /v1/embeddings with input as a string or array. Default model: mixerlead-ai/baai/bge-large-en-v1.5; multilingual: mixerlead-ai/baai/bge-m3.

Usage & Credits

Credits = (prompt + completion tokens) × model credit multiplier. Each response includes X-MixerLead-Tokens-Billed and usage.mixerlead_billed_tokens.

Rate Limits & Quotas

Per-key requests per minute; HTTP 429 with Retry-After when a limit or quota is reached.

Errors

Errors use the OpenAI envelope {"error": {"message", "type", "param", "code"}}: 401 invalid_api_key, 402 subscription_expired, 403 endpoint_not_allowed, 404 model_not_found, 429 rate_limit_exceeded / insufficient_quota / upstream_capacity.