API reference

Endpoint

Reasoning

Kimi K3 always thinks before answering. Control the depth with the top-level reasoning_effort field (low, high, max; default max). Thinking is returned inmessage.reasoning_content (streamed as delta.reasoning_content before any delta.content). For multi-turn and tool use, pass the full assistant message back, including reasoning_content and tool_calls.

from openai import OpenAI

client = OpenAI(base_url="https://k3.cybertino.io/v1", api_key="LOGITS_API_KEY")

stream = client.chat.completions.create(
    model="kimi-k3",
    stream=True,
    reasoning_effort="max",          # "low" | "high" | "max" (default "max")
    messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
for chunk in stream:
    d = chunk.choices[0].delta
    if getattr(d, "reasoning_content", None):
        print(d.reasoning_content, end="")   # thinking tokens, streamed first
    if d.content:
        print(d.content, end="")             # answer tokens

Tool calling and vision

Standard OpenAI tools / tool_choice; the response carries tool_calls. Image inputs use image_url content parts.

{
  "model": "kimi-k3",
  "messages": [{"role": "user", "content": "Weather in Tokyo?"}],
  "tools": [{"type": "function", "function": {"name": "get_weather",
             "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
}

Limits

Serving details

Weights: moonshotai/Kimi-K3, native MXFP4 (no additional quantization), FP8 KV cache. Engine: vLLM with DSpark speculative decoding, tensor parallel over 8x NVIDIA B300 per replica. Sampling defaults follow the model card;temperature, top_p and max_tokens are honored as sent.