API reference
Endpoint
- Base URL:
https://k3.cybertino.io/v1(OpenAI-compatible) - Auth:
Authorization: Bearer <key> - Model id:
kimi-k3(also acceptsmoonshotai/Kimi-K3) - Routes:
/chat/completions,/completions,/models
Reasoning
Kimi K3 always thinks before answering. Control the depth with the top-level reasoning_effort field (low, high, max; default max). Thinking is returned inmessage.reasoning_content (streamed as delta.reasoning_content before any delta.content). For multi-turn and tool use, pass the full assistant message back, including reasoning_content and tool_calls.
from openai import OpenAI
client = OpenAI(base_url="https://k3.cybertino.io/v1", api_key="LOGITS_API_KEY")
stream = client.chat.completions.create(
model="kimi-k3",
stream=True,
reasoning_effort="max", # "low" | "high" | "max" (default "max")
messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
for chunk in stream:
d = chunk.choices[0].delta
if getattr(d, "reasoning_content", None):
print(d.reasoning_content, end="") # thinking tokens, streamed first
if d.content:
print(d.content, end="") # answer tokensTool calling and vision
Standard OpenAI tools / tool_choice; the response carries tool_calls. Image inputs use image_url content parts.
{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "Weather in Tokyo?"}],
"tools": [{"type": "function", "function": {"name": "get_weather",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
}Limits
- Context: 1,048,576 tokens. Output: up to 131,072 tokens per request (set
max_tokens). - Free tier: 60 requests/min, 4 concurrent, $5 credit. Over-limit requests return HTTP 429; exhausted credit returns HTTP 400 with a budget message.
- Streaming requests can run for up to 60 minutes.
Serving details
Weights: moonshotai/Kimi-K3, native MXFP4 (no additional quantization), FP8 KV cache. Engine: vLLM with DSpark speculative decoding, tensor parallel over 8x NVIDIA B300 per replica. Sampling defaults follow the model card;temperature, top_p and max_tokens are honored as sent.