Kimi K3 inference API
Kimi K3, served fast, on NVIDIA B300.
The full 2.8T-parameter open-weights model in its native MXFP4 precision, with speculative decoding, behind an OpenAI-compatible endpoint. Pay per token. No minimums, no sales call.
Output speed
~200 tok/s
single stream, 10k-token prompt, measured on our B300 node
Time to first token
~0.5 s
10k-token prompt, in-region
Context window
1M tokens
native Kimi K3, 1,048,576
Drop-in for the OpenAI SDK
Base URL https://k3.cybertino.io/v1, model id kimi-k3. Streaming, tool calling, vision input and the reasoning_effort knob (low / high / max) all work as in the official API.
curl https://k3.cybertino.io/v1/chat/completions \
-H "Authorization: Bearer $LOGITS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"stream": true,
"reasoning_effort": "max",
"messages": [{"role": "user", "content": "Explain speculative decoding in two sentences."}]
}'Original weights
Native MXFP4 checkpoint from Moonshot, no extra quantization, Kimi K3 tool-call and reasoning parsers.
Simple pricing
$2.40 per 1M input tokens, $12.00 per 1M output tokens. Reasoning tokens are billed as output.
Dedicated hardware
8x NVIDIA B300 per replica, tensor parallel, speculative decoding tuned for interactive use.