Kimi K3 inference API

Kimi K3, served fast, on NVIDIA B300.

The full 2.8T-parameter open-weights model in its native MXFP4 precision, with speculative decoding, behind an OpenAI-compatible endpoint. Pay per token. No minimums, no sales call.

Output speed

~200 tok/s

single stream, 10k-token prompt, measured on our B300 node

Time to first token

~0.5 s

10k-token prompt, in-region

Context window

1M tokens

native Kimi K3, 1,048,576

Drop-in for the OpenAI SDK

Base URL https://k3.cybertino.io/v1, model id kimi-k3. Streaming, tool calling, vision input and the reasoning_effort knob (low / high / max) all work as in the official API.

curl https://k3.cybertino.io/v1/chat/completions \
  -H "Authorization: Bearer $LOGITS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "stream": true,
    "reasoning_effort": "max",
    "messages": [{"role": "user", "content": "Explain speculative decoding in two sentences."}]
  }'

Original weights

Native MXFP4 checkpoint from Moonshot, no extra quantization, Kimi K3 tool-call and reasoning parsers.

Simple pricing

$2.40 per 1M input tokens, $12.00 per 1M output tokens. Reasoning tokens are billed as output.

Dedicated hardware

8x NVIDIA B300 per replica, tensor parallel, speculative decoding tuned for interactive use.