AIREITER
API DOCSPRICING
TEMPLATES
MoonshotText Chat

Kimi K3 AI Chat Playground and API

Try Kimi K3 online for long-context codebases, research collections, document review, and agent memory through an OpenAI-compatible Chat Completions API.

InputOfficial $3.00 per 1M tokensAIReiter $1.50 per 1M tokensOutputOfficial $15.00 per 1M tokensAIReiter $7.50 per 1M tokensCache readOfficial $0.30 per 1M tokensAIReiter $0.15 per 1M tokens
Run with API
PlaygroundReadmeAPI

INPUT

1
2
3
4
5
6
7
8
9
10

Install the official OpenAI client — AIReiter speaks the same protocol, so only the base URL changes:

npm install openai

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AIREITER_API_KEY,
  baseURL: "https://aireiter.com/api/v1",
});

Run kimi-k3:

const response = await client.chat.completions.create({
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  });

console.log(response);

Stream the response instead:

const stream = await client.chat.completions.create({
  ...{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  },
  stream: true,
});

for await (const event of stream) {
  console.log(event);
}

Install the official OpenAI client — AIReiter speaks the same protocol, so only the base URL changes:

pip install openai

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIREITER_API_KEY"],
    base_url="https://aireiter.com/api/v1",
)

Run kimi-k3:

response = client.chat.completions.create(
      model = "kimi-k3",
      messages = [
        {
          role = "user",
          content = "Send a message"
        }
      ],
      max_tokens = 4096
)

print(response)

Stream the response instead:

stream = client.chat.completions.create(
      model = "kimi-k3",
      messages = [
        {
          role = "user",
          content = "Send a message"
        }
      ],
      max_tokens = 4096,
    stream=True,
)

for event in stream:
    print(event)

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Run kimi-k3 against AIReiter's API:

curl -s -X POST \
  -H "Authorization: Bearer $AIREITER_API_KEY" \
  -H "Content-Type: application/json" \
  "https://aireiter.com/api/v1/chat/completions" \
  -d '{
  "model": "kimi-k3",
  "messages": [
    {
      "role": "user",
      "content": "Send a message"
    }
  ],
  "max_tokens": 4096
}'

Add "stream": true to the body to receive the response as server-sent events.

OUTPUT

Example

A codebase you don't know is a risk, not a hurdle. The plan should burn down that risk in order: understand, run, trace, spike, then make the change small enough that "production-ready" is a property you can demonstrate, not a hope.

Here are the first five steps, with the evidence you collect at each.

Step 1 — Get it building and the baseline green (half day) Read the README, the manifest(s), the CI config, and any architecture docs. Boot the app locally, run the full test suite, run the linter/typechecker. Evidence: The exact commands that reproduce a clean build and a passing test suite from a fresh checkout (recorded output, not memory of it). A one-page map you wrote yourself: where the entry points are, the layers, the build/test toolchain, and the main risks to your feature (e.g. "this touches cron jobs and payments").

Step 2 — Trace the smallest end-to-end slice that resembles your feature (half day) Don't start with your feature. Pick a neighboring one that already works and follow it from entry to persistence. Note the conventions along the way: how errors surface, how config is injected, how logging is done, how tests are written. Evidence: An annotated trace — file → function → what it does — for that slice. A written list of "conventions I must follow" (not intuitions; things you observed in real code). A named shortlist of the 3–5 files you will actually touch. If you can't produce this trace, you're not ready to write code.

Step 3 — Spike the core path (day 2) Build the ugly version: hardcoded values, no error handling, no tests. The purpose is to confirm the path you traced in step 2 is real and to surface what you didn't know you didn't know. Evidence: A working spike demonstrating the feature's central data path, alongside a list of every assumption the spike broke and what you corrected. That correction list is the most valuable document in this whole plan.

Step 4 — Write the contract before the code (half of day 2 / day 3) Once the spike proves the path, pin down what production needs: the inputs/outputs, the error cases, where it sits in the conventions from step 2. Then write the tests — they'll be red, but they're the specification. Evidence: A one-to-two-page design doc, an agreed interface/API shape (with the team if there is one), and a red test suite that encodes intended behavior. If you can't write the contract without consulting the code, you haven't finished step 2.

Step 5 — Implement in small, verified increments (days 3–4) Replace the spike with the real thing in small commits, each one keeping the suite green, leaning on the existing patterns. Run lint/tests/typecheck per commit — CI, not just locally — and exercise the actual path against a real instance (staging or a local environment that isn't stubbed). Evidence: A branch with progressive commits, each green in CI; coverage on the new code; something that proves it works against reality (a test result, a log trace, a screenshot); and a review by at least one person who knows the codebase. The review counts as evidence — an unfamiliar codebase has tribal knowledge you cannot extract from the files alone.

Steps 6+ would be the things that actually make it "shipped": a migration plan and its rollback, feature flagging, observability, the release and post-release verification. But the first five get you to a reviewed, green, working slice in staging — which is the point at which you can say "this will work in production" with evidence behind it, instead of a guess.

{
  "model": "kimi-k3",
  "input": {
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  },
  "output": "A codebase you don't know is a risk, not a hurdle. The plan should burn down that risk in order: understand, run, trace, spike, then make the change small enough that \"production-ready\" is a property you can demonstrate, not a hope.\n\nHere are the first five steps, with the evidence you collect at each.\n\n**Step 1 — Get it building and the baseline green (half day)**\nRead the README, the manifest(s), the CI config, and any architecture docs. Boot the app locally, run the full test suite, run the linter/typechecker.\n*Evidence:* The exact commands that reproduce a clean build and a passing test suite from a fresh checkout (recorded output, not memory of it). A one-page map you wrote yourself: where the entry points are, the layers, the build/test toolchain, and the main risks to your feature (e.g. \"this touches cron jobs and payments\").\n\n**Step 2 — Trace the smallest end-to-end slice that resembles your feature (half day)**\nDon't start with your feature. Pick a neighboring one that already works and follow it from entry to persistence. Note the conventions along the way: how errors surface, how config is injected, how logging is done, how tests are written.\n*Evidence:* An annotated trace — file → function → what it does — for that slice. A written list of \"conventions I must follow\" (not intuitions; things you observed in real code). A named shortlist of the 3–5 files you will actually touch. If you can't produce this trace, you're not ready to write code.\n\n**Step 3 — Spike the core path (day 2)**\nBuild the ugly version: hardcoded values, no error handling, no tests. The purpose is to confirm the path you traced in step 2 is real and to surface what you didn't know you didn't know.\n*Evidence:* A working spike demonstrating the feature's central data path, alongside a list of every assumption the spike broke and what you corrected. That correction list is the most valuable document in this whole plan.\n\n**Step 4 — Write the contract before the code (half of day 2 / day 3)**\nOnce the spike proves the path, pin down what production needs: the inputs/outputs, the error cases, where it sits in the conventions from step 2. Then write the tests — they'll be red, but they're the specification.\n*Evidence:* A one-to-two-page design doc, an agreed interface/API shape (with the team if there is one), and a red test suite that encodes intended behavior. If you can't write the contract without consulting the code, you haven't finished step 2.\n\n**Step 5 — Implement in small, verified increments (days 3–4)**\nReplace the spike with the real thing in small commits, each one keeping the suite green, leaning on the existing patterns. Run lint/tests/typecheck per commit — CI, not just locally — and exercise the actual path against a real instance (staging or a local environment that isn't stubbed).\n*Evidence:* A branch with progressive commits, each green in CI; coverage on the new code; something that proves it works against reality (a test result, a log trace, a screenshot); and a review by at least one person who knows the codebase. The review counts as evidence — an unfamiliar codebase has tribal knowledge you cannot extract from the files alone.\n\nSteps 6+ would be the things that actually make it \"shipped\": a migration plan and its rollback, feature flagging, observability, the release and post-release verification. But the first five get you to a reviewed, green, working slice in staging — which is the point at which you can say \"this will work in production\" with evidence behind it, instead of a guess.",
  "metrics": {
    "input_tokens": 134,
    "output_tokens": 2354,
    "generated_in_seconds": 42.7
  },
  "example": true
}
Generated in
42.7 seconds
Input tokens
134
Output tokens
2354
Tokens per second
55.13 tokens / second
Time to first token
-

Model details

Use the same model key in Playground, API requests, and internal workflows.

Model ID
kimi-k3
Provider
Moonshot
Protocol
OpenAI Chat Completions
Context window
1,048,576 tokens
Max output
131,072 tokens
Input tokens
150 credits / 1M tokens
Output tokens
750 credits / 1M tokens
Cache read
15 credits / 1M tokens
Cache write
-

What You Can Do with Kimi K3

Choose Kimi K3 when context is the bottleneck and one request needs to carry a large working set of code, documents, evidence, or agent history.

Long-Context Review

Keep large codebases, document collections, or research evidence together in one working context.

Repository Analysis

Trace relationships across files and discuss changes with more of the project state available.

Research Synthesis

Compare claims across many notes and sources before producing a structured conclusion.

Agent Memory Evaluation

Review long tool traces and prior decisions to find where an automated workflow went wrong.

Kimi K3 Use Cases

Best suited to workflows where preserving more evidence in the prompt can avoid premature chunking, retrieval, or loss of project state.
01

Codebase Review

Analyze more repository context in a single request.

02

Long Document Sets

Review contracts, policies, reports, or research collections.

03

Agent Trace Analysis

Inspect long tool histories and retained state.

04

Context-Heavy Prototypes

Test whether more context improves results before building retrieval.

How to Use Kimi K3

Test the model in three straightforward steps.

01

Choose Your Settings

Set the response controls and upload options supported by the model.

02

Send a Prompt

Describe the task, add relevant context, and review the streamed response and token usage.

03

Connect the API

Use the documented endpoint and your API key to bring the same model into your product.

Build with the Kimi K3 API

Go from an interactive test to a production integration with predictable controls and usage reporting.

Familiar Protocols

Use the API protocol configured for this model, including streaming where available.

Usage Visibility

Track input tokens, output tokens, and consumed credits after each response.

Model-Specific Controls

Pass the supported generation parameters instead of relying on generic defaults.

One Account and Balance

Test and operate supported text models through the same AIReiter account and billing system.

Kimi K3 FAQ

Common questions about the online playground, pricing, and API access.

/ 01

What is Kimi K3 best for?

Use it when a request needs a large working set of code, documents, research evidence, or agent history.

/ 02

What context window is available for Kimi K3?

AIReiter lists Kimi K3 with a 1,048,576-token context window; validate client limits and timeouts before sending very large requests.

/ 03

Can I call Kimi K3 with an OpenAI-style client?

Yes. AIReiter exposes it through an OpenAI-compatible Chat Completions endpoint.

/ 04

How is Kimi K3 priced?

Current input, cache-read, and output token rates are displayed by AIReiter; confirm them before production use.

/ 05

When should I choose a smaller model instead?

Use a lighter model for short, stateless requests that do not benefit from Kimi K3 long-context capacity.

AIREITER

Questions? Contact us at
[email protected]

新速率有限公司NEWRATE LIMITED香港九龍花園街 2-16 號好景商業中心 2304 室Room 2304, Haojing Commercial Center, 2-16 Garden Street, Kowloon, Hong Kong

LLM

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1GLM-5.3 FlashGemini 3.6 Flash

AI Video

Gemini Omni 1.1 Flash ExtMiniMax H3Kling 3.0 Motion ControlKling 3.0 TurboKling 3.0

AI Image

GPT-Image 2.5Grok Imagine Image 2.0Midjourney V8.1Midjourney V7Z-Image Turbo

Blog

View All →

Company

Privacy PolicyTerms of ServiceRefund Policy

© 2026 AIReiter. All rights reserved.