AIREITER
DOCS APIPRECIOS
PLANTILLAS
MoonshotText Chat

Kimi K3 AI Chat Playground y API

Prueba Kimi K3 en línea para bases de código de contexto largo, colecciones de investigación, revisión de documentos y memoria de agentes a través de una API de Chat Completions compatible con OpenAI.

EntradaOficial $3.00 por 1 M de tokensAIReiter $1.50 por 1 M de tokensSalidaOficial $15.00 por 1 M de tokensAIReiter $7.50 por 1 M de tokensLectura de cachéOficial $0.30 por 1 M de tokensAIReiter $0.15 por 1 M de tokens
Ejecutar con API
PlaygroundReadmeAPI

ENTRADA

1
2
3
4
5
6
7
8
9
10

Install the official OpenAI client — AIReiter speaks the same protocol, so only the base URL changes:

npm install openai

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AIREITER_API_KEY,
  baseURL: "https://aireiter.com/api/v1",
});

Run kimi-k3:

const response = await client.chat.completions.create({
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  });

console.log(response);

Stream the response instead:

const stream = await client.chat.completions.create({
  ...{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  },
  stream: true,
});

for await (const event of stream) {
  console.log(event);
}

Install the official OpenAI client — AIReiter speaks the same protocol, so only the base URL changes:

pip install openai

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIREITER_API_KEY"],
    base_url="https://aireiter.com/api/v1",
)

Run kimi-k3:

response = client.chat.completions.create(
      model = "kimi-k3",
      messages = [
        {
          role = "user",
          content = "Send a message"
        }
      ],
      max_tokens = 4096
)

print(response)

Stream the response instead:

stream = client.chat.completions.create(
      model = "kimi-k3",
      messages = [
        {
          role = "user",
          content = "Send a message"
        }
      ],
      max_tokens = 4096,
    stream=True,
)

for event in stream:
    print(event)

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Run kimi-k3 against AIReiter's API:

curl -s -X POST \
  -H "Authorization: Bearer $AIREITER_API_KEY" \
  -H "Content-Type: application/json" \
  "https://aireiter.com/api/v1/chat/completions" \
  -d '{
  "model": "kimi-k3",
  "messages": [
    {
      "role": "user",
      "content": "Send a message"
    }
  ],
  "max_tokens": 4096
}'

Add "stream": true to the body to receive the response as server-sent events.

SALIDA

Example

A codebase you don't know is a risk, not a hurdle. The plan should burn down that risk in order: understand, run, trace, spike, then make the change small enough that "production-ready" is a property you can demonstrate, not a hope.

Here are the first five steps, with the evidence you collect at each.

Step 1 — Get it building and the baseline green (half day) Read the README, the manifest(s), the CI config, and any architecture docs. Boot the app locally, run the full test suite, run the linter/typechecker. Evidence: The exact commands that reproduce a clean build and a passing test suite from a fresh checkout (recorded output, not memory of it). A one-page map you wrote yourself: where the entry points are, the layers, the build/test toolchain, and the main risks to your feature (e.g. "this touches cron jobs and payments").

Step 2 — Trace the smallest end-to-end slice that resembles your feature (half day) Don't start with your feature. Pick a neighboring one that already works and follow it from entry to persistence. Note the conventions along the way: how errors surface, how config is injected, how logging is done, how tests are written. Evidence: An annotated trace — file → function → what it does — for that slice. A written list of "conventions I must follow" (not intuitions; things you observed in real code). A named shortlist of the 3–5 files you will actually touch. If you can't produce this trace, you're not ready to write code.

Step 3 — Spike the core path (day 2) Build the ugly version: hardcoded values, no error handling, no tests. The purpose is to confirm the path you traced in step 2 is real and to surface what you didn't know you didn't know. Evidence: A working spike demonstrating the feature's central data path, alongside a list of every assumption the spike broke and what you corrected. That correction list is the most valuable document in this whole plan.

Step 4 — Write the contract before the code (half of day 2 / day 3) Once the spike proves the path, pin down what production needs: the inputs/outputs, the error cases, where it sits in the conventions from step 2. Then write the tests — they'll be red, but they're the specification. Evidence: A one-to-two-page design doc, an agreed interface/API shape (with the team if there is one), and a red test suite that encodes intended behavior. If you can't write the contract without consulting the code, you haven't finished step 2.

Step 5 — Implement in small, verified increments (days 3–4) Replace the spike with the real thing in small commits, each one keeping the suite green, leaning on the existing patterns. Run lint/tests/typecheck per commit — CI, not just locally — and exercise the actual path against a real instance (staging or a local environment that isn't stubbed). Evidence: A branch with progressive commits, each green in CI; coverage on the new code; something that proves it works against reality (a test result, a log trace, a screenshot); and a review by at least one person who knows the codebase. The review counts as evidence — an unfamiliar codebase has tribal knowledge you cannot extract from the files alone.

Steps 6+ would be the things that actually make it "shipped": a migration plan and its rollback, feature flagging, observability, the release and post-release verification. But the first five get you to a reviewed, green, working slice in staging — which is the point at which you can say "this will work in production" with evidence behind it, instead of a guess.

{
  "model": "kimi-k3",
  "input": {
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Send a message"
      }
    ],
    "max_tokens": 4096
  },
  "output": "A codebase you don't know is a risk, not a hurdle. The plan should burn down that risk in order: understand, run, trace, spike, then make the change small enough that \"production-ready\" is a property you can demonstrate, not a hope.\n\nHere are the first five steps, with the evidence you collect at each.\n\n**Step 1 — Get it building and the baseline green (half day)**\nRead the README, the manifest(s), the CI config, and any architecture docs. Boot the app locally, run the full test suite, run the linter/typechecker.\n*Evidence:* The exact commands that reproduce a clean build and a passing test suite from a fresh checkout (recorded output, not memory of it). A one-page map you wrote yourself: where the entry points are, the layers, the build/test toolchain, and the main risks to your feature (e.g. \"this touches cron jobs and payments\").\n\n**Step 2 — Trace the smallest end-to-end slice that resembles your feature (half day)**\nDon't start with your feature. Pick a neighboring one that already works and follow it from entry to persistence. Note the conventions along the way: how errors surface, how config is injected, how logging is done, how tests are written.\n*Evidence:* An annotated trace — file → function → what it does — for that slice. A written list of \"conventions I must follow\" (not intuitions; things you observed in real code). A named shortlist of the 3–5 files you will actually touch. If you can't produce this trace, you're not ready to write code.\n\n**Step 3 — Spike the core path (day 2)**\nBuild the ugly version: hardcoded values, no error handling, no tests. The purpose is to confirm the path you traced in step 2 is real and to surface what you didn't know you didn't know.\n*Evidence:* A working spike demonstrating the feature's central data path, alongside a list of every assumption the spike broke and what you corrected. That correction list is the most valuable document in this whole plan.\n\n**Step 4 — Write the contract before the code (half of day 2 / day 3)**\nOnce the spike proves the path, pin down what production needs: the inputs/outputs, the error cases, where it sits in the conventions from step 2. Then write the tests — they'll be red, but they're the specification.\n*Evidence:* A one-to-two-page design doc, an agreed interface/API shape (with the team if there is one), and a red test suite that encodes intended behavior. If you can't write the contract without consulting the code, you haven't finished step 2.\n\n**Step 5 — Implement in small, verified increments (days 3–4)**\nReplace the spike with the real thing in small commits, each one keeping the suite green, leaning on the existing patterns. Run lint/tests/typecheck per commit — CI, not just locally — and exercise the actual path against a real instance (staging or a local environment that isn't stubbed).\n*Evidence:* A branch with progressive commits, each green in CI; coverage on the new code; something that proves it works against reality (a test result, a log trace, a screenshot); and a review by at least one person who knows the codebase. The review counts as evidence — an unfamiliar codebase has tribal knowledge you cannot extract from the files alone.\n\nSteps 6+ would be the things that actually make it \"shipped\": a migration plan and its rollback, feature flagging, observability, the release and post-release verification. But the first five get you to a reviewed, green, working slice in staging — which is the point at which you can say \"this will work in production\" with evidence behind it, instead of a guess.",
  "metrics": {
    "input_tokens": 134,
    "output_tokens": 2354,
    "generated_in_seconds": 42.7
  },
  "example": true
}
Generated in
42.7 seconds
Token de entrada
134
Token de salida
2354
Tokens per second
55.13 tokens / second
Time to first token
-

Detalles del modelo

Usa la misma clave de modelo en Playground, solicitudes API y flujos de trabajo internos.

ID del modelo
kimi-k3
Proveedor
Moonshot
Protocolo
OpenAI Chat Completions
Ventana de contexto
1,048,576 tokens
Salida máxima
131,072 tokens
Token de entrada
150 créditos / 1 M de tokens
Token de salida
750 créditos / 1 M de tokens
Lectura de caché
15 créditos / 1 M de tokens
Escritura de caché
-

Lo que puedes hacer con Kimi K3

Elige Kimi K3 cuando el contexto sea el cuello de botella y una sola solicitud necesite incluir un gran conjunto de trabajo de código, documentos, evidencia o historial del agente.

Revisión de contexto largo

Mantén juntas en un solo contexto de trabajo grandes bases de código, colecciones de documentos o evidencia de investigación.

Análisis de repositorios

Rastrea relaciones entre archivos y analiza cambios con más estado del proyecto disponible.

Síntesis de investigación

Compara afirmaciones en muchas notas y fuentes antes de producir una conclusión estructurada.

Evaluación de memoria de agentes

Revisa largos rastros de herramientas y decisiones previas para encontrar dónde salió mal un flujo de trabajo automatizado.

Casos de uso de Kimi K3

Ideal para flujos de trabajo en los que preservar más evidencia en el prompt puede evitar el troceado prematuro, la recuperación o la pérdida del estado del proyecto.
01

Revisión de bases de código

Analiza más contexto del repositorio en una sola solicitud.

02

Conjuntos largos de documentos

Revisa contratos, políticas, informes o colecciones de investigación.

03

Análisis de trazas de agentes

Inspecciona historiales largos de herramientas y el estado retenido.

04

Prototipos con mucho contexto

Prueba si más contexto mejora los resultados antes de construir la recuperación.

Cómo usar Kimi K3

Prueba el modelo en tres pasos sencillos.

01

Elige tu configuración

Configura los controles de respuesta y las opciones de carga compatibles con el modelo.

02

Envía una instrucción

Describe la tarea, añade el contexto relevante y revisa la respuesta en streaming y el uso de tokens.

03

Conecta la API

Usa el endpoint documentado y tu clave de API para llevar el mismo modelo a tu producto.

Desarrolla con la API de Kimi K3

Pasa de una prueba interactiva a una integración de producción con controles predecibles e informes de uso.

Protocolos familiares

Usa el protocolo API configurado para este modelo, incluida la transmisión en tiempo real donde esté disponible.

Visibilidad del uso

Haz seguimiento de los tokens de entrada, los tokens de salida y los créditos consumidos después de cada respuesta.

Controles específicos del modelo

Pasa los parámetros de generación compatibles en lugar de depender de valores predeterminados genéricos.

Una sola cuenta y saldo

Prueba y utiliza los modelos de texto compatibles a través de la misma cuenta de AIReiter y el mismo sistema de facturación.

Preguntas frecuentes sobre Kimi K3

Preguntas comunes sobre el playground en línea, los precios y el acceso a la API.

/ 01

¿Para qué es mejor Kimi K3?

Úsalo cuando una solicitud necesite un gran conjunto de trabajo de código, documentos, evidencia de investigación o historial del agente.

/ 02

¿Qué ventana de contexto está disponible para Kimi K3?

AIReiter enumera Kimi K3 con una ventana de contexto de 1,048,576 tokens; valida los límites y tiempos de espera del cliente antes de enviar solicitudes muy grandes.

/ 03

¿Puedo llamar a Kimi K3 con un cliente al estilo OpenAI?

Sí. AIReiter lo expone a través de un endpoint de Chat Completions compatible con OpenAI.

/ 04

¿Cómo se fija el precio de Kimi K3?

AIReiter muestra las tarifas actuales de input, cache-read y output tokens; confírmelas antes de usarlas en producción.

/ 05

¿Cuándo debería elegir un modelo más pequeño en su lugar?

Use un modelo más ligero para solicitudes breves y sin estado que no se benefician de la capacidad de contexto largo de Kimi K3.

AIREITER

¿Preguntas? Contáctanos en
[email protected]

新速率有限公司NEWRATE LIMITED香港九龍花園街 2-16 號好景商業中心 2304 室Room 2304, Haojing Commercial Center, 2-16 Garden Street, Kowloon, Hong Kong

LLM

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1GLM-5.3 FlashGemini 3.6 Flash

Video IA

Gemini Omni 1.1 Flash ExtMiniMax H3Kling 3.0 Motion ControlKling 3.0 TurboKling 3.0

Imagen IA

GPT-Image 2.5Grok Imagine Image 2.0Midjourney V8.1Midjourney V7Z-Image Turbo

Blog

Ver todo →

Compañía

Política de privacidadTérminos de servicioPolítica de reembolso

© 2026 AIReiter. Todos los derechos reservados.