Switch your client in three lines
Drop your Kimi API key into any OpenAI-compatible client and start generating uncensored text. This guide shows you exactly how to configure your SDK and understand the real token pricing.
Install OpenAI Python/Node SDK
Start by installing the standard OpenAI client library. This works because our API follows the same OpenAI-compatible interface. You do not need a special vendor SDK. Just run pip install openai for Python or npm install openai for Node.js. The library handles the JSON serialization and HTTP requests automatically.
Configure Base URL and API Key
Point your client to our endpoint. Set the base URL to https://api.kimiapi.top/v1 and provide your API key from the dashboard. This key authenticates your prepaid credit. You can generate a new key instantly via Google or email on the signup page. No credit card is required to start.
curl https://api.kimiapi.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Send Your First Chat Completion
Send a simple message to the uncensored model. This model is tuned to answer without refusals for lawful adult use. It supports a 64,000-token context window, though max output is 16,000 tokens. The response returns text directly. Errors and refusals do not consume your credit.
from openai import OpenAI
client = OpenAI(base_url="https://api.kimiapi.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Enable Streaming with SSE
For real-time responses, enable streaming. The SDK streams tokens via Server-Sent Events (SSE). The final chunk contains the total token usage for billing. This approach reduces perceived latency for long outputs. You can process each token as it arrives in your application logic.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Use JSON Mode for Structured Output
Ensure consistent machine-readable output by setting response_format to {"type": "json_object"}. This forces the model to return valid JSON. It is ideal for data extraction or configuration generation. The model will not break structure, making it easier to parse responses in your code.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.kimiapi.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Check Usage and Costs
Monitor your usage via the dashboard. Input costs are $0.25 per million tokens; output is $1.00 per million tokens. If you receive a 401, check your key. A 402 means you need to top up with USDT or USDC. A 429 indicates you hit the 300 requests per minute limit. Credit never expires, so top up when convenient.
Questions and answers
How does billing work for the Kimi API?
You pay only for real token usage. Input tokens cost $0.25 per million, and output tokens cost $1.00 per million. Errors and refusals are free. Credit is prepaid and never expires.
What crypto payments are accepted?
We accept USDT (TRC20) and USDC (Base). Top-ups range from $10 to $500. Payments are processed directly without cards or PayPal.
What is the context window limit?
The model supports a 64,000-token context window for both prompt and completion. However, the maximum output per request is 16,000 tokens unless you specify a lower limit.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.