Quickstart Guide for Uncensored LLM API
This quickstart guide helps you integrate our uncensored chat-completions API using standard OpenAI-compatible clients. You will find the base URL, authentication steps, code examples, and limits needed to start sending requests immediately.
Authentication & Headers
Our API uses standard API key authentication. You must include your key in the Authorization header for every request. New accounts receive a key immediately after signup on the Get API key page. If you lose your key, you can regenerate it at any time, which revokes the old one. Only one active key is allowed per account.
The base URL for all requests is https://api.llmapistatus.com/v1. This works with official OpenAI SDKs and any OpenAI-compatible client by simply updating the base_url configuration. There are no special headers required beyond authentication.
Chat Completions Endpoint
Send text in and get text out using the POST /v1/chat/completions endpoint. The model ID is always uncensored. This is an open-weight model run on our own GPU servers, tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model.
The API supports tool and function calling. You do not need to manage routing between different vendors; we serve a single, dedicated uncensored endpoint. For general comparisons of other providers, see our LLM Leaderboard.
curl https://api.llmapistatus.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Streaming Responses
For lower latency, you can enable streaming by setting stream: true in your request body. The API returns a Server-Sent Events (SSE) stream. This is ideal for chat interfaces where you want to display tokens as they are generated.
Streaming works with both the official OpenAI SDKs and generic HTTP clients. The stream continues until the model finishes or an error occurs. This method reduces the perceived wait time for users, making the uncensored model feel more responsive in production applications.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling Support
The API supports structured output via function calling. You can define tools in your request and the model will return structured JSON that matches your schema. This is useful for building agents or applications that need to interact with other systems.
Use the tools parameter in your chat completion request. The model will respond with tool calls when appropriate, allowing you to execute functions and feed the results back into the conversation. This feature is available on all requests without extra cost.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmapistatus.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Model Listing & Context Window
You can list available models using the GET /v1/models endpoint. This will return metadata including the model ID, which is uncensored. The model supports a context window of 100,000 tokens for both prompt and completion combined.
This large context window allows you to pass extensive system instructions, conversation history, or document content in a single request. It is suitable for long-form generation or processing large documents. Keep your total token count within the 100k limit to avoid errors.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmapistatus.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Rate Limits & Quotas
Each API key is limited to 300 requests per minute. The maximum request body size is 8 MB. If you exceed the rate limit, you will receive a 429 status code. If your API key is invalid, you will get a 401 error. If you have no prepaid credit, you will receive a 402 error.
Pricing is transparent: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credit is prepaid and never expires. You can top up from $10 by crypto (USDT or USDC). See the LLM API Pricing page for details on bonus credits.
API specifications
Use this table to decide whether the API fits your project before you buy credit.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Methods | POST /v1/chat/completions · GET /v1/models |
| API key | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.llmapistatus.com/v1 |
| JSON mode | response_format: {"type": "json_object"} |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Max context | 100,000 tokens (prompt + completion together) |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Streaming | Supported (stream: true), usage included at the end |
| Concurrency | 8 requests at the same time per key |
| Rate limit | 300 requests per minute per key |
| Request size | up to 8 MB per request |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Free trial | $0.50 for 7 days, no card |
| Volume bonus | +5% from $50, +10% from $100 |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Subscription | no monthly fee; paid credit does not expire |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Sign-in | Google or e-mail and password |
| Key management | one active key per account; a new key replaces the old one |
HTTP errors
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is this model the same as GPT-4 or Claude?
No. The model ID is "uncensored" and it is an open-weight model run on our own GPU servers. It is tuned for lawful adult use without content refusals but is a distinct model from GPT, Claude, Gemini, or other vendors.
How do I get started with trial credit?
Sign up on the Get API key page with just an email and password. You will receive $0.50 of trial credit valid for 7 days. No credit card or phone number is required to start.
What happens if I hit the rate limit?
If you exceed 300 requests per minute, the API returns a 429 status code. You should implement exponential backoff in your client. You can regenerate your key if needed, but the rate limit applies per key.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.