One API: Quickstart Guide
Integrate our uncensored LLM API with a single configuration change. Drop-in compatibility means you swap the base URL and API key, then keep your existing code.
Base URL & Authentication
Start by pointing your client to our dedicated endpoint. This openai compatible api uses standard authentication headers, so you only need to change two values in your configuration.
The base URL is https://api.onekeyllmapi.com/v1. Generate your unique API key on the signup page. No card is required to start, and the key is shown immediately upon registration. Keep this key secure; it authenticates all requests to the service.
Basic Completion
Send a simple text prompt to receive a direct response. This is the core of our ai api for developers offering. The model returns text without content refusals for lawful adult use, making it ideal for creative or unrestricted tasks.
- Model ID:
uncensored - Endpoint:
POST /v1/chat/completions - Context: 64,000 tokens
Make your first request with this example:
curl https://api.onekeyllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Use the official OpenAI Python library to interact with our service. The only changes needed are the base_url and api_key. This ensures your existing logic works without modification.
- Install via
pip install openai - Set environment variables or pass directly in the client init
- Call
chat.completions.create()as usual
Configure the client like this:
from openai import OpenAI
client = OpenAI(base_url="https://api.onekeyllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js SDK Usage
Integrate with JavaScript or TypeScript projects using the standard OpenAI Node package. Point the client to our URL and provide your key. This approach maintains compatibility with any OpenAI-compatible client structure.
- Install via
npm install openai - Initialize with custom base URL
- Use standard async/await patterns
Example initialization:
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.onekeyllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Receive responses token-by-token for lower perceived latency. Our streaming api supports Server-Sent Events (SSE), allowing you to display output as it generates. This is useful for chat interfaces or real-time applications.
- Set
stream: truein your request - Parse the SSE stream on the client
- Handle partial tokens efficiently
Enable streaming in your code:
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors & Context
Understand the constraints to build robust integrations. The API supports a 64k context api window, allowing long documents or conversations. Requests are limited to 8 MB body size.
Common errors include 401 for invalid keys, 402 when prepaid credit is exhausted, and 429 for exceeding 300 requests per minute. Pay-as-you-go credit never expires, and you can top up from $10. Tool calling is supported for function execution.
Capabilities and limits
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Spec | Value |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Model ID | uncensored |
| Base URL | https://api.onekeyllmapi.com/v1 |
| Authentication | Bearer token in the Authorization header |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Context window | 64,000 tokens, input and output combined |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Structured output | JSON object mode via response_format json_object |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Request size | up to 8 MB per request |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Requests per minute | 300/min per key |
| Parallel requests | 8 requests at the same time per key |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Subscription | no monthly fee; paid credit does not expire |
| Bonus credit | +5% from $50, +10% from $100 |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Keys | one active key per account; a new key replaces the old one |
| Sign-in | sign in with Google or with e-mail + password |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
HTTP errors
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is this an OpenAI model?
No. We serve a single, dedicated uncensored large language model run on our own GPU servers. It is not GPT, Claude, or any other vendor's model, but it is compatible with their API structure.
How does pricing work?
We use a pay-as-you-go model with no monthly fees. Input tokens cost $0.25 per 1M, and output tokens cost $1.00 per 1M. Prepaid credit never expires, and you can start with a $0.50 trial credit.
What is the context window size?
The API supports a 64,000 token context window for both prompt and completion. This allows for extensive conversations or long document processing within a single request.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.