EmbeddingApis/Docs
Uncensored text API for your pipeline
Embeddings API Quickstart Guide
Use this guide to integrate our uncensored chat-completions API into your text processing pipelines. The endpoint accepts standard OpenAI formats, allowing you to generate text, summarize documents, or extract data without content refusals.
Authentication
Access the API using a unique API key generated from your account dashboard. Include this key in the Authorization header of every request using the Bearer scheme. The base URL for all endpoints is https://api.embeddingapis.com/v1. This URL works with the official OpenAI SDKs and any client configured for OpenAI-compatible services. Ensure your key is kept secure, as it grants access to your prepaid credits.
Each account is limited to one active key, which can be regenerated at any time. Regenerating the key immediately revokes the previous one, so update your configuration promptly.
The Chat Completions Endpoint
Send text prompts to the POST /v1/chat/completions endpoint. The API serves a single uncensored model identified by the ID uncensored. This model is an open-weight LLM tuned to answer without content refusals for lawful adult, fictional, or security-research topics. It processes text in and outputs text, making it ideal for post-processing, summarization, or extraction steps in your pipeline.
Structure your request with a messages array containing your system and user prompts. The API returns a JSON response containing the generated text.
curl https://api.embeddingapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Use the official OpenAI Python library to interact with the API. Set the base_url to our endpoint and provide your API key. This approach allows you to leverage existing client code with minimal changes. The client handles serialization and network requests, returning structured responses that you can parse for your application logic.
Ensure you are using a recent version of the SDK to support all features, including streaming and tool calling. The model ID remains uncensored when invoking completions.
from openai import OpenAI
client = OpenAI(base_url="https://api.embeddingapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK Integration
For JavaScript or TypeScript environments, use the OpenAI Node.js SDK. Configure the client with the correct base URL and API key. This setup ensures compatibility with your existing Node-based pipelines. The SDK manages the connection to the POST /v1/chat/completions endpoint, allowing you to send prompts and receive text completions efficiently.
This integration is suitable for server-side text processing, such as generating captions, scripts, or extracting data from large text blocks.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.embeddingapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
For low-latency applications, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This is particularly useful for real-time user interfaces or when processing long texts where immediate feedback is valuable.
Parse the SSE stream to handle partial responses. The uncensored model supports this mode, allowing you to react to content as it is produced without waiting for the entire completion.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors, and Context Window
The context window supports 100,000 tokens for both prompt and completion combined. Requests exceeding the 8 MB body limit or the rate limit of 300 requests per minute will fail. Authentication errors return a 401 status if the key is invalid. When prepaid credit is exhausted, requests fail because no credits remain to cover the token usage. A 429 status indicates rate limit exhaustion.
A hard content limit always applies: requests containing sexual content involving minors are blocked, regardless of other settings. All other lawful adult content is processed without refusal.
Capabilities and limits
The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.
| Item | Value |
|---|---|
| Protocol | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Authentication | Authorization: Bearer YOUR_KEY |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model ID | uncensored |
| Base URL | https://api.embeddingapis.com/v1 |
| Streaming | Supported (stream: true), usage included at the end |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Max output | up to 16,000 tokens per request (default 2,048) |
| Context window | 100,000 tokens (prompt + completion together) |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| JSON mode | JSON object mode via response_format json_object |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Concurrency | 8 requests at the same time per key |
| Rate limit | 300/min per key |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Subscription | paid credit never expires, no subscription |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Trial credit | $0.50 for 7 days, no card |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Account | Google or e-mail and password |
| Key management | one active key per account; a new key replaces the old one |
| Content | adult content allowed; sexual content involving minors is refused |
Error codes
Every error is JSON with a type you can switch on. You are never charged for an error.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Does this API support embeddings generation?
No. This API is a text chat-completions endpoint. It does not generate embeddings, images, audio, or video. Use it for text generation tasks like summarization or extraction.
How is pricing calculated?
Pricing is pay-as-you-go prepaid credit. Input tokens cost $0.25 per 1M tokens, and output tokens cost $1.00 per 1M tokens. Credits do not expire.
Is the model uncensored?
Yes, the model does not refuse lawful adult, fictional, or controversial topics. However, sexual content involving minors is always blocked.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.