Get API key

EmbeddingApisDocs

Uncensored text API for your pipeline

Embeddings API Quickstart Guide

Use this guide to integrate our uncensored chat-completions API into your text processing pipelines. The endpoint accepts standard OpenAI formats, allowing you to generate text, summarize documents, or extract data without content refusals.

Authentication

Access the API using a unique API key generated from your account dashboard. Include this key in the Authorization header of every request using the Bearer scheme. The base URL for all endpoints is https://api.embeddingapis.com/v1. This URL works with the official OpenAI SDKs and any client configured for OpenAI-compatible services. Ensure your key is kept secure, as it grants access to your prepaid credits.

Each account is limited to one active key, which can be regenerated at any time. Regenerating the key immediately revokes the previous one, so update your configuration promptly.

The Chat Completions Endpoint

Send text prompts to the POST /v1/chat/completions endpoint. The API serves a single uncensored model identified by the ID uncensored. This model is an open-weight LLM tuned to answer without content refusals for lawful adult, fictional, or security-research topics. It processes text in and outputs text, making it ideal for post-processing, summarization, or extraction steps in your pipeline.

Structure your request with a messages array containing your system and user prompts. The API returns a JSON response containing the generated text.

curl https://api.embeddingapis.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Use the official OpenAI Python library to interact with the API. Set the base_url to our endpoint and provide your API key. This approach allows you to leverage existing client code with minimal changes. The client handles serialization and network requests, returning structured responses that you can parse for your application logic.

Ensure you are using a recent version of the SDK to support all features, including streaming and tool calling. The model ID remains uncensored when invoking completions.

from openai import OpenAI

client = OpenAI(base_url="https://api.embeddingapis.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node SDK Integration

For JavaScript or TypeScript environments, use the OpenAI Node.js SDK. Configure the client with the correct base URL and API key. This setup ensures compatibility with your existing Node-based pipelines. The SDK manages the connection to the POST /v1/chat/completions endpoint, allowing you to send prompts and receive text completions efficiently.

This integration is suitable for server-side text processing, such as generating captions, scripts, or extracting data from large text blocks.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.embeddingapis.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

For low-latency applications, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This is particularly useful for real-time user interfaces or when processing long texts where immediate feedback is valuable.

Parse the SSE stream to handle partial responses. The uncensored model supports this mode, allowing you to react to content as it is produced without waiting for the entire completion.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context Window

The context window supports 100,000 tokens for both prompt and completion combined. Requests exceeding the 8 MB body limit or the rate limit of 300 requests per minute will fail. Authentication errors return a 401 status if the key is invalid. When prepaid credit is exhausted, requests fail because no credits remain to cover the token usage. A 429 status indicates rate limit exhaustion.

A hard content limit always applies: requests containing sexual content involving minors are blocked, regardless of other settings. All other lawful adult content is processed without refusal.

Capabilities and limits

The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.

ItemValue
ProtocolOpenAI Chat Completions schema; official openai SDKs work unchanged
AuthenticationAuthorization: Bearer YOUR_KEY
MethodsPOST /v1/chat/completions · GET /v1/models
Model IDuncensored
Base URLhttps://api.embeddingapis.com/v1
StreamingSupported (stream: true), usage included at the end
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Max outputup to 16,000 tokens per request (default 2,048)
Context window100,000 tokens (prompt + completion together)
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
JSON modeJSON object mode via response_format json_object
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request size8 MB request body
Concurrency8 requests at the same time per key
Rate limit300/min per key
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Subscriptionpaid credit never expires, no subscription
How you paypay as you go from prepaid credit; nothing is charged for failed or refused requests
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Trial credit$0.50 for 7 days, no card
Bonus credit+5% on $50+, +10% on $100+
AccountGoogle or e-mail and password
Key managementone active key per account; a new key replaces the old one
Contentadult content allowed; sexual content involving minors is refused

Error codes

Every error is JSON with a type you can switch on. You are never charged for an error.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Does this API support embeddings generation?

No. This API is a text chat-completions endpoint. It does not generate embeddings, images, audio, or video. Use it for text generation tasks like summarization or extraction.

How is pricing calculated?

Pricing is pay-as-you-go prepaid credit. Input tokens cost $0.25 per 1M tokens, and output tokens cost $1.00 per 1M tokens. Credits do not expire.

Is the model uncensored?

Yes, the model does not refuse lawful adult, fictional, or controversial topics. However, sexual content involving minors is always blocked.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API keyRead the docs