Get API key

Hermes API: Switch your client in three lines

Get started with the hermes api in minutes using standard OpenAI SDKs. This quickstart guide covers configuration, basic requests, streaming, and error handling for your uncensored LLM integration.

per 1M input tokens
$0.25
Output tokens / 1M
$1.00
token context
100,000
trial credit
$0.50
requests per minute
300

Install the OpenAI SDK

Since our endpoint is OpenAI-compatible, you can use the official SDKs you already know. This approach lets you switch from other vendors without rewriting your application logic. We support the standard openai package for Python and Node.js environments.

Ensure you are using a recent version of the SDK to guarantee compatibility with streaming and tool calling features. The client library handles the JSON serialization and HTTP transport automatically, so you can focus on the model parameters rather than network details.

Configure Base URL & Key

Authentication is handled via a Bearer token in the Authorization header. You obtain your key by signing up with an email and password on our dashboard. No credit card is required for the trial, and you can regenerate your key at any time to revoke access.

Set the base URL to our dedicated endpoint. This ensures all requests go to our uncensored model infrastructure instead of the default OpenAI cloud. You can find your API key on the 'Get API key' page immediately after signup.

Send Your First Request

Send a standard chat completion request to test connectivity. The model ID is always uncensored. This model is tuned to answer without content refusals for lawful adult use, making it ideal for creative writing, roleplay, or unrestricted research.

  • Use POST /v1/chat/completions for text generation.
  • The context window supports up to 100,000 tokens for prompt plus completion.
  • Requests are charged based on input and output token counts.
curl https://api.hermesllmapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Enable Streaming (SSE)

For better user experience, enable streaming to receive tokens as they are generated. This reduces perceived latency, especially for longer responses. The server returns Server-Sent Events (SSE) that your client can parse in real-time.

Set the stream parameter to true in your request body. The SDK will yield chunks as they arrive. This is particularly useful for chat interfaces where you want to display text incrementally rather than waiting for the entire response to finish.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Use Tool/Function Calling

Our API supports function calling, allowing the model to output structured JSON for external tool execution. Define your functions in the tools parameter, and the model will return tool_calls in the response if appropriate.

This feature works exactly as documented in the OpenAI specification. You can parse the function name and arguments, execute your backend logic, and feed the results back into the conversation. It is fully compatible with standard OpenAI SDK tool-calling workflows.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.hermesllmapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Check Available Models

Query the models endpoint to verify connectivity and see available options. Currently, we offer a single dedicated model: uncensored. This simplicity avoids the complexity of model routing or generic infrastructure, ensuring predictable behavior.

You can also use this endpoint to debug authentication issues. If the request fails, check your API key validity and network connectivity. The response will list the model details, including the context window size and token limits.

from openai import OpenAI

client = OpenAI(base_url="https://api.hermesllmapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Handle Rate Limits & Errors

Our API enforces a limit of 300 requests per minute per key. If you exceed this, you will receive a 429 error. We also support an 8 MB request body limit for complex prompts or large context windows.

Common errors include 401 (invalid key) and 402 (insufficient credit). Prepaid credit never expires, so you can top up at your convenience. Credit bonuses are applied automatically for larger purchases: +5% from $50 and +10% from $100.

What the API supports

Everything the endpoint can and cannot do, in one place — check it before you top up.

SpecValue
API formatOpenAI Chat Completions schema; official openai SDKs work unchanged
Modeluncensored
MethodsPOST /v1/chat/completions · GET /v1/models
AuthenticationAuthorization: Bearer YOUR_KEY
Base URLhttps://api.hermesllmapi.com/v1
Max output16,000 tokens max; 2,048 if max_tokens is not set
JSON modeJSON object mode via response_format json_object
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Max context100,000 tokens (prompt + completion together)
StreamingYes — server-sent events; the last chunk carries token usage
Concurrencyup to 8 in parallel per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Requests per minute300/min per key
Request size8 MB request body
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Bonus credit+5% from $50, +10% from $100
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Free trial$0.50 of credit valid 7 days, no card needed
Billingprepaid credit, charged by real token usage; errors and refusals are free
Subscriptionno monthly fee; paid credit does not expire
Key managementone active key per account; a new key replaces the old one
Sign-inGoogle or e-mail and password
Contentuncensored for adults; the only hard rule: no sexual content involving minors

Error reference

Errors come back as JSON with a stable type; failed and refused requests are not billed.

HTTPTypeWhat to do
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly
01

Questions and answers

Does the hermes api support image or audio generation?

No, we currently offer a text-only chat-completions API. We do not support embeddings, image, audio, or video generation, nor do we offer fine-tuning. Our focus is on providing a straightforward, uncensored text model for standard OpenAI-compatible clients.

How does the uncensored model handle adult content?

The model does not refuse lawful adult, fictional, or controversial topics. However, there is a hard content limit: requests involving sexual content with minors are always blocked. This is the only strict restriction applied to your prompts.

What happens if I run out of credit?

Your API key remains active, but requests will return a 402 error indicating insufficient funds. You can top up your account with crypto (USDT or USDC) at any time. Your prepaid credit never expires, so you can add funds whenever you are ready to continue.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API keyRead the docs