Hermes API: Switch your client in three lines
Get started with the hermes api in minutes using standard OpenAI SDKs. This quickstart guide covers configuration, basic requests, streaming, and error handling for your uncensored LLM integration.
- per 1M input tokens
- $0.25
- Output tokens / 1M
- $1.00
- token context
- 100,000
- trial credit
- $0.50
- requests per minute
- 300
Install the OpenAI SDK
Since our endpoint is OpenAI-compatible, you can use the official SDKs you already know. This approach lets you switch from other vendors without rewriting your application logic. We support the standard openai package for Python and Node.js environments.
Ensure you are using a recent version of the SDK to guarantee compatibility with streaming and tool calling features. The client library handles the JSON serialization and HTTP transport automatically, so you can focus on the model parameters rather than network details.
Configure Base URL & Key
Authentication is handled via a Bearer token in the Authorization header. You obtain your key by signing up with an email and password on our dashboard. No credit card is required for the trial, and you can regenerate your key at any time to revoke access.
Set the base URL to our dedicated endpoint. This ensures all requests go to our uncensored model infrastructure instead of the default OpenAI cloud. You can find your API key on the 'Get API key' page immediately after signup.
Send Your First Request
Send a standard chat completion request to test connectivity. The model ID is always uncensored. This model is tuned to answer without content refusals for lawful adult use, making it ideal for creative writing, roleplay, or unrestricted research.
- Use
POST /v1/chat/completionsfor text generation. - The context window supports up to 100,000 tokens for prompt plus completion.
- Requests are charged based on input and output token counts.
curl https://api.hermesllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Enable Streaming (SSE)
For better user experience, enable streaming to receive tokens as they are generated. This reduces perceived latency, especially for longer responses. The server returns Server-Sent Events (SSE) that your client can parse in real-time.
Set the stream parameter to true in your request body. The SDK will yield chunks as they arrive. This is particularly useful for chat interfaces where you want to display text incrementally rather than waiting for the entire response to finish.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Use Tool/Function Calling
Our API supports function calling, allowing the model to output structured JSON for external tool execution. Define your functions in the tools parameter, and the model will return tool_calls in the response if appropriate.
This feature works exactly as documented in the OpenAI specification. You can parse the function name and arguments, execute your backend logic, and feed the results back into the conversation. It is fully compatible with standard OpenAI SDK tool-calling workflows.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.hermesllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Check Available Models
Query the models endpoint to verify connectivity and see available options. Currently, we offer a single dedicated model: uncensored. This simplicity avoids the complexity of model routing or generic infrastructure, ensuring predictable behavior.
You can also use this endpoint to debug authentication issues. If the request fails, check your API key validity and network connectivity. The response will list the model details, including the context window size and token limits.
from openai import OpenAI
client = OpenAI(base_url="https://api.hermesllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Handle Rate Limits & Errors
Our API enforces a limit of 300 requests per minute per key. If you exceed this, you will receive a 429 error. We also support an 8 MB request body limit for complex prompts or large context windows.
Common errors include 401 (invalid key) and 402 (insufficient credit). Prepaid credit never expires, so you can top up at your convenience. Credit bonuses are applied automatically for larger purchases: +5% from $50 and +10% from $100.
What the API supports
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Spec | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Authentication | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.hermesllmapi.com/v1 |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| JSON mode | JSON object mode via response_format json_object |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Max context | 100,000 tokens (prompt + completion together) |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Concurrency | up to 8 in parallel per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Requests per minute | 300/min per key |
| Request size | 8 MB request body |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Bonus credit | +5% from $50, +10% from $100 |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Subscription | no monthly fee; paid credit does not expire |
| Key management | one active key per account; a new key replaces the old one |
| Sign-in | Google or e-mail and password |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
Error reference
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Does the hermes api support image or audio generation?
No, we currently offer a text-only chat-completions API. We do not support embeddings, image, audio, or video generation, nor do we offer fine-tuning. Our focus is on providing a straightforward, uncensored text model for standard OpenAI-compatible clients.
How does the uncensored model handle adult content?
The model does not refuse lawful adult, fictional, or controversial topics. However, there is a hard content limit: requests involving sexual content with minors are always blocked. This is the only strict restriction applied to your prompts.
What happens if I run out of credit?
Your API key remains active, but requests will return a 402 error indicating insufficient funds. You can top up your account with crypto (USDT or USDC) at any time. Your prepaid credit never expires, so you can add funds whenever you are ready to continue.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.