Overview
All requests go to one base URL and authenticate with an API key sent as a bearer token.
| Base URL | https://api.zithertech.com/v1 |
| Authentication | Authorization: Bearer YOUR_API_KEY |
| Chat completions | POST /v1/chat/completions |
| List models | GET /v1/models |
Create API keys on the API keys page of the console. Keep keys secret: anyone with a key can spend your balance. If a key leaks, delete it in the console and create a new one.
Quickstart
The examples read the key from the ZITHER_API_KEY environment variable.
curl
curl https://api.zithertech.com/v1/chat/completions \
-H "Authorization: Bearer $ZITHER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Write a haiku about strings."}]
}'
Python
Install the SDK with pip install openai.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zithertech.com/v1",
api_key=os.environ["ZITHER_API_KEY"],
timeout=180,
)
resp = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Write a haiku about strings."}],
)
print(resp.choices[0].message.content)
Node.js
Install the SDK with npm install openai.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.zithertech.com/v1",
apiKey: process.env.ZITHER_API_KEY,
timeout: 180_000,
});
const resp = await client.chat.completions.create({
model: "qwen3.8-27b",
messages: [{ role: "user", content: "Write a haiku about strings." }],
});
console.log(resp.choices[0].message.content);
Streaming
Set stream to true to receive tokens as server-sent events while they are generated.
stream = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Explain vector databases in three sentences."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)
Models
Pass the model ID in the model field. GET /v1/models returns the models your key can use.
| Model ID | Context window | Input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
qwen3.8-27b | 65,536 tokens | $0.35 | $2.00 |
The context window covers the prompt and the completion together.
Parameters
The common Chat Completions parameters are supported:
modelandmessages(required)max_tokens: upper limit on generated tokenstemperature,top_p: sampling controlsstop: up to four stop sequencespresence_penalty,frequency_penaltyseed: best-effort reproducible samplingstreamandstream_options
Timeouts and cold starts
Models run on serverless GPUs that scale down when they are idle. If a model hasn't been used for a while, the first request can take up to about two minutes while the model loads. Later requests start right away.
Set your client timeout to at least 180 seconds, and use streaming in interactive apps so users see output as soon as it is generated.
Errors
Errors return a JSON body with a message that describes the problem.
| Status | Meaning | What to do |
|---|---|---|
401 | The API key is missing, invalid or expired, or has reached its spending limit. | Check the Authorization header and the key's settings in the console. |
403 | Your balance is too low, or the key isn't allowed to use this model. | Top up your balance or adjust the key. |
429 | Too many requests. | Retry with exponential backoff. |
5xx | A temporary problem reaching the model. | Retry with exponential backoff. Contact support if it continues. |
Billing
Online payment is coming soon. Until it launches, new accounts cannot add credits.
Usage is paid from prepaid credits. Every response includes a usage object with prompt_tokens and completion_tokens. The cost of a request is:
prompt_tokens × input price + completion_tokens × output price
The amount is deducted from your balance when the request completes. The usage log in the console lists the tokens and cost of every request. Prices, credit validity and refunds are covered in the Terms of Service.
Support
Email support@zithertech.com. For problems with a specific request, include the time, the model and the error message, but never your API key.