Documentation

Zither Tech API

The API follows the OpenAI Chat Completions format, so the official OpenAI SDKs and most OpenAI-compatible tools work by changing the base URL and API key.

Overview

All requests go to one base URL and authenticate with an API key sent as a bearer token.

Base URLhttps://api.zithertech.com/v1
AuthenticationAuthorization: Bearer YOUR_API_KEY
Chat completionsPOST /v1/chat/completions
List modelsGET /v1/models

Create API keys on the API keys page of the console. Keep keys secret: anyone with a key can spend your balance. If a key leaks, delete it in the console and create a new one.

Quickstart

The examples read the key from the ZITHER_API_KEY environment variable.

curl

Terminal
curl https://api.zithertech.com/v1/chat/completions \
  -H "Authorization: Bearer $ZITHER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Write a haiku about strings."}]
  }'

Python

Install the SDK with pip install openai.

main.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.zithertech.com/v1",
    api_key=os.environ["ZITHER_API_KEY"],
    timeout=180,
)

resp = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Write a haiku about strings."}],
)
print(resp.choices[0].message.content)

Node.js

Install the SDK with npm install openai.

main.mjs
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.zithertech.com/v1",
  apiKey: process.env.ZITHER_API_KEY,
  timeout: 180_000,
});

const resp = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Write a haiku about strings." }],
});
console.log(resp.choices[0].message.content);

Streaming

Set stream to true to receive tokens as server-sent events while they are generated.

stream.py
stream = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Explain vector databases in three sentences."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print("\n", chunk.usage)

Models

Pass the model ID in the model field. GET /v1/models returns the models your key can use.

Model IDContext windowInput / 1M tokensOutput / 1M tokens
qwen3.8-27b65,536 tokens$0.35$2.00

The context window covers the prompt and the completion together.

Parameters

The common Chat Completions parameters are supported:

  • model and messages (required)
  • max_tokens: upper limit on generated tokens
  • temperature, top_p: sampling controls
  • stop: up to four stop sequences
  • presence_penalty, frequency_penalty
  • seed: best-effort reproducible sampling
  • stream and stream_options

Timeouts and cold starts

Models run on serverless GPUs that scale down when they are idle. If a model hasn't been used for a while, the first request can take up to about two minutes while the model loads. Later requests start right away.

Set your client timeout to at least 180 seconds, and use streaming in interactive apps so users see output as soon as it is generated.

Errors

Errors return a JSON body with a message that describes the problem.

StatusMeaningWhat to do
401The API key is missing, invalid or expired, or has reached its spending limit.Check the Authorization header and the key's settings in the console.
403Your balance is too low, or the key isn't allowed to use this model.Top up your balance or adjust the key.
429Too many requests.Retry with exponential backoff.
5xxA temporary problem reaching the model.Retry with exponential backoff. Contact support if it continues.

Billing

Online payment is coming soon. Until it launches, new accounts cannot add credits.

Usage is paid from prepaid credits. Every response includes a usage object with prompt_tokens and completion_tokens. The cost of a request is:

prompt_tokens × input price + completion_tokens × output price

The amount is deducted from your balance when the request completes. The usage log in the console lists the tokens and cost of every request. Prices, credit validity and refunds are covered in the Terms of Service.

Support

Email support@zithertech.com. For problems with a specific request, include the time, the model and the error message, but never your API key.