API Docs

$…·Add funds
·

API Docs

OpenAI-compatible. Set base_url to https://api.eastx.ai/v1 and your existing SDK code works.

Looking for the Merchant Management API?

This page covers the end-user inference API — calls to /v1/* from any OpenAI-compatible SDK. If you’re a white-label tenant admin and want to manage brand, pricing, finance, or custom domains programmatically, use the separate Merchant Management API: token prefix sk-mgmt-*, base path /api/v1/merchant/*. Mint a token from the Console under API keys; the browsable OpenAPI spec lives at admin.eastx.ai/api/v1/merchant/docs.

Quickstart

One chat completions call — pick your favorite client.

curl https://api.eastx.ai/v1/chat/completions \
  -H "Authorization: Bearer $EASTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Authentication

All endpoints require an API key in the Authorization header:

Authorization: Bearer sk-eastx-<your-key>
  • • Keys are issued at /keys and shown to you ONCE — store them securely.
  • • Optional per-key spend cap and RPM limit can be set when creating a key.
  • • Revoke compromised keys immediately from the same page; revocation is instant.
GET/v1/models

List models

Returns the catalog of models your account can call. Filters out coming-soon and disabled models.

Example

curl https://api.eastx.ai/v1/models \
  -H "Authorization: Bearer $EASTX_API_KEY"
  • Each entry has OpenAI-shape fields (id, object, created, owned_by) plus extras: display_name, input_price_per_1m, output_price_per_1m. Most OpenAI SDKs ignore unknown fields.
GET/v1/models/{id}

Retrieve a model

Lookup a single model by id. Used by some SDKs during client initialization to validate the model name.

Example

curl https://api.eastx.ai/v1/models/claude-sonnet-4-6 \
  -H "Authorization: Bearer $EASTX_API_KEY"
  • 404 model_not_found if the id is unknown or disabled.
POST/v1/chat/completions

Chat completions

OpenAI-compatible chat-completions endpoint. Supports streaming (SSE) and non-streaming. Token usage is tracked at the upstream and your balance is debited at the end of each request.

Request body

{
  "model": "claude-sonnet-4-6",
  "messages": [
    {"role": "system", "content": "You are concise."},
    {"role": "user", "content": "Explain JWT in one sentence."}
  ],
  "max_tokens": 256,
  "temperature": 0.2,
  "stream": false
}

Example

curl https://api.eastx.ai/v1/chat/completions \
  -H "Authorization: Bearer $EASTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "messages": [{"role":"user","content":"hi"}],
    "max_tokens": 32
  }'
  • For streaming, set "stream": true. Each SSE event is "data: {...}" carrying an OpenAI-shape chunk with "choices[0].delta.content" for partial text and "choices[0].finish_reason" on the final non-DONE chunk. Consume until you see "data: [DONE]".
  • Vision and tool / function calling are passed through to upstream models that support them (e.g. Claude, qwen-vl-*, GPT-class). See the examples below.
  • Balance is pre-checked against worst-case cost (input + max_tokens × output_price). Insufficient balance returns 402.
  • Reasoning models (e.g. gpt-5.5) bill the chain-of-thought as output_tokens — your usage.completion_tokens may exceed the visible message length. The chain itself is not exposed in the response, mirroring OpenAI’s o-series behavior.

Tool / function calling

Request body

{
  "model": "claude-sonnet-4-6",
  "messages": [
    {"role": "user", "content": "What's the weather in Paris?"}
  ],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city.",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }]
}
  • The model decides whether to call the tool. When it does, the response’s "choices[0].message.tool_calls" lists the function name + JSON arguments, and "finish_reason" is "tool_calls". Send back a follow-up request with the result as a {"role":"tool", ...} message.
  • Not every upstream supports tools. Currently Anthropic Claude, Qwen-class, and most OpenAI-compatible models do; small open-source models often do not.

Vision (image inputs)

Request body

{
  "model": "claude-sonnet-4-6",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this image?"},
      {"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
    ]
  }]
}
  • Pass each user message’s "content" as an array of parts instead of a plain string. "image_url" accepts either an HTTPS URL or a "data:image/<mime>;base64,<...>" data URL.
  • Use a vision-capable model (e.g. claude-sonnet-4-6, qwen-vl-plus). Sending an image to a text-only model is upstream-dependent — it may 400 or silently drop the image.
POST/v1/embeddings

Embeddings

Generate vector embeddings for one or more strings. Only models with modality=embedding are accepted here (chat models return 400 wrong_endpoint).

Request body

{
  "model": "text-embedding-v3",
  "input": "Hello, world."
}

Example

curl https://api.eastx.ai/v1/embeddings \
  -H "Authorization: Bearer $EASTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-v3",
    "input": ["hello", "world"]
  }'
  • input accepts a single string or an array of strings.
  • Billing is by input tokens only (output_tokens = 0).
POST/v1/images/generations

Image generation

Generate images from a text prompt. We wrap Aliyun’s async task API so the client sees a synchronous response — typical latency is 3-10s. Charged per image regardless of resolution.

Request body

{
  "model": "qwen-image-plus",
  "prompt": "a red apple on a wooden table",
  "n": 1,
  "size": "1024x1024"
}

Example

curl https://api.eastx.ai/v1/images/generations \
  -H "Authorization: Bearer $EASTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-image-plus",
    "prompt": "a red apple on a wooden table",
    "n": 1,
    "size": "1024x1024"
  }'
  • Response shape: {"created":<ts>,"data":[{"url":"<signed-oss-url>","revised_prompt":"..."}]}.
  • n is capped at 4. Sizes supported depend on the model; 1024x1024 works everywhere.
  • URLs are temporary signed OSS links — download and re-host within ~1 hour.
  • Internal polling deadline is 90s; long jobs return 504 upstream_timeout.

Error codes

All errors return JSON of shape {"error":{"code":"...","message":"...","type":"..."}}. HTTP status maps to error.code as follows.

The type field groups codes into families so you can branch on it without pattern-matching every name: auth_error (key issues), invalid_request_error (your input), rate_limit (RPM caps), billing_error (balance / cap), upstream_error (provider side), server_error (rare — our side).

HTTPerror.codeWhen you see it
400bad_requestRequired field missing or malformed body. Inspect error.message for specifics.
400wrong_endpointModel modality does not match endpoint (e.g. chat model on /v1/embeddings).
401missing_keyNo Authorization header. Send "Authorization: Bearer sk-eastx-...".
401invalid_keyKey not found or revoked. Generate a new one at /keys.
402insufficient_balanceWorst-case cost exceeds account balance. Top up via /recharge.
402cap_reachedPer-key spend cap (balanceCapUsd) reached. Raise it or rotate the key.
403disabledAPI key was disabled by an admin.
403expiredAPI key expiresAt is in the past.
404model_not_foundUnknown model id. Call /v1/models to see what is available.
429rate_limit_exceededPer-key requests-per-minute (RPM) limit hit. Slow down or split across keys.
429tenant_rate_limit_exceededTenant-wide RPM ceiling reached — aggregated across every key in your tenant. Distinct from the per-key 429 above. Ask your tenant admin to raise the cap.
500pricing_misconfiguredThe model is missing required price config on our side (rare admin-side bug). Retry shortly; if it persists, contact support with the model id.
502upstream_unreachableCouldn't connect to the upstream provider. Usually transient — retry with backoff.
502upstream_errorUpstream returned an error. The response message has upstream details + upstream_status.
502image_generation_failedUpstream image task ended without success (returned FAILED or empty). Retry, change the prompt, or try a different image model.
503upstream_unavailableAll upstream tokens for the channel are in cooldown or disabled. We are aware.
504upstream_timeoutImage generation took longer than 90s. Submit again; same prompt may succeed.

Rate limits & quotas

  • Per-key RPM: each API key has a requests-per-minute limit (default 60). Exceeding it returns 429 rate_limit_exceeded. Raise the limit on a key in the /keys page.
  • Tenant-wide RPM ceiling: on white-label tenants, the tenant admin can set a tenant-level RPM cap that applies to all keys in the tenant combined. Hitting it returns 429 with tenant_rate_limit_exceeded — a distinct code from the per-key 429 so you can tell the two apart in your retry logic. Splitting traffic across more keys doesn’t help here; ask your tenant admin to raise the cap.
  • Account balance: every successful request is debited from your USD balance. Insufficient balance returns 402; top up via /recharge.
  • Per-key spend cap: optional balanceCapUsd cuts off a key after a cumulative spend. Useful for embedding short-lived keys in client apps.
  • Upstream cooldowns: if all upstream tokens for a channel are throttled or erroring, the gateway returns 503 upstream_unavailable. Retry on backoff.