Skip to content
VexAPI Documentation

VexAPI by Arioron

Build with intelligence

VexAPI is one API for Arioron models: text, vision, documents, tool calling, structured output, embeddings, agents and more. It speaks both the OpenAI and the Anthropic wire formats, so the official SDKs you already use work by changing one line.

Base URL

https://api.arioron.com/v1

One platform, three wire formats

Pick whichever matches your code. All three run on the same models with identical capabilities, and every request shows up in the same usage and logs.

EndpointWire formatBest for
POST /v1/responsesOpenAI ResponsesNew projects. Built-in web search, server-side conversation state.
POST /v1/chat/completionsOpenAI Chat CompletionsExisting OpenAI integrations and the wider ecosystem of tools built on it.
POST /v1/messagesAnthropic MessagesExisting Anthropic SDK integrations.

Quickstart

1. Create an API key on the API keys page and export it:

export VEX_API_KEY="ar_..."

2. Install the SDK you prefer. VexAPI works with the official OpenAI and Anthropic SDKs as-is.

pip install openai

3. Make your first request:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.arioron.com/v1",
    api_key=os.environ["VEX_API_KEY"],
)

response = client.responses.create(
    model="perceptix-vex-amber-2.0",
    input="Write a one-sentence bedtime story about a lighthouse.",
)
print(response.output_text)
Anthropic SDK base URL. The Anthropic SDK appends /v1/messages itself, so its base URL is https://api.arioron.com, without /v1.

Authentication

Send your key in the Authorization header. The x-api-key header is also accepted, which is what the Anthropic SDKs send. Keys are never accepted in the URL, because URLs end up in proxy and access logs.

Authorization: Bearer ar_...
# or
x-api-key: ar_...

Keeping keys safe

  • Keys are secrets. Call VexAPI from your server, never from a browser or mobile app.
  • Each key belongs to a project, so its usage, logs and rate limits are tracked separately.
  • Create read-only keys for dashboards and reporting. They can call GET endpoints only.
  • Give keys an expiry and rotate them. Expired or revoked keys stop working immediately.
StatusCodeMeaning
401missing_api_keyNo key was sent.
401invalid_api_keyThe key doesn't exist or was revoked.
401expired_api_keyThe key passed its expiry date.
403insufficient_permissionsA read-only key tried to write.

Models

Pass a model id as model. List models programmatically with GET /v1/models.

ModelContextMax outputCapabilities
perceptix-vex-amber-2.0 1,048,576 65,536 Vision Tools Structured Reasoning Streaming
perceptix-vex-amber-latest
Strongest Amber Model Yet
1,048,576 65,536 Vision Tools Structured Reasoning Streaming
photon-lite-latest
Latest series of Photon Lite models
1,048,576 65,536 Vision Tools Structured Reasoning Streaming
vex-embed
General-purpose text embeddings for search, clustering and retrieval. Supports reduced dimensions (e.g. 768, 1536).
—3,072 dims Embeddings

Admins can also define model aliases: a stable name you call, routed to the healthiest of several models.

Responses

POST/v1/responses

The recommended interface. Send a string or a list of input items and get back typed output items. VexAPI stores responses by default, so you can continue a conversation by passing previous_response_id instead of resending history.

response = client.responses.create(
    model="perceptix-vex-amber-2.0",
    instructions="You are a concise travel assistant.",
    input="Three things to do in Lisbon in one day.",
)
print(response.output_text)

# Continue the conversation without resending history
follow_up = client.responses.create(
    model="perceptix-vex-amber-2.0",
    previous_response_id=response.id,
    input="Which of those is best on a rainy day?",
)
print(follow_up.output_text)

Request body

modelstringrequired

Model id or alias.

inputstring | arrayrequired

A string, or a list of items: messages (role user, assistant, system or developer, with text, input_image or input_file parts), function_call and function_call_output.

instructionsstring

System instructions for this response. Not carried over to responses that chain from it.

previous_response_idstring

Continue from a stored response. Its full input and output become this request's history.

toolsarray

Function tools ({"type": "function", "name", "description", "parameters"}) and the built-in {"type": "web_search"}.

tool_choicestring | object

auto (default), none, required, or {"type": "function", "name": ...}.

text.formatobject

{"type": "text"} (default), {"type": "json_object"}, or {"type": "json_schema", "name", "schema"}. See structured outputs.

reasoning.effortstring

minimal, low, medium or high. See reasoning.

max_output_tokensinteger

Upper bound on output tokens, including reasoning tokens.

temperaturenumber

0 to 2. Higher is more varied.

top_pnumber

0 to 1. Nucleus sampling.

streamboolean

Stream typed server-sent events. See streaming.

storeboolean

Default true. When false, the response isn't retrievable or chainable, and request/response bodies aren't kept in your logs.

metadataobject

Up to you. Echoed back on the response.

Response object

{
  "id": "resp_8f3c…",
  "object": "response",
  "created_at": 1759567200,
  "status": "completed",
  "model": "perceptix-vex-amber-2.0",
  "output": [
    {
      "type": "message",
      "id": "msg_2b1a…",
      "status": "completed",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "1. Ride Tram 28…", "annotations": [] }]
    }
  ],
  "usage": {
    "input_tokens": 24,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 162,
    "output_tokens_details": { "reasoning_tokens": 41 },
    "total_tokens": 186
  },
  "previous_response_id": null,
  "error": null,
  "incomplete_details": null
}

status is completed, or incomplete with incomplete_details.reason set to max_output_tokens or content_filter. Output items are message, function_call and web_search_call.

Add {"type": "web_search"} to tools and the model searches the live web on its own. VexAPI runs the searches server-side and returns a web_search_call item, with the query and sources, before the final message. Tokens for every model turn are included in usage.

response = client.responses.create(
    model="perceptix-vex-amber-2.0",
    tools=[{"type": "web_search"}],
    input="What changed in the latest Python release?",
)
for item in response.output:
    if item.type == "web_search_call":
        print("searched:", item.action.query)
print(response.output_text)

Retrieve and delete

GET/v1/responses/{response_id}
DELETE/v1/responses/{response_id}

Stored responses are private to your account. Deleting one also stops it being used as a previous_response_id.

Chat Completions

POST/v1/chat/completions

Fully compatible with the OpenAI Chat Completions format, so existing code, frameworks and tools built on it work unchanged.

completion = client.chat.completions.create(
    model="perceptix-vex-amber-2.0",
    messages=[
        {"role": "developer", "content": "You answer in one sentence."},
        {"role": "user", "content": "Why is the sky blue?"},
    ],
)
print(completion.choices[0].message.content)
print(completion.usage.prompt_tokens, completion.usage.completion_tokens)

Request body

modelstringrequired

Model id or alias.

messagesarrayrequired

Roles system, developer, user, assistant and tool. User content can mix text, image_url and file parts.

tools / tool_choicearray / string | object

Function tools. The model returns tool_calls; you run them and send back tool messages. See tool calling.

response_formatobject

{"type": "json_object"} or {"type": "json_schema", "json_schema": {"name", "schema"}}.

max_tokensinteger

Caps visible output. When set without reasoning_effort, reasoning is kept minimal so short limits don't truncate answers.

max_completion_tokensinteger

Caps output including reasoning tokens.

reasoning_effortstring

minimal, low, medium or high.

temperature / top_pnumber

Sampling controls (0–2 and 0–1).

stopstring | array

Up to 4 stop sequences.

stream / stream_optionsboolean / object

Stream chunks. Set stream_options.include_usage for a final usage chunk.

storeboolean

Default false. When true, request and response bodies are kept in your logs for inspection.

Differences from OpenAI. n must be 1, and log probabilities and audio input/output aren't supported. seed, user, penalties and logit_bias are accepted but have no effect.

Messages

POST/v1/messages
POST/v1/messages/count_tokens

Compatible with the Anthropic Messages API. Point the Anthropic SDK at VexAPI and use an Arioron model id. Errors on these endpoints use Anthropic's error shape, so the SDK raises the exception classes you already handle.

import anthropic

client = anthropic.Anthropic(base_url="https://api.arioron.com", api_key=os.environ["VEX_API_KEY"])

message = client.messages.create(
    model="perceptix-vex-amber-2.0",
    max_tokens=1024,
    system="You only answer in French.",
    messages=[{"role": "user", "content": "Say hello."}],
)
print(message.content[0].text, message.usage.output_tokens)

count = client.messages.count_tokens(
    model="perceptix-vex-amber-2.0",
    messages=[{"role": "user", "content": "How many tokens is this?"}],
)
print(count.input_tokens)
modelstringrequired

Model id or alias.

max_tokensintegerrequired

Maximum visible output tokens.

messagesarrayrequired

user and assistant turns with text, image, document, tool_use and tool_result blocks. Image and document sources can be base64, url or file.

systemstring | array

System prompt.

tools / tool_choicearray / object

Tools with input_schema; tool_choice types auto, any, tool and none.

thinkingobject

{"type": "enabled", "budget_tokens": n} turns on extended reasoning. Off by default, so max_tokens is all visible output.

temperature / top_p / stop_sequences / stream—

As in the Anthropic API.

Differences from Anthropic. Server tools (such as Anthropic's web search tool) and prompt caching aren't available on this endpoint; cache_control is ignored. For built-in web search, use Responses.

Embeddings

POST/v1/embeddings

Turn text into vectors for search, clustering and retrieval. Vectors are L2-normalized, so cosine similarity is a dot product.

result = client.embeddings.create(
    model="vex-embed",
    input=["The quick brown fox", "A fast auburn fox"],
    dimensions=768,  # optional: smaller vectors, cheaper storage
)
vectors = [d.embedding for d in result.data]
modelstringrequired

vex-embed.

inputstring | arrayrequired

One string or up to 2,048 strings.

dimensionsinteger

Output size, 1 to 3,072. Defaults to 3,072. 768 and 1,536 are good trade-offs.

encoding_formatstring

float (default) or base64 (little-endian float32). The OpenAI Python SDK uses base64 automatically.

Tool calling

Describe your functions with JSON Schema. When the model wants one, it returns a call with arguments instead of text. You run the function and send the result back, and the model uses it to answer. VexAPI never runs your functions; only you do.

import json

tools = [{
    "type": "function",
    "function": {
        "name": "get_order_status",
        "description": "Look up the status of a customer order",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
        },
    },
}]
messages = [{"role": "user", "content": "Where is order A-1001?"}]

first = client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)          # {"order_id": "A-1001"}

messages.append(first.choices[0].message)
messages.append({
    "role": "tool",
    "tool_call_id": call.id,
    "content": json.dumps({"status": "shipped", "eta": "Oct 9"}),  # your function's result
})
final = client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=messages, tools=tools)
print(final.choices[0].message.content)

Use tool_choice to force a call (required), a specific function, or none at all. The model may request several calls in one turn; answer each by its id.

Structured outputs

Give a JSON Schema and the output is guaranteed to parse and match it. Use json_object when you only need valid JSON.

from pydantic import BaseModel

class Weather(BaseModel):
    city: str
    temp_c: float

completion = client.chat.completions.create(
    model="perceptix-vex-amber-2.0",
    messages=[{"role": "user", "content": "Make up today's weather for Paris."}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "weather", "schema": Weather.model_json_schema(), "strict": True},
    },
)
weather = Weather.model_validate_json(completion.choices[0].message.content)

Images & files

Models marked Vision read images, PDFs and text files alongside your prompt. Provide them as a public URL, a data: URL, or a file_id from the Files API.

completion = client.chat.completions.create(
    model="perceptix-vex-amber-2.0",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What's in this image?"},
        {"type": "image_url", "image_url": {"url": "https://example.com/photo.png"}},
    ]}],
)

# A PDF you uploaded earlier
response = client.responses.create(
    model="perceptix-vex-amber-2.0",
    input=[{"role": "user", "content": [
        {"type": "input_file", "file_id": "file_abc123"},
        {"type": "input_text", "text": "Summarize this contract in five bullets."},
    ]}],
)
URL inputs. VexAPI downloads image URLs itself, up to 20 MB and over http(s) only. URLs that point at private or internal networks are refused with invalid_image_url.

Streaming

Set stream: true to receive tokens as they're generated, as server-sent events in each format's native event style.

# Responses: typed events
with client.responses.stream(model="perceptix-vex-amber-2.0", input="Write a haiku about rain.") as stream:
    for event in stream:
        if event.type == "response.output_text.delta":
            print(event.delta, end="", flush=True)
    final = stream.get_final_response()

# Chat Completions: chunks, plus a usage chunk at the end
for chunk in client.chat.completions.create(
    model="perceptix-vex-amber-2.0",
    messages=[{"role": "user", "content": "Write a haiku about rain."}],
    stream=True,
    stream_options={"include_usage": True},
):
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
EndpointEvents
/v1/responsesresponse.created, response.output_item.added, response.output_text.delta, response.function_call_arguments.done, response.web_search_call.completed, response.completed (or .incomplete / .failed). Every event has a sequence_number.
/v1/chat/completionschat.completion.chunk objects, then data: [DONE].
/v1/messagesmessage_start, content_block_start/_delta/_stop, message_delta, message_stop.

Reasoning

Models marked Reasoning can think before answering, which helps with math, code and multi-step problems. Reasoning tokens are billed as output and reported separately in reasoning_tokens.

  • Responses: reasoning: {"effort": "low"}. max_output_tokens includes reasoning.
  • Chat Completions: reasoning_effort. max_completion_tokens includes reasoning. The older max_tokens caps only visible text and keeps reasoning minimal.
  • Messages: off unless you pass thinking, so max_tokens is all visible output.

Use minimal for fast, cheap answers to simple prompts, and high when quality matters more than latency.

Migrating from OpenAI

Keep the official openai SDK. Change two lines and pick an Arioron model:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.arioron.com/v1",          # was: https://api.openai.com/v1
    api_key=os.environ["VEX_API_KEY"],
)

client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=[...])   # was: gpt-…
client.responses.create(model="perceptix-vex-amber-2.0", input="...")
client.embeddings.create(model="vex-embed", input="...")          # was: text-embedding-…

What works the same: Chat Completions and Responses formats, streaming, function calling, JSON mode and JSON Schema outputs, vision, previous_response_id, embeddings with dimensions, error classes and status codes. Model ids that don't exist on VexAPI return a 404 that lists the models that do.

Not available yet: n > 1, log probabilities and audio in chat. The file_search and code_interpreter built-in tools work in the Responses API.

Migrating from Anthropic

import anthropic

client = anthropic.Anthropic(
    base_url="https://api.arioron.com",       # was: https://api.anthropic.com
    api_key=os.environ["VEX_API_KEY"],
)
client.messages.create(model="perceptix-vex-amber-2.0", max_tokens=1024, messages=[...])

What works the same: the Messages format, system prompts, streaming, tool use with tool_use/tool_result, images and PDFs, count_tokens, extended thinking, and errors that map to the SDK's exception classes. Not available on this endpoint: server tools and prompt caching.

Files

POST/v1/files
GET/v1/files
GET/v1/files/{file_id}
GET/v1/files/{file_id}/content
DELETE/v1/files/{file_id}

Upload documents and images to reference by file_id in model inputs, batches and fine-tuning. Up to 25 MB per file. Supported types: PDF, TXT, DOCX, CSV, JSON, PNG, JPEG, GIF, WebP, MP3, WAV, MP4 and MOV.

uploaded = client.files.create(file=open("contract.pdf", "rb"), purpose="assistants")
print(uploaded.id)   # file_…

content = client.files.content(uploaded.id).read()   # download bytes, e.g. batch results

Agents

POST/v1/agents
POST/v1/agents/{agent_id}/runs
GET/v1/agents/{agent_id}/runs/{run_id}
POST/v1/agents/{agent_id}/runs/{run_id}/cancel

An agent is a saved model, instructions and set of built-in tools. Each run loops through model calls and tool calls until it has an answer, and every step is recorded in the run's trace so you can see exactly what it did. Runs can execute synchronously or in the background, with step limits, timeouts and cancellation.

# Create an agent that may only use web search
curl https://api.arioron.com/v1/agents -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
  "name": "Research assistant",
  "model": "perceptix-vex-amber-2.0",
  "instructions": "Research the question, then answer in five bullets with sources.",
  "tools": ["web_search"],
  "max_steps": 8
}'

# Run it in the background
curl https://api.arioron.com/v1/agents/agent_…/runs -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
  "input": "What were the biggest announcements at this year's PyCon?",
  "background": true
}'
# → 202 {"id": "run_…", "status": "queued", …}

# Poll for the result and the full step-by-step trace
curl https://api.arioron.com/v1/agents/agent_…/runs/run_… -H "Authorization: Bearer $VEX_API_KEY"
name, modelstringrequired

Display name and model id.

instructionsstring

The agent's system prompt.

toolsarray

Built-in tools the agent may use: web_search, web_scrape, extract_article, get_page_metadata, get_weather, get_current_time, calculate, convert_units. Calls to any other tool are refused.

max_steps / timeout_secondsinteger

Defaults 8 and 120. The run stops when either is reached.

memory_enabledboolean

Pull relevant memories scoped to this agent into each run.

Run status moves from queued to running to completed, failed or cancelled. Pass conversation_id to run an agent on top of a stored conversation. Subscribe to agent_run.completed with a webhook instead of polling.

Conversations

POST/v1/conversations
GET/v1/conversations/{conversation_id}
POST/v1/conversations/{conversation_id}/messages
DELETE/v1/conversations/{conversation_id}

Persistent message threads stored by VexAPI. Post a user message with generate_response: true and a model to append the assistant's reply in the same call.

curl https://api.arioron.com/v1/conversations/conv_…/messages -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
  "role": "user",
  "content": "What is our refund window?",
  "generate_response": true,
  "model": "perceptix-vex-amber-2.0"
}'

Memory

POST/v1/memories
GET/v1/memories
PATCH/v1/memories/{memory_id}
DELETE/v1/memories/{memory_id}

Long-term facts your app or agents can recall by meaning, not keywords. Scope each memory to your account (user), a project, an agent or a conversation; agents with memory_enabled retrieve their own memories automatically.

curl https://api.arioron.com/v1/memories -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
  "content": "The customer prefers email over phone calls.",
  "scope": "agent",
  "owner_ref": "agent_…"
}'

Batch

POST/v1/batches
GET/v1/batches/{batch_id}
POST/v1/batches/{batch_id}/cancel

Run thousands of requests in the background. Upload a JSONL file with one request per line, create a batch, and download a JSONL file of results when it finishes. Each line is processed exactly like the live endpoint, and one bad line doesn't fail the batch.

{"custom_id": "q1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "perceptix-vex-amber-2.0", "messages": [{"role": "user", "content": "Classify: 'love it'"}]}}
{"custom_id": "q2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "perceptix-vex-amber-2.0", "messages": [{"role": "user", "content": "Classify: 'never again'"}]}}

Supported endpoints: /v1/chat/completions, /v1/embeddings and /v1/moderations. Status moves through validating, queued and running to completed, cancelled or failed. A completed batch can still contain failed lines; request_counts.failed says how many. Each output line is {"custom_id", "status_code", "response", "error"}. Token usage is billed to the key that created the batch.

Webhooks

POST/v1/webhooks
GET/v1/webhooks
PATCH/v1/webhooks/{webhook_id}
DELETE/v1/webhooks/{webhook_id}

Get a signed HTTP POST when work finishes instead of polling. Events: agent_run.completed, batch.completed, or * for everything. Failed deliveries are retried with backoff, up to 3 attempts.

The signing secret is shown once, in the create response. Verify every delivery before trusting it:

import hashlib, hmac

def verify(raw_body: bytes, signature: str, secret: str) -> bool:
    expected = hmac.new(secret.encode(), raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, signature)

# Flask example
@app.post("/hooks/vexapi")
def vexapi_hook():
    if not verify(request.get_data(), request.headers["X-VexAPI-Signature"], WEBHOOK_SECRET):
        abort(401)
    event = request.get_json()   # {"type": "batch.completed", "created": …, "data": {…}}
    return "", 204

Background jobs

GET/v1/jobs
GET/v1/jobs/{job_id}
POST/v1/jobs/{job_id}/cancel

Batches, background agent runs, evaluations and webhook deliveries all run on one durable job queue. These endpoints give a single view of that work: status, progress, attempts and errors. Filter with ?type= and ?status=. Jobs survive restarts: if a worker stops mid-job, the job is picked up again automatically.

Evaluations

POST/v1/evals
POST/v1/evals/{eval_id}/runs
GET/v1/evals/{eval_id}/runs/{run_id}

Measure a model against your own test cases before you ship a prompt or switch models. Grade with exact_match, contains, or llm_judge, which uses a model to score each answer from 0 to 1 against the expected output. Runs report a pass rate plus per-case output, score and latency.

curl https://api.arioron.com/v1/evals -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
  "name": "Capital cities",
  "grader_type": "contains",
  "cases": [
    {"input": "Capital of Australia?", "expected_output": "Canberra"},
    {"input": "Capital of Canada?", "expected_output": "Ottawa"}
  ]
}'
curl https://api.arioron.com/v1/evals/eval_…/runs -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "perceptix-vex-amber-2.0"}'

Code interpreter

POST/v1/containers
POST/v1/containers/{id}/files
POST/v1/containers/{id}/execute
POST/v1/containers/{id}/files/edit

Models can write and run Python to calculate, analyse data and create files. Code runs in a secure sandbox (Python 3.14 compiled to WebAssembly): it has no internet access, can't see anything outside its own workspace, and every run is limited in time (30 s by default, up to 120 s) and memory (256 MB). The standard library is available; numpy and pandas are not.

A container is that workspace. Upload files into it, run code, edit files in place, and download what the code produced. Containers are private to you and expire after 20 idle minutes (configurable with expires_after.minutes).

containerstring | object

Responses tool: "auto" (a fresh container per request), {"type": "auto", "file_ids": [...]} to preload files, or an existing container id.

codestringrequired

/execute: Python source. Relative paths are in /workspace. Returns stdout, stderr, exit_code, status (completed, failed or timeout) and the files it wrote.

opstringrequired

/files/edit: view (with optional view_range), create (file_text), str_replace (old_str must appear exactly once, new_str), insert (insert_line, new_str), append or delete.

response = client.responses.create(
    model="perceptix-vex-amber-2.0",
    input="Using sales.csv, which region grew fastest from Jan to Feb? Save the growth table as growth.csv.",
    tools=[{"type": "code_interpreter", "container": {"type": "auto", "file_ids": [sales_csv.id]}}],
    include=["code_interpreter_call.outputs"],
)
print(response.output_text)
# Download files the code wrote:
call = next(o for o in response.output if o.type == "code_interpreter_call")
data = client.containers.files.content.retrieve("growth.csv", container_id=call.container_id)
Agents get the same sandbox. Give an agent the execute_python, create_file, read_file, edit_file and list_files tools and it works in your own private workspace.

Realtime voice

GET/v1/realtime (WebSocket)

Talk to a model with your voice, and hear it answer, over one WebSocket. With vex-realtime, VexAPI detects when you stop talking, the model listens to what you said, and the reply is spoken with any library voice, sentence by sentence as it's written. The protocol follows OpenAI's Realtime API: stream microphone audio as input_audio_buffer.append events and receive spoken replies as response.output_audio.delta, with live transcripts of both sides. The model answers when you stop talking, and talking over it interrupts the reply (input_audio_buffer.speech_started tells your client to stop playback). Audio is 16-bit PCM, mono, 24 kHz, base64-encoded, in both directions.

modelquery

vex-realtime (default). gpt-realtime is accepted as an alias.

session.instructionsstring

System instructions for the conversation.

session.audio.output.voicestring

Any library voice (OpenAI voice names are mapped). Default af_heart.

session.output_modalitiesarray

["audio"] (default) or ["text"] for text replies.

session.audio.input.turn_detectionobject | null

{"type": "server_vad", "silence_duration_ms": 700} (default) or null to control turns yourself with input_audio_buffer.commit and response.create.

// Server-side (Node 22+). In a browser, use ?api_key= only with short-lived keys.
const ws = new WebSocket("wss://api.arioron.com/v1/realtime?model=vex-realtime",
  { headers: { Authorization: `Bearer ${process.env.VEX_API_KEY}` } });

ws.onopen = () => ws.send(JSON.stringify({
  type: "session.update",
  session: { instructions: "You are a friendly travel agent.", audio: { output: { voice: "bm_george" } } },
}));

// Send microphone audio (PCM16, 24 kHz) as it arrives:
function onMicChunk(pcm16) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: Buffer.from(pcm16).toString("base64") }));
}

ws.onmessage = ({ data }) => {
  const event = JSON.parse(data);
  if (event.type === "response.output_audio.delta") playPcm(Buffer.from(event.delta, "base64"));
  if (event.type === "response.output_audio_transcript.delta") process.stdout.write(event.delta);
  if (event.type === "input_audio_buffer.speech_started") stopPlayback();   // user interrupted
};
Text works too. Set output_modalities to ["text"], then send conversation.item.create with an input_text message and response.create to get response.output_text.delta events.

Audio

Speech in both directions, with the same request shapes as OpenAI's audio API, so the OpenAI SDKs work unchanged.

Voices

GET/v1/voices
GET/v1/voices/{voice_id}/preview

Browse and listen to every voice, and try your own text, on the Voices page in the console. There are two kinds of voice, listed together by GET /v1/voices (filter with ?category= and ?language=):

  • Library voices (38): natural open-source voices generated on VexAPI's own servers, in English (American and British), Spanish, French, Hindi, Italian and Portuguese. Fast, with no usage quota. Ids look like af_heart: the first letter is the language, the second the gender.
  • Expressive voices (30), such as Kore and Puck: multilingual, and they follow instructions about tone and emotion. A little slower.

Each voice has a preview_url that returns a short MP3 sample.

Text to speech

POST/v1/audio/speech
modelstring

vex-speech (default). tts-1, tts-1-hd and gpt-4o-mini-tts are accepted as aliases.

inputstringrequired

Text to speak, up to 4,096 characters.

voicestring

A voice id from the voice library, such as af_heart (default), bm_george or ff_siwis; an expressive voice such as Kore or Puck; or an OpenAI voice name (alloy, nova, onyx, …), which maps to a library voice.

instructionsstring

Expressive voices only: how to say it (tone, emotion, accent, pace). For example "Speak warmly, like a flight attendant."

speednumber

0.25 to 4.0.

response_formatstring

mp3 (default), wav, or pcm (16-bit, 24 kHz, mono).

speech = client.audio.speech.create(
    model="vex-speech",
    voice="Kore",
    input="Your order has shipped and will arrive on Thursday.",
    instructions="Friendly and upbeat.",
)
speech.write_to_file("shipped.mp3")

Speech to text

POST/v1/audio/transcriptions
POST/v1/audio/translations

Upload audio as multipart/form-data (mp3, mp4, m4a, wav, webm, ogg, flac, aac or aiff, up to 20 MB). Transcriptions keep the spoken language; translations return English.

filefilerequired

The audio to transcribe.

modelstring

vex-transcribe (default). whisper-1 and gpt-4o-transcribe are accepted as aliases.

languagestring

ISO 639-1 code of the spoken language, if you know it.

promptstring

Names, terms or spellings the transcript should respect.

response_formatstring

json (default), text, srt, vtt, or verbose_json (adds language, duration and timestamped segments).

with open("meeting.m4a", "rb") as audio:
    transcript = client.audio.transcriptions.create(model="vex-transcribe", file=audio)
print(transcript.text)

Images

POST/v1/images/generations
POST/v1/images/edits

Generate images from a prompt, or edit images you upload (change a background, add or remove objects, combine up to four images). Request and response shapes match OpenAI's Images API.

promptstringrequired

What to create or change.

modelstring

vex-image (default). gpt-image-1, dall-e-3 and dall-e-2 are accepted as aliases.

ninteger

1 to 4 images.

sizestring

For example 1024x1024, 1792x1024 or 1024x1792. Mapped to the nearest supported aspect ratio.

aspect_ratiostring

Instead of size: 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9 and more.

response_formatstring

b64_json (default) or url. URLs stay valid for one hour.

image / maskfile

Edits only (multipart): one to four image files, and an optional PNG mask.

import base64

result = client.images.generate(model="vex-image", prompt="A watercolor fox in a snowy forest", size="1792x1024")
with open("fox.png", "wb") as f:
    f.write(base64.b64decode(result.data[0].b64_json))
Image generation is not enabled on this deployment yet. Both endpoints return 503 capability_not_enabled until it is.

Moderations

POST/v1/moderations

Check text for harassment, hate, sexual and dangerous content before you show or act on it.

curl https://api.arioron.com/v1/moderations -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
  -d '{"input": "Text to check"}'
# → {"results": [{"flagged": false, "categories": {"harassment": false, …}}]}

Browser sessions

POST/v1/browser/sessions
POST/v1/browser/sessions/{session_id}/actions

A real browser (Chromium) for your app or agent. Pages run JavaScript, and each session keeps its own cookies, history and open page between actions, so you can sign in, click through, fill forms and read the result. Sessions close after 15 minutes idle; the next action reopens the last page.

urlstring

Create: open this page straight away.

allowlistarray

Create: domains the session may navigate to, for example ["example.com"]. Subdomains are included.

typestringrequired

Action: navigate, click, double_click, hover, type, select, key, scroll, wait, back, forward, reload, elements, extract, screenshot, evaluate, drag.

ref / selector / textint / string

What to act on: a ref number from elements (most reliable), a CSS selector, or visible text. Mouse actions also take x and y.

formatstring

extract: markdown (default), text, html or links.

screenshotboolean

Any action: also return a PNG of the result. screenshot itself takes full_page.

import os
import requests

API = "https://api.arioron.com/v1"
H = {"Authorization": f"Bearer {os.environ['VEX_API_KEY']}"}

session = requests.post(f"{API}/browser/sessions", headers=H, json={"url": "https://news.ycombinator.com"}).json()
act = lambda **a: requests.post(f"{API}/browser/sessions/{session['id']}/actions", headers=H, json=a).json()

elements = act(type="elements")["result"]["elements"]   # [{"ref": 0, "tag": "a", "label": "new", ...}, ...]
act(type="click", text="past")
page = act(type="extract", format="markdown")["result"]
print(page["title"], page["content"][:500])

requests.delete(f"{API}/browser/sessions/{session['id']}", headers=H)
Network safety. Pages can only reach public internet addresses. Requests to localhost, private networks or cloud metadata are blocked, and downloads are disabled.

Computer use

POST/v1/computer/sessions
POST/v1/computer/sessions/{session_id}/actions

A screen your model controls with the mouse and keyboard, for computer-use agents. The environment is a browser window. Every action returns a fresh screenshot, so a vision model can look, decide the next coordinates, and act, in a loop.

displayobject

Create: {"width": 1280, "height": 800} (320–2560 each).

start_urlstring

Create: page to open first.

permissionsarray

Create: limit the session to some of screenshot, mouse, keyboard, navigate. Default: all.

typestringrequired

Action: screenshot, click, double_click, move, drag, scroll, type, keypress, wait, navigate.

x, yinteger

Screen coordinates for mouse actions. click also takes button; drag takes a path of points; scroll takes scroll_x and scroll_y.

text / keysstring / array

type takes text; keypress takes keys, such as ["CTRL", "L"] or "Enter".

session = requests.post(f"{API}/computer/sessions", headers=H,
                        json={"display": {"width": 1024, "height": 768}, "start_url": "https://www.wikipedia.org"}).json()
shot = session["screenshot"]["data"]          # base64 PNG: show it to your model

act = lambda **a: requests.post(f"{API}/computer/sessions/{session['id']}/actions", headers=H, json=a).json()
act(type="click", x=450, y=493)                 # coordinates your model chose from the screenshot
act(type="type", text="Mount Fuji")
result = act(type="keypress", keys=["ENTER"])
print(result["current_url"])                    # https://en.wikipedia.org/wiki/Mount_Fuji

Model aliases

POST/v1/model_aliases
GET/v1/model_aliases

An alias is a model name you control, like vex-fast, that routes to an ordered list of real models. Call the alias in any endpoint. If the first model is unhealthy (a high recent error rate), requests fail over to the next. Use routing_strategy: "latency" to prefer the fastest target. Creating aliases requires an admin key.

Usage & logs

GET/v1/usage
GET/v1/logs
GET/v1/logs/{request_id}

Every request is logged with its endpoint, model, status, latency and exact token counts. Usage aggregates them by day and by model (?days=30). Logs accept ?endpoint=, ?model= and ?limit=. Both are also on the Usage and Logs pages.

Projects

GET/v1/projects
POST/v1/projects
DELETE/v1/projects/{project_id}

Projects group API keys, files, agents and usage, so separate apps or environments stay separate. Resources created with a project-scoped key belong to that project.

Errors

Errors use standard HTTP status codes and one JSON shape. param names the field that caused the problem when there is one, and request_id identifies the request in your logs.

{
  "error": {
    "type": "invalid_request_error",
    "code": "model_not_found",
    "message": "The model 'gpt-4o' does not exist or is not available. Available models: perceptix-vex-amber-2.0, …",
    "param": "model",
    "request_id": "req_6b1f0c…"
  }
}
StatusTypeWhat to do
400invalid_request_errorFix the request. param and message say what's wrong.
401authentication_errorCheck the API key.
403permission_errorThe key can't do this, for example a read-only key writing.
404not_found_errorThe resource, model or route doesn't exist.
429rate_limit_errorWait for Retry-After seconds, then retry.
500, 502server_errorRetry with backoff. Include the request ID if you contact support.
501invalid_request_error / provider_unavailableThe capability isn't available yet. See availability.
503overloaded_errorThe model is temporarily over capacity. Retry shortly.

The official SDKs retry 429 and 5xx responses automatically.

Rate limits

Limits apply per API key, per minute and per day, and depend on your plan. Read requests (GET) don't count. Every counted response reports where you stand:

HeaderMeaning
x-ratelimit-limit-requestsRequests allowed per minute.
x-ratelimit-remaining-requestsRequests left in the current minute.
x-ratelimit-limit-requests-dayRequests allowed per day.
x-ratelimit-remaining-requests-dayRequests left today.
Retry-AfterOn a 429: seconds until you can retry.

Request IDs

Every response has an x-request-id header, and errors repeat it in the body. Look it up on the Logs page to see the request's model, latency, token counts and error. x-processing-ms reports how long VexAPI spent on the request.

raw = client.chat.completions.with_raw_response.create(model="perceptix-vex-amber-2.0", messages=[...])
print(raw.headers["x-request-id"])   # req_…
completion = raw.parse()

Availability

What's live today. Capabilities marked not yet available have their endpoints in place and return 501 provider_unavailable, so you can build against them now and they'll start working without code changes.

CapabilityStatus
Text generation, streaming, tool calling, structured outputAvailable
Vision, PDFs and text filesAvailable
Built-in web search (Responses)Available
EmbeddingsAvailable
Agents, conversations, memoryAvailable
Batch, webhooks, background jobsAvailable
EvaluationsAvailable
Text-to-speech: 38 library voices + 30 expressive voicesAvailable
Speech-to-text and translationAvailable
ModerationsAvailable
Browser sessions (JavaScript, clicks, forms, screenshots)Available
Computer useAvailable
File search (vector stores)Available
Code interpreter and file editing (sandboxed)Available
Realtime voiceAvailable
Image generation and editingNot yet available
Fine-tuningNot yet available