VexAPI by Arioron
Build with intelligence
VexAPI is one API for Arioron models: text, vision, documents, tool calling, structured output, embeddings, agents and more. It speaks both the OpenAI and the Anthropic wire formats, so the official SDKs you already use work by changing one line.
base_url, keep your code.
Switch from AnthropicSame SDK, same Messages API.
API referenceEvery endpoint, parameter and event.
Base URL
https://api.arioron.com/v1One platform, three wire formats
Pick whichever matches your code. All three run on the same models with identical capabilities, and every request shows up in the same usage and logs.
| Endpoint | Wire format | Best for |
|---|---|---|
POST /v1/responses | OpenAI Responses | New projects. Built-in web search, server-side conversation state. |
POST /v1/chat/completions | OpenAI Chat Completions | Existing OpenAI integrations and the wider ecosystem of tools built on it. |
POST /v1/messages | Anthropic Messages | Existing Anthropic SDK integrations. |
Quickstart
1. Create an API key on the API keys page and export it:
export VEX_API_KEY="ar_..."
2. Install the SDK you prefer. VexAPI works with the official OpenAI and Anthropic SDKs as-is.
pip install openai
npm install openai
pip install anthropic
3. Make your first request:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.arioron.com/v1",
api_key=os.environ["VEX_API_KEY"],
)
response = client.responses.create(
model="perceptix-vex-amber-2.0",
input="Write a one-sentence bedtime story about a lighthouse.",
)
print(response.output_text)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.arioron.com/v1",
apiKey: process.env.VEX_API_KEY,
});
const response = await client.responses.create({
model: "perceptix-vex-amber-2.0",
input: "Write a one-sentence bedtime story about a lighthouse.",
});
console.log(response.output_text);
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://api.arioron.com",
api_key=os.environ["VEX_API_KEY"],
)
message = client.messages.create(
model="perceptix-vex-amber-2.0",
max_tokens=1024,
messages=[{"role": "user", "content": "Write a one-sentence bedtime story about a lighthouse."}],
)
print(message.content[0].text)
curl https://api.arioron.com/v1/responses \
-H "Authorization: Bearer $VEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "perceptix-vex-amber-2.0",
"input": "Write a one-sentence bedtime story about a lighthouse."
}'
/v1/messages itself, so its base URL is https://api.arioron.com, without /v1.Authentication
Send your key in the Authorization header. The x-api-key header is also
accepted, which is what the Anthropic SDKs send. Keys are never accepted in the URL, because URLs end up
in proxy and access logs.
Authorization: Bearer ar_...
# or
x-api-key: ar_...
Keeping keys safe
- Keys are secrets. Call VexAPI from your server, never from a browser or mobile app.
- Each key belongs to a project, so its usage, logs and rate limits are tracked separately.
- Create read-only keys for dashboards and reporting. They can call
GETendpoints only. - Give keys an expiry and rotate them. Expired or revoked keys stop working immediately.
| Status | Code | Meaning |
|---|---|---|
| 401 | missing_api_key | No key was sent. |
| 401 | invalid_api_key | The key doesn't exist or was revoked. |
| 401 | expired_api_key | The key passed its expiry date. |
| 403 | insufficient_permissions | A read-only key tried to write. |
Models
Pass a model id as model. List models programmatically with GET /v1/models.
| Model | Context | Max output | Capabilities |
|---|---|---|---|
perceptix-vex-amber-2.0 |
1,048,576 | 65,536 | Vision Tools Structured Reasoning Streaming |
perceptix-vex-amber-latestStrongest Amber Model Yet |
1,048,576 | 65,536 | Vision Tools Structured Reasoning Streaming |
photon-lite-latestLatest series of Photon Lite models |
1,048,576 | 65,536 | Vision Tools Structured Reasoning Streaming |
vex-embedGeneral-purpose text embeddings for search, clustering and retrieval. Supports reduced dimensions (e.g. 768, 1536). |
— | 3,072 dims | Embeddings |
Admins can also define model aliases: a stable name you call, routed to the healthiest of several models.
Responses
/v1/responsesThe recommended interface. Send a string or a list of input items and get back typed output items.
VexAPI stores responses by default, so you can continue a conversation by passing
previous_response_id instead of resending history.
response = client.responses.create(
model="perceptix-vex-amber-2.0",
instructions="You are a concise travel assistant.",
input="Three things to do in Lisbon in one day.",
)
print(response.output_text)
# Continue the conversation without resending history
follow_up = client.responses.create(
model="perceptix-vex-amber-2.0",
previous_response_id=response.id,
input="Which of those is best on a rainy day?",
)
print(follow_up.output_text)
const response = await client.responses.create({
model: "perceptix-vex-amber-2.0",
instructions: "You are a concise travel assistant.",
input: "Three things to do in Lisbon in one day.",
});
console.log(response.output_text);
const followUp = await client.responses.create({
model: "perceptix-vex-amber-2.0",
previous_response_id: response.id,
input: "Which of those is best on a rainy day?",
});
console.log(followUp.output_text);
curl https://api.arioron.com/v1/responses \
-H "Authorization: Bearer $VEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "perceptix-vex-amber-2.0",
"instructions": "You are a concise travel assistant.",
"input": "Three things to do in Lisbon in one day."
}'
Request body
modelstringrequiredModel id or alias.
inputstring | arrayrequiredA string, or a list of items: messages (role user, assistant, system or developer, with text, input_image or input_file parts), function_call and function_call_output.
instructionsstringSystem instructions for this response. Not carried over to responses that chain from it.
previous_response_idstringContinue from a stored response. Its full input and output become this request's history.
toolsarrayFunction tools ({"type": "function", "name", "description", "parameters"}) and the built-in {"type": "web_search"}.
tool_choicestring | objectauto (default), none, required, or {"type": "function", "name": ...}.
text.formatobject{"type": "text"} (default), {"type": "json_object"}, or {"type": "json_schema", "name", "schema"}. See structured outputs.
reasoning.effortstringminimal, low, medium or high. See reasoning.
max_output_tokensintegerUpper bound on output tokens, including reasoning tokens.
temperaturenumber0 to 2. Higher is more varied.
top_pnumber0 to 1. Nucleus sampling.
streambooleanStream typed server-sent events. See streaming.
storebooleanDefault true. When false, the response isn't retrievable or chainable, and request/response bodies aren't kept in your logs.
metadataobjectUp to you. Echoed back on the response.
Response object
{
"id": "resp_8f3c…",
"object": "response",
"created_at": 1759567200,
"status": "completed",
"model": "perceptix-vex-amber-2.0",
"output": [
{
"type": "message",
"id": "msg_2b1a…",
"status": "completed",
"role": "assistant",
"content": [{ "type": "output_text", "text": "1. Ride Tram 28…", "annotations": [] }]
}
],
"usage": {
"input_tokens": 24,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens": 162,
"output_tokens_details": { "reasoning_tokens": 41 },
"total_tokens": 186
},
"previous_response_id": null,
"error": null,
"incomplete_details": null
}
status is completed, or incomplete with
incomplete_details.reason set to max_output_tokens or
content_filter. Output items are message, function_call and
web_search_call.
Built-in web search
Add {"type": "web_search"} to tools and the model searches the live web on
its own. VexAPI runs the searches server-side and returns a web_search_call item, with the
query and sources, before the final message. Tokens for every model turn are included in
usage.
response = client.responses.create(
model="perceptix-vex-amber-2.0",
tools=[{"type": "web_search"}],
input="What changed in the latest Python release?",
)
for item in response.output:
if item.type == "web_search_call":
print("searched:", item.action.query)
print(response.output_text)
Retrieve and delete
/v1/responses/{response_id}/v1/responses/{response_id}Stored responses are private to your account. Deleting one also stops it being used as a
previous_response_id.
Chat Completions
/v1/chat/completionsFully compatible with the OpenAI Chat Completions format, so existing code, frameworks and tools built on it work unchanged.
completion = client.chat.completions.create(
model="perceptix-vex-amber-2.0",
messages=[
{"role": "developer", "content": "You answer in one sentence."},
{"role": "user", "content": "Why is the sky blue?"},
],
)
print(completion.choices[0].message.content)
print(completion.usage.prompt_tokens, completion.usage.completion_tokens)
const completion = await client.chat.completions.create({
model: "perceptix-vex-amber-2.0",
messages: [
{ role: "developer", content: "You answer in one sentence." },
{ role: "user", content: "Why is the sky blue?" },
],
});
console.log(completion.choices[0].message.content);
curl https://api.arioron.com/v1/chat/completions \
-H "Authorization: Bearer $VEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "perceptix-vex-amber-2.0",
"messages": [{"role": "user", "content": "Why is the sky blue?"}]
}'
Request body
modelstringrequiredModel id or alias.
messagesarrayrequiredRoles system, developer, user, assistant and tool. User content can mix text, image_url and file parts.
tools / tool_choicearray / string | objectFunction tools. The model returns tool_calls; you run them and send back tool messages. See tool calling.
response_formatobject{"type": "json_object"} or {"type": "json_schema", "json_schema": {"name", "schema"}}.
max_tokensintegerCaps visible output. When set without reasoning_effort, reasoning is kept minimal so short limits don't truncate answers.
max_completion_tokensintegerCaps output including reasoning tokens.
reasoning_effortstringminimal, low, medium or high.
temperature / top_pnumberSampling controls (0–2 and 0–1).
stopstring | arrayUp to 4 stop sequences.
stream / stream_optionsboolean / objectStream chunks. Set stream_options.include_usage for a final usage chunk.
storebooleanDefault false. When true, request and response bodies are kept in your logs for inspection.
n must be 1, and log probabilities and audio input/output aren't supported. seed, user, penalties and logit_bias are accepted but have no effect.Messages
/v1/messages/v1/messages/count_tokensCompatible with the Anthropic Messages API. Point the Anthropic SDK at VexAPI and use an Arioron model id. Errors on these endpoints use Anthropic's error shape, so the SDK raises the exception classes you already handle.
import anthropic
client = anthropic.Anthropic(base_url="https://api.arioron.com", api_key=os.environ["VEX_API_KEY"])
message = client.messages.create(
model="perceptix-vex-amber-2.0",
max_tokens=1024,
system="You only answer in French.",
messages=[{"role": "user", "content": "Say hello."}],
)
print(message.content[0].text, message.usage.output_tokens)
count = client.messages.count_tokens(
model="perceptix-vex-amber-2.0",
messages=[{"role": "user", "content": "How many tokens is this?"}],
)
print(count.input_tokens)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://api.arioron.com", apiKey: process.env.VEX_API_KEY });
const message = await client.messages.create({
model: "perceptix-vex-amber-2.0",
max_tokens: 1024,
messages: [{ role: "user", content: "Say hello." }],
});
console.log(message.content[0].text);
modelstringrequiredModel id or alias.
max_tokensintegerrequiredMaximum visible output tokens.
messagesarrayrequireduser and assistant turns with text, image, document, tool_use and tool_result blocks. Image and document sources can be base64, url or file.
systemstring | arraySystem prompt.
tools / tool_choicearray / objectTools with input_schema; tool_choice types auto, any, tool and none.
thinkingobject{"type": "enabled", "budget_tokens": n} turns on extended reasoning. Off by default, so max_tokens is all visible output.
temperature / top_p / stop_sequences / stream—As in the Anthropic API.
cache_control is ignored. For built-in web search, use Responses.Embeddings
/v1/embeddingsTurn text into vectors for search, clustering and retrieval. Vectors are L2-normalized, so cosine similarity is a dot product.
result = client.embeddings.create(
model="vex-embed",
input=["The quick brown fox", "A fast auburn fox"],
dimensions=768, # optional: smaller vectors, cheaper storage
)
vectors = [d.embedding for d in result.data]
curl https://api.arioron.com/v1/embeddings \
-H "Authorization: Bearer $VEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "vex-embed", "input": "The quick brown fox"}'
modelstringrequiredvex-embed.
inputstring | arrayrequiredOne string or up to 2,048 strings.
dimensionsintegerOutput size, 1 to 3,072. Defaults to 3,072. 768 and 1,536 are good trade-offs.
encoding_formatstringfloat (default) or base64 (little-endian float32). The OpenAI Python SDK uses base64 automatically.
Tool calling
Describe your functions with JSON Schema. When the model wants one, it returns a call with arguments instead of text. You run the function and send the result back, and the model uses it to answer. VexAPI never runs your functions; only you do.
import json
tools = [{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the status of a customer order",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
},
}]
messages = [{"role": "user", "content": "Where is order A-1001?"}]
first = client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments) # {"order_id": "A-1001"}
messages.append(first.choices[0].message)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps({"status": "shipped", "eta": "Oct 9"}), # your function's result
})
final = client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=messages, tools=tools)
print(final.choices[0].message.content)
tools = [{
"type": "function",
"name": "get_order_status",
"description": "Look up the status of a customer order",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}]
first = client.responses.create(model="perceptix-vex-amber-2.0", tools=tools, input="Where is order A-1001?")
call = next(item for item in first.output if item.type == "function_call")
final = client.responses.create(
model="perceptix-vex-amber-2.0",
tools=tools,
previous_response_id=first.id,
input=[{
"type": "function_call_output",
"call_id": call.call_id,
"output": '{"status": "shipped", "eta": "Oct 9"}',
}],
)
print(final.output_text)
Use tool_choice to force a call (required), a specific function, or none at
all. The model may request several calls in one turn; answer each by its id.
Structured outputs
Give a JSON Schema and the output is guaranteed to parse and match it. Use json_object
when you only need valid JSON.
from pydantic import BaseModel
class Weather(BaseModel):
city: str
temp_c: float
completion = client.chat.completions.create(
model="perceptix-vex-amber-2.0",
messages=[{"role": "user", "content": "Make up today's weather for Paris."}],
response_format={
"type": "json_schema",
"json_schema": {"name": "weather", "schema": Weather.model_json_schema(), "strict": True},
},
)
weather = Weather.model_validate_json(completion.choices[0].message.content)
response = client.responses.create(
model="perceptix-vex-amber-2.0",
input="Make up today's weather for Paris.",
text={"format": {
"type": "json_schema",
"name": "weather",
"schema": {
"type": "object",
"properties": {"city": {"type": "string"}, "temp_c": {"type": "number"}},
"required": ["city", "temp_c"],
"additionalProperties": False,
},
}},
)
print(response.output_text) # {"city": "Paris", "temp_c": 15.5}
Images & files
Models marked Vision read images, PDFs and text files alongside your prompt. Provide them as a
public URL, a data: URL, or a file_id from the Files API.
completion = client.chat.completions.create(
model="perceptix-vex-amber-2.0",
messages=[{"role": "user", "content": [
{"type": "text", "text": "What's in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.png"}},
]}],
)
# A PDF you uploaded earlier
response = client.responses.create(
model="perceptix-vex-amber-2.0",
input=[{"role": "user", "content": [
{"type": "input_file", "file_id": "file_abc123"},
{"type": "input_text", "text": "Summarize this contract in five bullets."},
]}],
)
invalid_image_url.Streaming
Set stream: true to receive tokens as they're generated, as server-sent events in each
format's native event style.
# Responses: typed events
with client.responses.stream(model="perceptix-vex-amber-2.0", input="Write a haiku about rain.") as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
final = stream.get_final_response()
# Chat Completions: chunks, plus a usage chunk at the end
for chunk in client.chat.completions.create(
model="perceptix-vex-amber-2.0",
messages=[{"role": "user", "content": "Write a haiku about rain."}],
stream=True,
stream_options={"include_usage": True},
):
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
| Endpoint | Events |
|---|---|
/v1/responses | response.created, response.output_item.added, response.output_text.delta, response.function_call_arguments.done, response.web_search_call.completed, response.completed (or .incomplete / .failed). Every event has a sequence_number. |
/v1/chat/completions | chat.completion.chunk objects, then data: [DONE]. |
/v1/messages | message_start, content_block_start/_delta/_stop, message_delta, message_stop. |
Reasoning
Models marked Reasoning can think before answering, which helps with math, code and
multi-step problems. Reasoning tokens are billed as output and reported separately in
reasoning_tokens.
- Responses:
reasoning: {"effort": "low"}.max_output_tokensincludes reasoning. - Chat Completions:
reasoning_effort.max_completion_tokensincludes reasoning. The oldermax_tokenscaps only visible text and keeps reasoning minimal. - Messages: off unless you pass
thinking, somax_tokensis all visible output.
Use minimal for fast, cheap answers to simple prompts, and high when quality
matters more than latency.
Migrating from OpenAI
Keep the official openai SDK. Change two lines and pick an Arioron model:
from openai import OpenAI
client = OpenAI(
base_url="https://api.arioron.com/v1", # was: https://api.openai.com/v1
api_key=os.environ["VEX_API_KEY"],
)
client.chat.completions.create(model="perceptix-vex-amber-2.0", messages=[...]) # was: gpt-…
client.responses.create(model="perceptix-vex-amber-2.0", input="...")
client.embeddings.create(model="vex-embed", input="...") # was: text-embedding-…
What works the same: Chat Completions and Responses formats, streaming, function calling, JSON mode and
JSON Schema outputs, vision, previous_response_id, embeddings with
dimensions, error classes and status codes. Model ids that don't exist on VexAPI return a
404 that lists the models that do.
Not available yet: n > 1, log probabilities and audio in chat. The
file_search and code_interpreter built-in tools work in the Responses API.
Migrating from Anthropic
import anthropic
client = anthropic.Anthropic(
base_url="https://api.arioron.com", # was: https://api.anthropic.com
api_key=os.environ["VEX_API_KEY"],
)
client.messages.create(model="perceptix-vex-amber-2.0", max_tokens=1024, messages=[...])
What works the same: the Messages format, system prompts, streaming, tool use with
tool_use/tool_result, images and PDFs, count_tokens, extended
thinking, and errors that map to the SDK's exception classes. Not available on this endpoint: server
tools and prompt caching.
Files
/v1/files/v1/files/v1/files/{file_id}/v1/files/{file_id}/content/v1/files/{file_id}Upload documents and images to reference by file_id in model inputs, batches and
fine-tuning. Up to 25 MB per file. Supported types: PDF, TXT, DOCX, CSV, JSON, PNG, JPEG, GIF, WebP,
MP3, WAV, MP4 and MOV.
uploaded = client.files.create(file=open("contract.pdf", "rb"), purpose="assistants")
print(uploaded.id) # file_…
content = client.files.content(uploaded.id).read() # download bytes, e.g. batch results
Agents
/v1/agents/v1/agents/{agent_id}/runs/v1/agents/{agent_id}/runs/{run_id}/v1/agents/{agent_id}/runs/{run_id}/cancelAn agent is a saved model, instructions and set of built-in tools. Each run loops through model calls and tool calls until it has an answer, and every step is recorded in the run's trace so you can see exactly what it did. Runs can execute synchronously or in the background, with step limits, timeouts and cancellation.
# Create an agent that may only use web search
curl https://api.arioron.com/v1/agents -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
"name": "Research assistant",
"model": "perceptix-vex-amber-2.0",
"instructions": "Research the question, then answer in five bullets with sources.",
"tools": ["web_search"],
"max_steps": 8
}'
# Run it in the background
curl https://api.arioron.com/v1/agents/agent_…/runs -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
"input": "What were the biggest announcements at this year's PyCon?",
"background": true
}'
# → 202 {"id": "run_…", "status": "queued", …}
# Poll for the result and the full step-by-step trace
curl https://api.arioron.com/v1/agents/agent_…/runs/run_… -H "Authorization: Bearer $VEX_API_KEY"
name, modelstringrequiredDisplay name and model id.
instructionsstringThe agent's system prompt.
toolsarrayBuilt-in tools the agent may use: web_search, web_scrape, extract_article, get_page_metadata, get_weather, get_current_time, calculate, convert_units. Calls to any other tool are refused.
max_steps / timeout_secondsintegerDefaults 8 and 120. The run stops when either is reached.
memory_enabledbooleanPull relevant memories scoped to this agent into each run.
Run status moves from queued to running to completed,
failed or cancelled. Pass conversation_id to run an agent on
top of a stored conversation. Subscribe to
agent_run.completed with a webhook instead of polling.
Conversations
/v1/conversations/v1/conversations/{conversation_id}/v1/conversations/{conversation_id}/messages/v1/conversations/{conversation_id}Persistent message threads stored by VexAPI. Post a user message with
generate_response: true and a model to append the assistant's reply in the
same call.
curl https://api.arioron.com/v1/conversations/conv_…/messages -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
"role": "user",
"content": "What is our refund window?",
"generate_response": true,
"model": "perceptix-vex-amber-2.0"
}'
Memory
/v1/memories/v1/memories/v1/memories/{memory_id}/v1/memories/{memory_id}Long-term facts your app or agents can recall by meaning, not keywords. Scope each memory to your
account (user), a project, an agent or a
conversation; agents with memory_enabled retrieve their own memories
automatically.
curl https://api.arioron.com/v1/memories -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
"content": "The customer prefers email over phone calls.",
"scope": "agent",
"owner_ref": "agent_…"
}'
Batch
/v1/batches/v1/batches/{batch_id}/v1/batches/{batch_id}/cancelRun thousands of requests in the background. Upload a JSONL file with one request per line, create a batch, and download a JSONL file of results when it finishes. Each line is processed exactly like the live endpoint, and one bad line doesn't fail the batch.
{"custom_id": "q1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "perceptix-vex-amber-2.0", "messages": [{"role": "user", "content": "Classify: 'love it'"}]}}
{"custom_id": "q2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "perceptix-vex-amber-2.0", "messages": [{"role": "user", "content": "Classify: 'never again'"}]}}
batch_file = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
batch = client.batches.create(input_file_id=batch_file.id, endpoint="/v1/chat/completions",
completion_window="24h")
batch = client.batches.retrieve(batch.id)
if batch.status == "completed": # failures: batch.request_counts.failed
results = client.files.content(batch.output_file_id).text
Supported endpoints: /v1/chat/completions, /v1/embeddings and
/v1/moderations. Status moves through validating, queued and
running to completed, cancelled or failed. A
completed batch can still contain failed lines; request_counts.failed says how many. Each output line is
{"custom_id", "status_code", "response", "error"}. Token usage is billed to the key that
created the batch.
Webhooks
/v1/webhooks/v1/webhooks/v1/webhooks/{webhook_id}/v1/webhooks/{webhook_id}Get a signed HTTP POST when work finishes instead of polling. Events: agent_run.completed,
batch.completed, or * for everything. Failed deliveries are retried with
backoff, up to 3 attempts.
The signing secret is shown once, in the create response. Verify every delivery before trusting it:
import hashlib, hmac
def verify(raw_body: bytes, signature: str, secret: str) -> bool:
expected = hmac.new(secret.encode(), raw_body, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, signature)
# Flask example
@app.post("/hooks/vexapi")
def vexapi_hook():
if not verify(request.get_data(), request.headers["X-VexAPI-Signature"], WEBHOOK_SECRET):
abort(401)
event = request.get_json() # {"type": "batch.completed", "created": …, "data": {…}}
return "", 204
import crypto from "node:crypto";
function verify(rawBody, signature, secret) {
const expected = crypto.createHmac("sha256", secret).update(rawBody).digest("hex");
return crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(signature));
}
Background jobs
/v1/jobs/v1/jobs/{job_id}/v1/jobs/{job_id}/cancelBatches, background agent runs, evaluations and webhook deliveries all run on one durable job queue.
These endpoints give a single view of that work: status, progress, attempts and errors. Filter with
?type= and ?status=. Jobs survive restarts: if a worker stops mid-job, the
job is picked up again automatically.
Evaluations
/v1/evals/v1/evals/{eval_id}/runs/v1/evals/{eval_id}/runs/{run_id}Measure a model against your own test cases before you ship a prompt or switch models. Grade with
exact_match, contains, or llm_judge, which uses a model to score
each answer from 0 to 1 against the expected output. Runs report a pass rate plus per-case output,
score and latency.
curl https://api.arioron.com/v1/evals -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" -d '{
"name": "Capital cities",
"grader_type": "contains",
"cases": [
{"input": "Capital of Australia?", "expected_output": "Canberra"},
{"input": "Capital of Canada?", "expected_output": "Ottawa"}
]
}'
curl https://api.arioron.com/v1/evals/eval_…/runs -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "perceptix-vex-amber-2.0"}'
File search
/v1/vector_stores/v1/vector_stores/{id}/files/v1/vector_stores/{id}/searchLet models answer from your documents. Add uploaded files to a vector store: VexAPI
extracts their text (PDF, Word, Markdown, text, CSV, JSON, HTML), splits it into overlapping chunks and
indexes them by meaning. Then search the store directly, or give the file_search tool to
the Responses API and the model searches whenever it needs to. Indexing runs in the background; a
file's status becomes completed (or failed with a reason).
file_idsarrayCreate: files to index straight away.
attributesobjectAdding a file: key/value pairs you can filter searches by.
querystringrequiredSearch: what to look for.
max_num_resultsintegerSearch: 1–50 (default 10).
filtersobjectSearch: {"type": "eq", "key": "team", "value": "hr"}, comparisons ne gt gte lt lte, or {"type": "and"|"or", "filters": [...]}.
ranking_options.score_thresholdnumberSearch: drop results scoring below this (0–1).
store = client.vector_stores.create(name="Handbook", file_ids=[handbook.id])
response = client.responses.create(
model="perceptix-vex-amber-2.0",
input="How many vacation days do new hires get?",
tools=[{"type": "file_search", "vector_store_ids": [store.id]}],
include=["file_search_call.results"], # optional: return the matched passages
)
print(response.output_text)
curl https://api.arioron.com/v1/vector_stores/$STORE_ID/search -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
-d '{"query": "vacation policy", "max_num_results": 3}'
Code interpreter
/v1/containers/v1/containers/{id}/files/v1/containers/{id}/execute/v1/containers/{id}/files/editModels can write and run Python to calculate, analyse data and create files. Code runs in a secure sandbox (Python 3.14 compiled to WebAssembly): it has no internet access, can't see anything outside its own workspace, and every run is limited in time (30 s by default, up to 120 s) and memory (256 MB). The standard library is available; numpy and pandas are not.
A container is that workspace. Upload files into it, run code, edit files in place, and
download what the code produced. Containers are private to you and expire after 20 idle minutes
(configurable with expires_after.minutes).
containerstring | objectResponses tool: "auto" (a fresh container per request), {"type": "auto", "file_ids": [...]} to preload files, or an existing container id.
codestringrequired/execute: Python source. Relative paths are in /workspace. Returns stdout, stderr, exit_code, status (completed, failed or timeout) and the files it wrote.
opstringrequired/files/edit: view (with optional view_range), create (file_text), str_replace (old_str must appear exactly once, new_str), insert (insert_line, new_str), append or delete.
response = client.responses.create(
model="perceptix-vex-amber-2.0",
input="Using sales.csv, which region grew fastest from Jan to Feb? Save the growth table as growth.csv.",
tools=[{"type": "code_interpreter", "container": {"type": "auto", "file_ids": [sales_csv.id]}}],
include=["code_interpreter_call.outputs"],
)
print(response.output_text)
# Download files the code wrote:
call = next(o for o in response.output if o.type == "code_interpreter_call")
data = client.containers.files.content.retrieve("growth.csv", container_id=call.container_id)
curl https://api.arioron.com/v1/containers/$CONTAINER_ID/execute -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
-d '{"code": "import statistics\nprint(statistics.median([3, 1, 4, 1, 5]))"}'
# {"status": "completed", "exit_code": 0, "stdout": "3\n", "stderr": "", "files": [], ...}
execute_python, create_file, read_file, edit_file and list_files tools and it works in your own private workspace.Realtime voice
/v1/realtime (WebSocket)Talk to a model with your voice, and hear it answer, over one WebSocket. With vex-realtime,
VexAPI detects when you stop talking, the model listens to what you said, and the reply is spoken
with any library voice, sentence by sentence as it's written. The protocol follows OpenAI's
Realtime API: stream microphone audio as input_audio_buffer.append events and receive
spoken replies as response.output_audio.delta, with live transcripts of both sides. The
model answers when you stop talking, and talking over it interrupts the reply
(input_audio_buffer.speech_started tells your client to stop playback). Audio is 16-bit
PCM, mono, 24 kHz, base64-encoded, in both directions.
modelqueryvex-realtime (default). gpt-realtime is accepted as an alias.
session.instructionsstringSystem instructions for the conversation.
session.audio.output.voicestringAny library voice (OpenAI voice names are mapped). Default af_heart.
session.output_modalitiesarray["audio"] (default) or ["text"] for text replies.
session.audio.input.turn_detectionobject | null{"type": "server_vad", "silence_duration_ms": 700} (default) or null to control turns yourself with input_audio_buffer.commit and response.create.
// Server-side (Node 22+). In a browser, use ?api_key= only with short-lived keys.
const ws = new WebSocket("wss://api.arioron.com/v1/realtime?model=vex-realtime",
{ headers: { Authorization: `Bearer ${process.env.VEX_API_KEY}` } });
ws.onopen = () => ws.send(JSON.stringify({
type: "session.update",
session: { instructions: "You are a friendly travel agent.", audio: { output: { voice: "bm_george" } } },
}));
// Send microphone audio (PCM16, 24 kHz) as it arrives:
function onMicChunk(pcm16) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: Buffer.from(pcm16).toString("base64") }));
}
ws.onmessage = ({ data }) => {
const event = JSON.parse(data);
if (event.type === "response.output_audio.delta") playPcm(Buffer.from(event.delta, "base64"));
if (event.type === "response.output_audio_transcript.delta") process.stdout.write(event.delta);
if (event.type === "input_audio_buffer.speech_started") stopPlayback(); // user interrupted
};
output_modalities to ["text"], then send conversation.item.create with an input_text message and response.create to get response.output_text.delta events.Audio
Speech in both directions, with the same request shapes as OpenAI's audio API, so the OpenAI SDKs work unchanged.
Voices
/v1/voices/v1/voices/{voice_id}/previewBrowse and listen to every voice, and try your own text, on the
Voices page in the console. There are two kinds of
voice, listed together by GET /v1/voices (filter with ?category=
and ?language=):
- Library voices (38): natural open-source voices generated on VexAPI's own servers,
in English (American and British), Spanish, French, Hindi, Italian and Portuguese. Fast, with no
usage quota. Ids look like
af_heart: the first letter is the language, the second the gender. - Expressive voices (30), such as
KoreandPuck: multilingual, and they followinstructionsabout tone and emotion. A little slower.
Each voice has a preview_url that returns a short MP3 sample.
Text to speech
/v1/audio/speechmodelstringvex-speech (default). tts-1, tts-1-hd and gpt-4o-mini-tts are accepted as aliases.
inputstringrequiredText to speak, up to 4,096 characters.
voicestringA voice id from the voice library, such as af_heart (default), bm_george or ff_siwis; an expressive voice such as Kore or Puck; or an OpenAI voice name (alloy, nova, onyx, …), which maps to a library voice.
instructionsstringExpressive voices only: how to say it (tone, emotion, accent, pace). For example "Speak warmly, like a flight attendant."
speednumber0.25 to 4.0.
response_formatstringmp3 (default), wav, or pcm (16-bit, 24 kHz, mono).
speech = client.audio.speech.create(
model="vex-speech",
voice="Kore",
input="Your order has shipped and will arrive on Thursday.",
instructions="Friendly and upbeat.",
)
speech.write_to_file("shipped.mp3")
curl https://api.arioron.com/v1/audio/speech -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "vex-speech", "voice": "Puck", "input": "Hello there!"}' --output hello.mp3
Speech to text
/v1/audio/transcriptions/v1/audio/translationsUpload audio as multipart/form-data (mp3, mp4, m4a, wav, webm, ogg, flac, aac or aiff,
up to 20 MB). Transcriptions keep the spoken language; translations return English.
filefilerequiredThe audio to transcribe.
modelstringvex-transcribe (default). whisper-1 and gpt-4o-transcribe are accepted as aliases.
languagestringISO 639-1 code of the spoken language, if you know it.
promptstringNames, terms or spellings the transcript should respect.
response_formatstringjson (default), text, srt, vtt, or verbose_json (adds language, duration and timestamped segments).
with open("meeting.m4a", "rb") as audio:
transcript = client.audio.transcriptions.create(model="vex-transcribe", file=audio)
print(transcript.text)
curl https://api.arioron.com/v1/audio/transcriptions -H "Authorization: Bearer $VEX_API_KEY" \
-F file=@meeting.m4a -F response_format=srt
Images
/v1/images/generations/v1/images/editsGenerate images from a prompt, or edit images you upload (change a background, add or remove objects, combine up to four images). Request and response shapes match OpenAI's Images API.
promptstringrequiredWhat to create or change.
modelstringvex-image (default). gpt-image-1, dall-e-3 and dall-e-2 are accepted as aliases.
ninteger1 to 4 images.
sizestringFor example 1024x1024, 1792x1024 or 1024x1792. Mapped to the nearest supported aspect ratio.
aspect_ratiostringInstead of size: 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9 and more.
response_formatstringb64_json (default) or url. URLs stay valid for one hour.
image / maskfileEdits only (multipart): one to four image files, and an optional PNG mask.
import base64
result = client.images.generate(model="vex-image", prompt="A watercolor fox in a snowy forest", size="1792x1024")
with open("fox.png", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))
503 capability_not_enabled until it is.Moderations
/v1/moderationsCheck text for harassment, hate, sexual and dangerous content before you show or act on it.
curl https://api.arioron.com/v1/moderations -H "Authorization: Bearer $VEX_API_KEY" -H "Content-Type: application/json" \
-d '{"input": "Text to check"}'
# → {"results": [{"flagged": false, "categories": {"harassment": false, …}}]}
Browser sessions
/v1/browser/sessions/v1/browser/sessions/{session_id}/actionsA real browser (Chromium) for your app or agent. Pages run JavaScript, and each session keeps its own cookies, history and open page between actions, so you can sign in, click through, fill forms and read the result. Sessions close after 15 minutes idle; the next action reopens the last page.
urlstringCreate: open this page straight away.
allowlistarrayCreate: domains the session may navigate to, for example ["example.com"]. Subdomains are included.
typestringrequiredAction: navigate, click, double_click, hover, type, select, key, scroll, wait, back, forward, reload, elements, extract, screenshot, evaluate, drag.
ref / selector / textint / stringWhat to act on: a ref number from elements (most reliable), a CSS selector, or visible text. Mouse actions also take x and y.
formatstringextract: markdown (default), text, html or links.
screenshotbooleanAny action: also return a PNG of the result. screenshot itself takes full_page.
import os
import requests
API = "https://api.arioron.com/v1"
H = {"Authorization": f"Bearer {os.environ['VEX_API_KEY']}"}
session = requests.post(f"{API}/browser/sessions", headers=H, json={"url": "https://news.ycombinator.com"}).json()
act = lambda **a: requests.post(f"{API}/browser/sessions/{session['id']}/actions", headers=H, json=a).json()
elements = act(type="elements")["result"]["elements"] # [{"ref": 0, "tag": "a", "label": "new", ...}, ...]
act(type="click", text="past")
page = act(type="extract", format="markdown")["result"]
print(page["title"], page["content"][:500])
requests.delete(f"{API}/browser/sessions/{session['id']}", headers=H)
Computer use
/v1/computer/sessions/v1/computer/sessions/{session_id}/actionsA screen your model controls with the mouse and keyboard, for computer-use agents. The environment is a browser window. Every action returns a fresh screenshot, so a vision model can look, decide the next coordinates, and act, in a loop.
displayobjectCreate: {"width": 1280, "height": 800} (320–2560 each).
start_urlstringCreate: page to open first.
permissionsarrayCreate: limit the session to some of screenshot, mouse, keyboard, navigate. Default: all.
typestringrequiredAction: screenshot, click, double_click, move, drag, scroll, type, keypress, wait, navigate.
x, yintegerScreen coordinates for mouse actions. click also takes button; drag takes a path of points; scroll takes scroll_x and scroll_y.
text / keysstring / arraytype takes text; keypress takes keys, such as ["CTRL", "L"] or "Enter".
session = requests.post(f"{API}/computer/sessions", headers=H,
json={"display": {"width": 1024, "height": 768}, "start_url": "https://www.wikipedia.org"}).json()
shot = session["screenshot"]["data"] # base64 PNG: show it to your model
act = lambda **a: requests.post(f"{API}/computer/sessions/{session['id']}/actions", headers=H, json=a).json()
act(type="click", x=450, y=493) # coordinates your model chose from the screenshot
act(type="type", text="Mount Fuji")
result = act(type="keypress", keys=["ENTER"])
print(result["current_url"]) # https://en.wikipedia.org/wiki/Mount_Fuji
Model aliases
/v1/model_aliases/v1/model_aliasesAn alias is a model name you control, like vex-fast, that routes to an ordered list of
real models. Call the alias in any endpoint. If the first model is unhealthy (a high recent error
rate), requests fail over to the next. Use routing_strategy: "latency" to prefer the
fastest target. Creating aliases requires an admin key.
Usage & logs
/v1/usage/v1/logs/v1/logs/{request_id}Every request is logged with its endpoint, model, status, latency and exact token counts. Usage
aggregates them by day and by model (?days=30). Logs accept ?endpoint=,
?model= and ?limit=. Both are also on the
Usage and
Logs pages.
Projects
/v1/projects/v1/projects/v1/projects/{project_id}Projects group API keys, files, agents and usage, so separate apps or environments stay separate. Resources created with a project-scoped key belong to that project.
Errors
Errors use standard HTTP status codes and one JSON shape. param names the field that
caused the problem when there is one, and request_id identifies the request in your logs.
{
"error": {
"type": "invalid_request_error",
"code": "model_not_found",
"message": "The model 'gpt-4o' does not exist or is not available. Available models: perceptix-vex-amber-2.0, …",
"param": "model",
"request_id": "req_6b1f0c…"
}
}
| Status | Type | What to do |
|---|---|---|
| 400 | invalid_request_error | Fix the request. param and message say what's wrong. |
| 401 | authentication_error | Check the API key. |
| 403 | permission_error | The key can't do this, for example a read-only key writing. |
| 404 | not_found_error | The resource, model or route doesn't exist. |
| 429 | rate_limit_error | Wait for Retry-After seconds, then retry. |
| 500, 502 | server_error | Retry with backoff. Include the request ID if you contact support. |
| 501 | invalid_request_error / provider_unavailable | The capability isn't available yet. See availability. |
| 503 | overloaded_error | The model is temporarily over capacity. Retry shortly. |
The official SDKs retry 429 and 5xx responses automatically.
Rate limits
Limits apply per API key, per minute and per day, and depend on your plan. Read requests
(GET) don't count. Every counted response reports where you stand:
| Header | Meaning |
|---|---|
x-ratelimit-limit-requests | Requests allowed per minute. |
x-ratelimit-remaining-requests | Requests left in the current minute. |
x-ratelimit-limit-requests-day | Requests allowed per day. |
x-ratelimit-remaining-requests-day | Requests left today. |
Retry-After | On a 429: seconds until you can retry. |
Request IDs
Every response has an x-request-id header, and errors repeat it in the body. Look it up on
the Logs page to see the request's model, latency,
token counts and error. x-processing-ms reports how long VexAPI spent on the request.
raw = client.chat.completions.with_raw_response.create(model="perceptix-vex-amber-2.0", messages=[...])
print(raw.headers["x-request-id"]) # req_…
completion = raw.parse()
Availability
What's live today. Capabilities marked not yet available have their endpoints in place and
return 501 provider_unavailable, so you can build against them now and they'll start
working without code changes.
| Capability | Status |
|---|---|
| Text generation, streaming, tool calling, structured output | Available |
| Vision, PDFs and text files | Available |
| Built-in web search (Responses) | Available |
| Embeddings | Available |
| Agents, conversations, memory | Available |
| Batch, webhooks, background jobs | Available |
| Evaluations | Available |
| Text-to-speech: 38 library voices + 30 expressive voices | Available |
| Speech-to-text and translation | Available |
| Moderations | Available |
| Browser sessions (JavaScript, clicks, forms, screenshots) | Available |
| Computer use | Available |
| File search (vector stores) | Available |
| Code interpreter and file editing (sandboxed) | Available |
| Realtime voice | Available |
| Image generation and editing | Not yet available |
| Fine-tuning | Not yet available |