Skip to content
VexAPI

Arioron · VexAPI

Build with
intelligence.


One API for Arioron models: reasoning, vision, documents, tool calling, structured output, embeddings and agents. It speaks the OpenAI and Anthropic formats, so the SDK you already use works by changing one line.

from openai import OpenAI

client = OpenAI(base_url="https://api.arioron.com/v1", api_key=VEX_API_KEY)

response = client.responses.create(
    model="vex-photon-lite",
    input="Explain quantum entanglement to a curious "
          "teenager in two sentences.",
)
print(response.output_text)

Output · captured from a real request

Imagine you have two special coins that are linked, so that if one lands on heads, the other instantly lands on tails, no matter how far apart they are. Quantum entanglement is like that, but with tiny particles whose fates become intertwined, so that measuring one tells you something definite about the other, even if they're light-years away.

Request
req_394f0560e6c7…
Tokens
12 in · 73 out
Latency
1.34 s
Status
200 OK

Scroll

Compatibility

Your existing code already works.

VexAPI implements the OpenAI Chat Completions and Responses formats and the Anthropic Messages format, verified against both official SDKs: streaming, tool calling, structured output, vision and error handling.

Point the client at VexAPI and choose an Arioron model. Nothing else changes.

OpenAI SDKResponses · Chat
- client = OpenAI(base_url="https://api.openai.com/v1")+ client = OpenAI(base_url="https://api.arioron.com/v1")  client.chat.completions.create(
-     model="gpt-…",+     model="perceptix-vex-amber",      messages=messages, tools=tools, stream=True)
Anthropic SDKMessages
- client = Anthropic(base_url="https://api.anthropic.com")+ client = Anthropic(base_url="https://api.arioron.com")  client.messages.create(
-     model="claude-…",+     model="perceptix-vex-amber",      max_tokens=1024, messages=messages, tools=tools)

Models

Arioron models.

Long context, multimodal input and tool use across the lineup. Choose depth or speed per request with reasoning_effort.

perceptix-vex-amber-2.0

Arioron model.

Context
1,048,576
Max output
65,536
Input
text · image
Reasoning
Yes

perceptix-vex-amber-latest

Strongest Amber Model Yet

Context
1,048,576
Max output
65,536
Input
text · image
Reasoning
Yes

photon-lite-latest

Latest series of Photon Lite models

Context
1,048,576
Max output
65,536
Input
text · image
Reasoning
Yes

vex-embed

Text embeddings for search, clustering and retrieval.

Dimensions
up to 3,072
Input
text
Output
normalized vectors

Observability

Nothing happens out of sight.

Every request is logged with its model, latency, status and exact token counts, measured by the model service, never estimated. Every response carries an x-request-id you can paste into your logs.

POST /v1/responses
200 · 12 → 73 tokens · 1.34 s
POST /v1/embeddings
200 · 3 tokens · 0.31 s
POST /v1/messages
404 · model_not_found

Production

Infrastructure, not features.

The parts you only notice when they're missing: scoped keys, predictable limits, safe URL handling and background work that survives a restart.

Keys

Scoped by design

Project-scoped, read-only and expiring keys, accepted in headers only, never in URLs.

Limits

Predictable

Per-key rate limits, reported in headers on every response, with a Retry-After when you hit one.

Network

Safe fetching

Images and pages are fetched from public addresses only. Internal networks are refused.

Jobs

Durable

Batches and agent runs survive restarts and retry automatically. Webhooks are HMAC-signed.