Custom models
A custom model wraps a catalogue model with your system prompt, an optional knowledge base and optional conversation memory, behind its own model id.
What a custom model is
- You create it in the dashboard: pick a base model from the catalogue (default
openai/gpt-4o-mini), write a system prompt (up to 10,000 characters), optionally upload documents and choose a memory mode. - You call it like any model, on
POST /v1/chat/completionsonly (streaming included)./v1/messages,/v1/responsesand the other endpoints do not resolve custom ids. - It is private: only its owner, or members of the organization that owns it, can call it with their API keys.
Model id and versions
| model value | What runs |
|---|---|
| acme/support-bot | The current (live) configuration |
| acme/support-bot:1.0 | The saved version 1.0 |
| acme/support-bot-v1 or acme/support-bot-1.0 | Same as :1.0 (version suffix) |
| acme/support-bot:latest | The version tagged latest |
- The id is
<namespace>/<name>: the namespace is your organization's slug (or your handle when you are not in an organization); the name uses lowercase letters, digits and hyphens. - Version tags look like
1.0,2.3orlatest;v2and2are read as2.0. - Avoid names that end in
-<number>(for examplebot-2): the suffix is read as a version. - If a key restricts
allowed_models, list the custom id or its base model.
Calling a custom model
……The response is a normal chat completion. Its model field is the base model that answered, not the custom id.
What the gateway adds to your request
Before calling the base model, the gateway prepends, in this order:
- Your system prompt (from the live configuration or from the pinned version).
- Knowledge-base excerpts relevant to the last user message, as a second system message, when the model has documents.
- The stored conversation history for your
session_id, when memory is enabled. - Then the messages you sent.
Conversation memory (session_id)
- Send
session_id(a string you choose) in the body, or through the OpenAI SDK'sextra_body. Without it, nothing is remembered. - With memory enabled on the model, the gateway replays the last 20 messages of that session and, after a successful answer, stores your messages plus the assistant reply.
- History expires after the model's memory TTL: 24 hours by default, configurable from 5 minutes to 7 days.
- A session is shared by everyone who can call the model with the same
session_id: use one unguessable id per end user and conversation. - To start fresh, use a new
session_id. Failed or interrupted streams are not saved.
Knowledge base (RAG)
- Upload up to 20 documents (10 MB each) in the dashboard. Text is split into chunks of about 1,000 words with 200 words of overlap and embedded with
openai/text-embedding-3-small; the embedding of your documents is billed to the model owner. - For each request, the last user message is used as the query: up to 5 matching chunks are added (the best 3 when none is close enough).
- Images and audio files are stored as attachments but are not searched.
- Retrieval itself is not billed separately, but the excerpts are part of the prompt and count as input tokens.
Billing
A custom model is billed at its base model's price, on the real token usage of the enriched request: your messages plus the system prompt, knowledge-base excerpts and memory the gateway added. Keep prompts and memory short to keep cost down. Usual billing rules apply (Pricing and billing).
Errors
| HTTP | error.code | When |
|---|---|---|
| 404 | model_not_found | Unknown custom id or version |
| 503 | model_unavailable | The custom model is disabled |
| 403 | model_not_allowed | You are not the owner or a member of the owning organization, or the key's allowed_models excludes it |
| 401 | invalid_api_key | Missing or invalid key |
After these checks, the base model's errors apply as usual (402, 429, 503, …), and their messages name the base model. See API errors.