Serverless Inference
Works with your OpenAI client
Point your existing OpenAI client at Berget AI. Same endpoints, same SDKs, and same streaming.
Read the docsapp.ts
const client = new OpenAI({
baseURL: 'https://api.berget.ai/v1',
apiKey: 'YOUR_API_KEY'
});
const res = await client.chat.completions.create({
model: 'gemma-4-31B-it',
messages: [{ role: 'user', content: 'Deploy app' }]
});Models
Pick the right model for the job
Every model we run, with live pricing and capabilities straight from the API. Whichever you choose, your traffic never leaves Sweden.
Loading models…
Simple, transparent pricing
Pay per token, or pick a plan for predictable spend. No egress fees.
* Rate limits are shared per account.
Want to compare pricing across all our models in detail? See pricing
Deployment options
Choose between shared and dedicated deployment, depending on latency needs, traffic patterns, and how much infrastructure control you need.