Security
No shared GPUs
On the shared API your requests are isolated logically, but the silicon is shared with other customers. Dedicated inference removes that question entirely. Your model, your hardware, nothing else on the machine.
Performance
Nobody else's peak is your problem
A shared model can be busiest exactly when you need it most. With dedicated inference nobody else is on the model. Capacity is constant, latency is predictable, and you set the priorities.
Price
When a dedicated GPU makes sense
At low volume, paying per token is your cheapest option. Past a few billion tokens a month, reserving the GPUs costs less than the metered API. Pick a model to see where the lines cross. Smaller models start around €2,500 per GPU per month.
Assumes 3 input tokens per output token at today's API prices. Your own input/output mix decides where the lines cross for you.
To get an exact quote, contact sales
Expertise
Hardware is the least of it
KV-cache tuning, quantisation, throughput optimisation, monitoring, and capacity planning are all included. You keep control, we carry the toil.
Your model
Any model, including yours
A legacy model we no longer host in the shared fleet, a fine-tune you own, or something in between. If our hardware can run it, we'll deploy it for you.
SLA
Guaranteed delivery, in writing
Public endpoints are best effort. Dedicated inference comes with a contracted SLA covering agreed availability, support response times, and remedies if we miss them.
Deployment options
Choose between shared and dedicated deployment, depending on latency needs, traffic patterns, and how much infrastructure control you need.