A - Private
SOVAPI STRUCTURE
Live measured from this run.
Run the same prompt. Compare latency, throughput, quality and real request cost.
A - Private
Live measured from this run.
B - Challenger
Live measured from this run.
Neutral - run the reference prompt to score accuracy.
Run the same prompt to compare time to first token, throughput and request cost. Results are measured for each request.
Results published by the model developers, according to their evaluation protocols.
Qwen research Google Gemma Anthropic publications DeepSeek publications
Compare your own workloads in the Arena.
Gemma 4 31B BF16 · H200
$0.12 · $0.38Per million input and output tokens. No cache-read tariff.
Qwen 3.8-27B FP8 · H200
$0.12 · $0.38Per million input and output tokens.
Qwen cache read
$0.04 cachedPer million cached input tokens.
Trial grant
1 BTrial tokens for developers starting from API ACCESS.
Reference / expected values. SovAPI cost is recalculated from live token usage. Challenger cost remains unavailable when the official API response does not expose a monetary amount.
| Metric | SovAPI | Challenger | How it is shown |
|---|---|---|---|
| Gemma 4 31B BF16 · H200 | $0.12 / M input · $0.38 / M output | -- when unavailable | No cache-read tariff |
| Qwen 3.8-27B FP8 · H200 | $0.12 / M input · $0.38 / M output | -- when unavailable | Model-specific live cost |
| Qwen cache read | $0.04 cached | -- when unavailable | Never inferred from uncached input |
| Whisper Large v3 | $0.00048 / audio minute | -- | Transcription in 99+ languages |
| Kokoro | $0.62 / M characters | -- | FR/EN speech synthesis; other languages on request |
| BGE-M3 | $0.01 / M input tokens | -- | Embeddings |
| Per-request cost | Calculated from real tokens | -- when unavailable | Only a server-reported amount is displayed |
| Trial grant | 1 B tokens | None | API ACCESS onboarding |
| Billing | Measured usage | Provider-specific | No SovAPI commitment |
Step 1
Use GitHub or Google OAuth, then generate your trial SovAPI key.
Create my API keyStep 2
Keep your OpenAI SDK, replace only base_url and api_key.
Step 3
One public model, automatic routing by modality.
Pay-as-you-go
| Model / service | Usage price | Details |
|---|---|---|
| Gemma 4 31B BF16 · H200 | $0.12 in · $0.38 out | Text, vision and extraction |
| Qwen 3.8-27B FP8 · H200 | $0.12 in · $0.38 out · $0.04 cached | Code, reasoning, vision and tool calling |
| Whisper Large v3 | $0.00048 / audio minute | Transcription, 99+ languages |
| Kokoro | $0.62 / M characters | Speech synthesis FR/EN · Other languages on request |
| BGE-M3 | $0.01 / M input tokens | Embeddings |
EU hosting · GDPR native · No data training · OpenAI-compatible API
curl https://api.sovinfra.ai/v1/chat/completions \
-H "Authorization: Bearer $SOVINFRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"sovapi","messages":[{"role":"user","content":"Hello"}]}'
Use sovapi for automatic routing, or pin a generation model directly:
gemma-4-31b for text and vision, and qwen3.8-27b for code, reasoning, vision
and tool calling. Whisper, BGE-M3 and Kokoro remain available through SovAPI.
Direct model call
Text, vision and structured extraction.
curl https://api.sovinfra.ai/v1/chat/completions \
-H "Authorization: Bearer $SOVINFRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gemma-4-31b","messages":[{"role":"user","content":"Summarize this request."}]}'
Direct model call
Code, reasoning, vision and tool calling.
curl https://api.sovinfra.ai/v1/chat/completions \
-H "Authorization: Bearer $SOVINFRA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Write a Python health check."}]}'
from openai import OpenAI
client = OpenAI(
base_url="https://api.sovinfra.ai/v1",
api_key="sk-sov-...",
)
response = client.chat.completions.create(
model="sovapi",
messages=[{"role": "user", "content": "Explain SovAPI in one sentence."}],
)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.sovinfra.ai/v1",
apiKey: process.env.SOVINFRA_API_KEY,
});
Private inference, models hosted in Europe, server-side secrets, and logs without prompts or keys. Legal notices and terms are available from the footer.