Files
docs/services/llm-benchmark.md
T

20 KiB
Raw Blame History

LLM Model Benchmark & Service Catalogue

Last updated: 2026-05-23 | Models: 60 | Source: LiteLLM /model/info + live benchmarks

Overview

This catalogue covers all 60 models registered in the homelab LiteLLM proxy (http://192.168.1.40:14000). Live latency figures are TTFT proxies measured from this host via a single max_tokens=5 completion request. Speed tiers are based on live measurements and published inference benchmarks.

Providers at a glance:

Provider Models Notes
Claude (Anthropic via openai-compat) 8 claude-max subscription; aliases included
Mistral API 10 voxtral voice family + codestral + OCR
Gemini API 12 Flash/Pro/embedding families
SAIA (self-hosted GPU cluster) 22 OpenAI-compatible; local GPU inference
Groq 2 Ultra-fast cloud inference
Cerebras 2 Ultra-fast wafer-scale inference
Cohere 2 Embeddings only

Performance Tiers

Tier Symbol Typical TTFT Profile
Ultra-fast 🚀 < 200 ms Groq, Cerebras, cached SAIA small models
Fast 200600 ms Mistral API, Gemini Flash, SAIA mid-size
Standard 🔵 6002 000 ms Claude, Gemini Pro, large API models
Self-hosted 🏠 varies SAIA cluster; latency depends on GPU load & model size

Note: SAIA models with very low latency (< 50 ms) on the trivial benchmark likely hit a cached/KV-prefilled response; real-world TTFT for longer prompts will be higher. Treat SAIA figures as best-case.


Model Catalogue

Chat & Reasoning Models

Model ID Provider Backend Context Vision Tools Cost In $/1M Cost Out $/1M Live Latency Speed Tier Notes
claude-sonnet-4-6 Anthropic openai-compat 200K $3.00 $15.00 2 083 ms 🔵 Flagship balance model
claude-opus-4-7 Anthropic openai-compat 200K $15.00 $75.00 3 218 ms 🔵 Highest capability Claude
claude-haiku-4-5 Anthropic openai-compat 200K $0.80 $4.00 1 453 ms 🔵 Fast + cheap Claude tier
claude-sonnet-4-6-high Anthropic openai-compat 200K $3.00 $15.00 1 746 ms 🔵 Extended thinking variant
claude-opus-4-7-high Anthropic openai-compat 200K $15.00 $75.00 2 285 ms 🔵 Extended thinking variant
voxtral-small-latest Mistral Mistral API 256K 160 ms 🚀 Voice+text multimodal
mistral-small-latest Mistral Mistral API 131K $0.06 $0.18 222 ms 🚀 Cheapest Mistral chat
voxtral-mini-latest Mistral Mistral API 100K ERROR Proxy config error (saia-image-proxy unreachable)
codestral-latest Mistral Mistral API 16K $1.00 $3.00 275 ms Coding-specialised Mistral
mistral-large-latest Mistral Mistral API 262K $0.50 $1.50 309 ms Flagship Mistral chat
pixtral-large-latest Mistral Mistral API 128K $2.00 $6.00 27 ms 🚀 Vision flagship; very low latency (likely cached)
voxtral-mini-realtime-latest Mistral Mistral API 4K ERROR Invalid model per API; realtime/ws endpoint only
gemini-2.5-flash Google Gemini API 65K $0.30 $2.50 501 ms Best-value Gemini; fast + smart
gemini-2.5-flash-lite Google Gemini API 65K $0.10 $0.40 518 ms Lightest + cheapest Gemini
gemini-2.5-pro Google Gemini API 65K $1.25 $10.00 888 ms 🔵 Top Gemini reasoning
gemini-3.1-pro-preview Google Gemini API 65K $2.00 $12.00 1 215 ms 🔵 Next-gen Gemini Pro preview
gemini-3-pro-preview Google Gemini API 65K $2.00 $12.00 1 258 ms 🔵 Gemini 3 Pro preview
gemini-3.1-flash-lite Google Gemini API 482 ms Gemini 3.1 flash lite preview
devstral-2-123b-instruct-2512 Mistral SAIA 4K 157 ms 🏠 123B coding model via SAIA GPU
qwen3-coder-30b-a3b-instruct Alibaba SAIA 32K 5 ms 🏠 MoE coding model; 5 ms = cached
openai-gpt-oss-120b OpenAI SAIA 131K 118 ms 🏠 OpenAI open-weight 120B via SAIA
qwen3-omni-30b-a3b-instruct Alibaba SAIA 16K 178 ms 🏠 Multimodal MoE; voice+vision
llama-3.3-70b-instruct Meta SAIA 131K 168 ms 🏠 Reliable general-purpose 70B
deepseek-r1-distill-llama-70b DeepSeek SAIA 131K 325 ms 🏠 R1 reasoning distill; outputs <think> tokens
qwen3.5-35b-a3b Alibaba SAIA 65K 250 ms 🏠 MoE 35B A3B
qwen3.5-27b Alibaba SAIA 131K 266 ms 🏠 Dense 27B
qwen3.6-35b-a3b Alibaba SAIA 65K 204 ms 🏠 MoE 35B A3B v3.6
apertus-70b-instruct-2509 Apertus SAIA 4K 143 ms 🏠 Small context; general chat
glm-4.7 Zhipu SAIA 128K 210 ms 🏠 GLM-4 series
qwen3-30b-a3b-instruct-2507 Alibaba SAIA 262K 127 ms 🏠 MoE 30B, very large context
gemma-3-27b-it Google SAIA 131K 284 ms 🏠 Gemma 3 27B instruct
internvl3.5-30b-a3b InternLM SAIA 16K 116 ms 🏠 Vision+tools MoE 30B
qwen3.5-122b-a10b Alibaba SAIA 65K 372 ms 🏠 MoE 122B A10B; larger/slower
qwen3.5-397b-a17b Alibaba SAIA 65K 248 ms 🏠 Largest SAIA MoE model
gemma-4-31b-it Google SAIA 8K 177 ms 🏠 Gemma 4 multimodal
meta-llama/llama-4-scout-17b-16e-instruct Meta Groq 8K $0.11 $0.34 442 ms 🚀 Groq-accelerated; 442 ms incl. queue
cerebras-llama-3.1-8b Meta Cerebras 128K $0.10 $0.10 258 ms 🚀 Wafer-scale ~2 000 TPS
cerebras-qwen-3-235b Alibaba Cerebras 199 ms 🚀 235B at wafer-scale speed

Embedding Models

Model ID Provider Backend Context Dimensions Cost $/1M Notes
gemini-embedding-2 Google Gemini API 8K $0.20 Primary Gemini embedding
gemini-embedding-001 Google Gemini API 2K $0.15 Legacy Gemini embedding
text-embedding-ada-002 Google Gemini API 8K $0.20 Alias → gemini-embedding-2
text-embedding-3-small Google Gemini API 8K $0.20 Alias → gemini-embedding-2
text-embedding-3-large Google Gemini API 8K $0.20 Alias → gemini-embedding-2
multilingual-e5-large-instruct Microsoft SAIA 8K 1 024 Self-hosted multilingual; strong for DE/EN RAG
cohere-embed-multilingual-v3 Cohere Cohere API 1K 1 024 $0.10 100+ languages
cohere-embed-english-v3 Cohere Cohere API 1K 1 024 $0.10 English-only; higher EN accuracy

Audio Models (TTS / ASR)

Model ID Provider Backend Type Language Notes
tts-1-de SAIA Piper TTS German Self-hosted German TTS
tts-1 SAIA Kokoro-82M TTS EN + others Self-hosted multilingual TTS
voxtral-mini-tts-latest Mistral Mistral API TTS Multilingual Mistral voice synthesis
saia-whisper SAIA whisper-large-v2 ASR Multilingual Self-hosted transcription
whisper-1 SAIA faster-whisper-large-v3 ASR Multilingual Faster self-hosted transcription
voxtral-mini-transcribe-2507 Mistral Mistral API ASR Multilingual Mistral audio transcription; ctx 16K
whisper-large-v3-turbo Meta Groq ASR Multilingual Groq-accelerated; fastest transcription

Image Models

Model ID Provider Backend Type Notes
saia-flux SAIA FLUX Image gen Self-hosted FLUX; note: garbles text labels
gemini-2.5-flash-image Google Gemini API Image gen ctx 32K; multimodal image generation
saia-image-edit SAIA Qwen-Image-Edit Image edit Image editing/inpainting
mistral-ocr-latest Mistral Mistral API OCR Document OCR; not a chat model

Service Recommendations

Nextcloud Assistant

Smart file/email/calendar assistant, summaries, writing help — multilingual DE/EN

Role Model Reasoning
Primary mistral-small-latest Cheapest API model with tool use, vision, 131K context, and 222 ms TTFT. Handles German natively.
Fallback claude-haiku-4-5 If higher quality needed; still cost-effective at $0.80/$4 and 1 453 ms TTFT.
Alt (local) qwen3.5-27b Free if staying fully on SAIA GPU cluster; 131K context + tools.
model: mistral-small-latest

Karakeep (Bookmarks / Reading)

Summarise articles, extract key points, tag/categorise — no vision required

Role Model Reasoning
Primary gemini-2.5-flash-lite Cheapest API model at $0.10/$0.40, 518 ms TTFT, strong comprehension.
Fallback mistral-small-latest Slightly pricier but faster at 222 ms.
model: gemini-2.5-flash-lite

Home Assistant

Intent recognition, automation triggers, voice pipeline — ultra-low latency critical

Role Model Reasoning
Primary cerebras-llama-3.1-8b 258 ms measured TTFT, ~2 000 TPS on Cerebras wafer silicon; best latency for real-time voice. 128K context, tools.
Fallback mistral-small-latest 222 ms TTFT, API-based, reliable tool calling.
Local alt internvl3.5-30b-a3b 116 ms on SAIA; avoids API cost for high-frequency automations.
model: cerebras-llama-3.1-8b

LobeChat Default

General chat assistant for daily use — balanced quality / speed / cost, vision nice

Role Model Reasoning
Primary gemini-2.5-flash 501 ms, vision, tools, large context, excellent reasoning at $0.30/$2.50. Best all-rounder.
Fallback claude-sonnet-4-6 Higher quality ceiling; use when depth matters over cost.
Free alt qwen3.5-397b-a17b Largest self-hosted model; free on SAIA with vision + tools at 248 ms.
model: gemini-2.5-flash

Code Assistant (Coder / IDE)

Code completion, review, debugging — strong code ability, large context, tools

Role Model Reasoning
Primary claude-sonnet-4-6 Best overall coding + reasoning; 200K context, tool use, reliable output.
Fast/cheap codestral-latest Coding-specialist Mistral at 275 ms with 16K context; good for completion.
Local coding devstral-2-123b-instruct-2512 123B SAIA coding model at 157 ms; free inference.
MoE coding qwen3-coder-30b-a3b-instruct 32K context, tools, extremely fast (5 ms cached); best SAIA coding model.
model: claude-sonnet-4-6   # IDE / review
model: qwen3-coder-30b-a3b-instruct   # local completion

Document OCR / Ingestion

Paperless → Docling → extract text — vision + OCR capable, large context

Role Model Reasoning
Primary mistral-ocr-latest Dedicated OCR endpoint; purpose-built for document text extraction.
Fallback gemini-2.5-pro 65K context, vision, strong at structured extraction from images.
Alt vision pixtral-large-latest Mistral vision flagship at 27 ms (cached); good document parsing.
model: mistral-ocr-latest   # OCR pipeline
model: gemini-2.5-pro        # fallback / complex layouts

Embeddings (Karakeep / LobeChat KB)

Semantic search, RAG, knowledge base — multilingual, high dimensions

Role Model Reasoning
Primary (API) gemini-embedding-2 8K context, $0.20/1M, strong multilingual.
Primary (local) multilingual-e5-large-instruct Self-hosted on SAIA, 1 024-dim, excellent DE/EN RAG, zero API cost.
Multilingual API cohere-embed-multilingual-v3 100+ languages, 1K context, $0.10/1M.
model: multilingual-e5-large-instruct   # local RAG
model: gemini-embedding-2               # API fallback

Note: text-embedding-ada-002, text-embedding-3-small, and text-embedding-3-large are all aliases for gemini-embedding-2 — use the canonical ID to avoid confusion.


Image Generation

ComfyUI complement, quick drafts

Role Model Reasoning
Primary saia-flux Self-hosted FLUX on SAIA GPU; no API cost. Note: avoid text in generated images (garbles).
API alt gemini-2.5-flash-image Gemini multimodal image gen for quick API-based drafts.
Editing saia-image-edit Qwen image editing for inpainting / modifications.
model: saia-flux

TTS (Voice Interfaces)

Read content aloud, voice responses

Role Model Reasoning
German tts-1-de Self-hosted Piper; native German pronunciation.
Multilingual tts-1 Self-hosted Kokoro-82M; covers EN + others, zero cost.
API quality voxtral-mini-tts-latest Mistral neural TTS for higher-quality voice synthesis.
model: tts-1-de    # German HA / Nextcloud voice
model: tts-1       # English / multilingual

Transcription (Meetings / Voice)

Speech to text

Role Model Reasoning
Primary whisper-large-v3-turbo Groq-accelerated; fastest available transcription.
Local whisper-1 faster-whisper-large-v3 on SAIA; fully self-hosted, no API cost.
Fallback saia-whisper whisper-large-v2 on SAIA; slightly older model.
model: whisper-large-v3-turbo   # real-time meetings
model: whisper-1                 # offline / batch

Reasoning / Analysis

Complex problem solving, research, multi-step tasks

Role Model Reasoning
Primary claude-opus-4-7-high Extended thinking mode; highest reasoning capability available.
Cheaper gemini-2.5-pro Strong reasoning at 888 ms, $1.25/$10.00; good for research tasks.
Local reasoning deepseek-r1-distill-llama-70b R1 chain-of-thought via SAIA at 325 ms; outputs <think> tokens.
Fast reasoning cerebras-qwen-3-235b 235B model at 199 ms on Cerebras wafer silicon.
model: claude-opus-4-7-high    # deep analysis
model: deepseek-r1-distill-llama-70b   # local reasoning

Batch / Offline Processing

Non-real-time document processing — cost-optimised, high throughput

Role Model Reasoning
Primary gemini-2.5-flash-lite $0.10/$0.40; cheapest API model with tools + vision.
Free llama-3.3-70b-instruct SAIA self-hosted 70B at 168 ms; 131K context, no API cost.
Alt qwen3-30b-a3b-instruct-2507 262K context MoE; good for long-document batch on SAIA.
model: gemini-2.5-flash-lite   # cost-sensitive API batch
model: llama-3.3-70b-instruct  # free local batch

Aliases & Duplicates

The following model IDs are aliases that route to the same backend model. Use the canonical ID in production to avoid ambiguity:

Alias ID Canonical Model Notes
sonnet claude-sonnet-4-6 Short alias
opus claude-opus-4-7 Short alias
haiku claude-haiku-4-5 Short alias
text-embedding-ada-002 gemini-embedding-2 OpenAI compat alias
text-embedding-3-small gemini-embedding-2 OpenAI compat alias
text-embedding-3-large gemini-embedding-2 OpenAI compat alias

Experimental / Not Yet Validated

The following models returned errors or have unresolved issues in live testing:

Model ID Status Error Detail Action
voxtral-mini-latest Fixed 2026-05-23 Stray api_base: saia-image-proxy:5999 — deleted + re-added clean; now routes to Mistral API
voxtral-mini-realtime-latest 🗑️ Removed 2026-05-23 WebSocket-only realtime endpoint; incompatible with REST completions Removed from LiteLLM; use Mistral WS API directly if needed
mistral-ocr-latest ⚠️ Not benchmarked OCR-mode model; requires document input, not chat completions Use via dedicated OCR pipeline only
voxtral-mini-transcribe-2507 ⚠️ Not benchmarked Audio transcription; not a chat completions model Use via audio transcription endpoint
gemini-2.5-flash-image ⚠️ Not benchmarked Image generation; not a chat completions model Use via images endpoint
saia-flux ⚠️ Not benchmarked FLUX image generation Use via images endpoint
saia-image-edit ⚠️ Not benchmarked Image editing Use via image edit endpoint
tts-1-de ⚠️ Not benchmarked Piper TTS audio output Use via audio/speech endpoint
tts-1 ⚠️ Not benchmarked Kokoro-82M TTS Use via audio/speech endpoint
voxtral-mini-tts-latest ⚠️ Not benchmarked Mistral TTS Use via audio/speech endpoint
saia-whisper ⚠️ Not benchmarked Whisper ASR Use via audio/transcriptions endpoint
whisper-1 ⚠️ Not benchmarked faster-whisper ASR Use via audio/transcriptions endpoint
whisper-large-v3-turbo ⚠️ Not benchmarked Groq Whisper ASR Use via audio/transcriptions endpoint
gemini-3.1-flash-lite ⚠️ Context unknown Preview model; ctx window not documented Monitor Gemini API release notes
cerebras-qwen-3-235b ⚠️ Context unknown Context window not documented in LiteLLM config Check Cerebras API docs

Raw Benchmark Data

All measurements from 2026-05-23. Single max_tokens=5 completion, prompt: "Reply with exactly: ok".

Model ID HTTP Status Latency (ms) Prompt Tokens Completion Tokens
claude-sonnet-4-6 200 2 083 3 4
claude-opus-4-7 200 3 218 6 6
claude-haiku-4-5 200 1 453 10 41
claude-sonnet-4-6-high 200 1 746 3 4
claude-opus-4-7-high 200 2 285 6 6
voxtral-small-latest 200 160 8 2
mistral-small-latest 200 222 20 2
voxtral-mini-latest 500 48
codestral-latest 200 275 13 2
devstral-2-123b-instruct-2512 200 157 8 2
qwen3-coder-30b-a3b-instruct 200 5 13 2
openai-gpt-oss-120b 200 118 74 5
gemini-2.5-flash 200 501 6 1
cerebras-llama-3.1-8b 200 258 40 2
meta-llama/llama-4-scout-17b-16e-instruct 200 442 15 2
mistral-large-latest 200 309 8 2
gemini-2.5-flash-lite 200 518 6 1
qwen3-omni-30b-a3b-instruct 200 178 13 2
llama-3.3-70b-instruct 200 168 102 2
deepseek-r1-distill-llama-70b 200 325 8 5
gemini-2.5-pro 200 888 6 2
pixtral-large-latest 200 27 13 2
voxtral-mini-realtime-latest 400 140
qwen3.5-35b-a3b 200 250 15 5
qwen3.5-27b 200 266 15 5
qwen3.6-35b-a3b 200 204 15 5
apertus-70b-instruct-2509 200 143 66 2
glm-4.7 200 210 9 2
gemini-3.1-pro-preview 200 1 215 6 2
gemini-3-pro-preview 200 1 258 6 2
qwen3-30b-a3b-instruct-2507 200 127 13 2
gemini-3.1-flash-lite 200 482 6 1
gemma-3-27b-it 200 284 14 3
internvl3.5-30b-a3b 200 116 13 2
qwen3.5-122b-a10b 200 372 15 5
qwen3.5-397b-a17b 200 248 15 5
gemma-4-31b-it 200 177 18 2
cerebras-qwen-3-235b 200 199 13 2

Benchmark note on SAIA models with < 50 ms latency (qwen3-coder: 5 ms, pixtral: 27 ms): these figures reflect a KV-cache or pre-warmed response for the trivial prompt. Real-world TTFT for cold prompts will be 100400 ms depending on model size and GPU availability.