Paddock Benchmarks
We benchmark Paddock against other engines and publish everything: the numbers, the setup, and the raw data. Pick a model below to see how every engine ran it on the same GPU.
How We Test
Same Machine, One at a Time
Every engine runs on the same machine and GPU, never at the same time, driven by the same client code. We take the best of the timed rounds after a warmup.
Unique Inputs
Every test input is unique, so an engine can't reuse earlier work and look faster than it really is.
Best Settings, Newest Builds
Each engine runs its latest release with the fastest setup it supports on the hardware.
Benchmarks by Model
Text Generation
Gemma 4 31B2026-07-20
up to 9.4xvs llama.cpp
up to 2.2xvs vLLM / SGLang
RTX 6000 Max-Q
gpt-oss2026-07-20
up to 3.2xvs llama.cpp
up to 2.7xvs vLLM / SGLang
RTX 6000 Max-Q · RTX 5090
Qwen3.5 9B2026-07-20
up to 3.5xvs llama.cpp
up to 2.0xvs vLLM / SGLang
RTX 5090 · RTX 6000 Max-Q
Qwen3.6 35B-A3B2026-07-20
up to 4.4xvs llama.cpp
up to 2.0xvs vLLM / SGLang
RTX 6000 Max-Q
Embeddings & Reranking
Run your own numbers
Paddock ships a benchmark harness, so you can check our claims on your own machine.