← All benchmarks

Qwen3 Embedding

Paddock vs vLLM vs SGLang vs llama.cpp · Embeddings · RTX 5090 (32 GB) · 2026-07-20

Qwen3 text-embedding model for retrieval, classification and ranking across 100+ languages.

By AlibabaOpen model card ↗

RTX 5090 · 2026-07-20

Up to 3.5x

vs vLLM

Faster in 13 of 13 tests

Up to 6.3x

vs SGLang

Faster in 13 of 13 tests

Up to 18.7x

vs llama.cpp

Faster in 13 of 13 tests

We measure throughput end to end and take the best timed round after a warmup. Each model is tested several ways, from one client sending large requests to many clients hitting the server at once. Every input is unique, so caching can't inflate anyone's numbers.

Measured Throughput

Each engine ran the same tests, in texts/s. Longer bars are better; the number next to a competitor is how many times faster Paddock was.

PaddockvLLMSGLangllama.cpp

Qwen3-Embedding-0.6B

vs vLLM: 1.6-3.5xvs SGLang: 3.5-6.3xvs llama.cpp: 13.3-18.7x
1 client · 32 texts per request
Paddock
3,380.1 texts/s
vLLM
973.23.47x
SGLang
534.16.33x
llama.cpp
253.013.36x
1 client · 128 texts per request
Paddock
3,661.0 texts/s
vLLM
1,286.02.85x
SGLang
710.75.15x
llama.cpp
196.118.67x
4 clients · 32 texts per request
Paddock
5,073.5 texts/s
vLLM
1,849.52.74x
SGLang
982.25.17x
llama.cpp
279.918.13x
8 clients · 128 texts per request
Paddock
3,704.1 texts/s
vLLM
2,291.31.62x
SGLang
1,051.43.52x
llama.cpp
278.613.3x
16 clients · 8 texts per request
Paddock
3,888.9 texts/s
vLLM
2,095.61.86x
SGLang
914.54.25x
llama.cpp
265.214.66x

Qwen3-Embedding-4B

vs vLLM: 1.3-2.1xvs SGLang: 1.7-3.2xvs llama.cpp: 3.9-5.4x
1 client · 32 texts per request
Paddock
719.4 texts/s
vLLM
340.52.11x
SGLang
223.33.22x
llama.cpp
161.64.45x
1 client · 128 texts per request
Paddock
709.8 texts/s
vLLM
466.41.52x
SGLang
253.42.8x
llama.cpp
132.45.36x
4 clients · 32 texts per request
Paddock
747.5 texts/s
vLLM
561.11.33x
SGLang
437.41.71x
llama.cpp
188.93.96x
16 clients · 8 texts per request
Paddock
732.9 texts/s
vLLM
561.41.31x
SGLang
440.41.66x
llama.cpp
186.23.94x

Qwen3-Embedding-8B

vs vLLM: 1.9-2.3xvs SGLang: 2-4.1xvs llama.cpp: 4.2-5.6x
1 client · 32 texts per request
Paddock
560.4 texts/s
vLLM
241.62.32x
SGLang
136.04.12x
llama.cpp
125.04.48x
1 client · 128 texts per request
Paddock
590.2 texts/s
vLLM
271.42.17x
SGLang
150.83.91x
llama.cpp
105.95.57x
4 clients · 32 texts per request
Paddock
619.7 texts/s
vLLM
318.31.95x
SGLang
275.82.25x
llama.cpp
138.14.49x
16 clients · 8 texts per request
Paddock
590.9 texts/s
vLLM
311.81.9x
SGLang
289.12.04x
llama.cpp
141.84.17x

Throughput Under Load

The same results drawn as lines, from a single client on the left to the most concurrent test on the right. The shaded area is the gap between Paddock and the strongest competitor at each point, and the small numbers are Paddock's speedup there.

Qwen3-Embedding-0.6B

PaddockvLLMSGLangllama.cpp
texts/s01,5003,0004,5006,0001 client32 per request1 client128 per request4 clients32 per request8 clients128 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 973.2 texts/sSGLang - 1 client · 32 texts per request: 534.1 texts/sllama.cpp - 1 client · 32 texts per request: 253.0 texts/sPaddock - 1 client · 32 texts per request: 3,380.1 texts/s3.5×vLLM - 1 client · 128 texts per request: 1,286.0 texts/sSGLang - 1 client · 128 texts per request: 710.7 texts/sllama.cpp - 1 client · 128 texts per request: 196.1 texts/sPaddock - 1 client · 128 texts per request: 3,661.0 texts/s2.8×vLLM - 4 clients · 32 texts per request: 1,849.5 texts/sSGLang - 4 clients · 32 texts per request: 982.2 texts/sllama.cpp - 4 clients · 32 texts per request: 279.9 texts/sPaddock - 4 clients · 32 texts per request: 5,073.5 texts/s2.7×vLLM - 8 clients · 128 texts per request: 2,291.3 texts/sSGLang - 8 clients · 128 texts per request: 1,051.4 texts/sllama.cpp - 8 clients · 128 texts per request: 278.6 texts/sPaddock - 8 clients · 128 texts per request: 3,704.1 texts/s1.6×vLLM - 16 clients · 8 texts per request: 2,095.6 texts/sSGLang - 16 clients · 8 texts per request: 914.5 texts/sllama.cpp - 16 clients · 8 texts per request: 265.2 texts/sPaddock - 16 clients · 8 texts per request: 3,888.9 texts/s1.9×+3,224 texts/s

Qwen3-Embedding-4B

PaddockvLLMSGLangllama.cpp
texts/s02004006008001 client32 per request1 client128 per request4 clients32 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 340.5 texts/sSGLang - 1 client · 32 texts per request: 223.3 texts/sllama.cpp - 1 client · 32 texts per request: 161.6 texts/sPaddock - 1 client · 32 texts per request: 719.4 texts/s2.1×vLLM - 1 client · 128 texts per request: 466.4 texts/sSGLang - 1 client · 128 texts per request: 253.4 texts/sllama.cpp - 1 client · 128 texts per request: 132.4 texts/sPaddock - 1 client · 128 texts per request: 709.8 texts/s1.5×vLLM - 4 clients · 32 texts per request: 561.1 texts/sSGLang - 4 clients · 32 texts per request: 437.4 texts/sllama.cpp - 4 clients · 32 texts per request: 188.9 texts/sPaddock - 4 clients · 32 texts per request: 747.5 texts/s1.3×vLLM - 16 clients · 8 texts per request: 561.4 texts/sSGLang - 16 clients · 8 texts per request: 440.4 texts/sllama.cpp - 16 clients · 8 texts per request: 186.2 texts/sPaddock - 16 clients · 8 texts per request: 732.9 texts/s1.3×+379 texts/s

Qwen3-Embedding-8B

PaddockvLLMSGLangllama.cpp
texts/s02004006008001 client32 per request1 client128 per request4 clients32 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 241.6 texts/sSGLang - 1 client · 32 texts per request: 136.0 texts/sllama.cpp - 1 client · 32 texts per request: 125.0 texts/sPaddock - 1 client · 32 texts per request: 560.4 texts/s2.3×vLLM - 1 client · 128 texts per request: 271.4 texts/sSGLang - 1 client · 128 texts per request: 150.8 texts/sllama.cpp - 1 client · 128 texts per request: 105.9 texts/sPaddock - 1 client · 128 texts per request: 590.2 texts/s2.2×vLLM - 4 clients · 32 texts per request: 318.3 texts/sSGLang - 4 clients · 32 texts per request: 275.8 texts/sllama.cpp - 4 clients · 32 texts per request: 138.1 texts/sPaddock - 4 clients · 32 texts per request: 619.7 texts/s1.9×vLLM - 16 clients · 8 texts per request: 311.8 texts/sSGLang - 16 clients · 8 texts per request: 289.1 texts/sllama.cpp - 16 clients · 8 texts per request: 141.8 texts/sPaddock - 16 clients · 8 texts per request: 590.9 texts/s1.9×+319 texts/s

Test Environment

Hardware

GPU
RTX 5090, 32 GB

Engine Versions

Paddock
pre-release build
vLLM
0.25.1. embed: --runner pooling. rerank: seq-classification via --hf-overrides {architectures:[Qwen3ForSequenceClassification]
SGLang
0.5.15.post1 (install --prerelease=allow)
llama.cpp
latest master @ 178a6c44 (~b10069)

Each engine ran the best setup it supports on this hardware, with everything on the GPU.

RTX PRO 6000 Blackwell · 2026-07-11

Up to 2.1x

vs vLLM

Faster in 15 of 15 tests

Up to 3x

vs SGLang

Faster in 15 of 15 tests

Up to 6.3x

vs mistral.rs

Faster in 15 of 15 tests

Up to 15.9x

vs llama.cpp

Faster in 15 of 15 tests

We measure throughput end to end and take the best timed round after a warmup. Each model is tested several ways, from one client sending large requests to many clients hitting the server at once. Every input is unique, so caching can't inflate anyone's numbers.

Measured Throughput

Each engine ran the same tests, in texts/s. Longer bars are better; the number next to a competitor is how many times faster Paddock was.

PaddockvLLMSGLangmistral.rsllama.cpp

Qwen3-Embedding-0.6B

vs vLLM: 1.3-2.1xvs SGLang: 1.5-3xvs mistral.rs: 5.4-6.3xvs llama.cpp: 11.8-15.9x
1 client · 32 texts per request
Paddock
1,200.7 texts/s
vLLM
849.51.41x
SGLang
521.42.3x
mistral.rs
219.15.48x
llama.cpp
89.413.43x
1 client · 128 texts per request
Paddock
1,915.7 texts/s
vLLM
918.52.09x
SGLang
632.03.03x
mistral.rs
307.26.24x
llama.cpp
120.715.87x
4 clients · 32 texts per request
Paddock
1,946.7 texts/s
vLLM
1,331.31.46x
SGLang
1,096.41.78x
mistral.rs
306.86.35x
llama.cpp
141.813.73x
8 clients · 128 texts per request
Paddock
2,044.1 texts/s
vLLM
1,426.31.43x
SGLang
1,271.71.61x
mistral.rs
345.75.91x
llama.cpp
144.614.14x
16 clients · 8 texts per request
Paddock
1,748.7 texts/s
vLLM
1,347.31.3x
SGLang
1,129.11.55x
mistral.rs
321.35.44x
llama.cpp
148.711.76x

Qwen3-Embedding-4B

vs vLLM: 1.2-1.4xvs SGLang: 1.2-2xvs mistral.rs: 4.7-5.3xvs llama.cpp: 4-5.2x
1 client · 32 texts per request
Paddock
284.2 texts/s
vLLM
207.21.37x
SGLang
142.42.0x
mistral.rs
53.95.27x
llama.cpp
54.95.18x
1 client · 128 texts per request
Paddock
341.8 texts/s
vLLM
249.21.37x
SGLang
178.91.91x
mistral.rs
66.75.12x
llama.cpp
69.54.92x
4 clients · 32 texts per request
Paddock
346.9 texts/s
vLLM
273.71.27x
SGLang
263.41.32x
mistral.rs
68.55.06x
llama.cpp
86.04.03x
8 clients · 128 texts per request
Paddock
342.9 texts/s
vLLM
279.71.23x
SGLang
277.51.24x
mistral.rs
72.74.72x
llama.cpp
80.44.26x
16 clients · 8 texts per request
Paddock
338.2 texts/s
vLLM
270.71.25x
SGLang
270.61.25x
mistral.rs
70.54.8x
llama.cpp
85.03.98x

Qwen3-Embedding-8B

vs vLLM: 1.6-1.7xvs SGLang: 1.6-2.3xvs mistral.rs: 5.6-6.2xvs llama.cpp: 3.8-4.3x
1 client · 32 texts per request
Paddock
191.6 texts/s
vLLM
122.61.56x
SGLang
92.52.07x
mistral.rs
33.55.72x
llama.cpp
45.74.19x
1 client · 128 texts per request
Paddock
240.2 texts/s
vLLM
145.41.65x
SGLang
102.92.33x
mistral.rs
40.35.96x
llama.cpp
56.54.25x
4 clients · 32 texts per request
Paddock
244.7 texts/s
vLLM
152.01.61x
SGLang
148.71.65x
mistral.rs
40.95.98x
llama.cpp
64.03.82x
8 clients · 128 texts per request
Paddock
249.4 texts/s
vLLM
153.01.63x
SGLang
159.81.56x
mistral.rs
44.35.63x
llama.cpp
60.04.16x
16 clients · 8 texts per request
Paddock
243.0 texts/s
vLLM
151.51.6x
SGLang
150.21.62x
mistral.rs
39.06.23x
llama.cpp
63.73.81x

Throughput Under Load

The same results drawn as lines, from a single client on the left to the most concurrent test on the right. The shaded area is the gap between Paddock and the strongest competitor at each point, and the small numbers are Paddock's speedup there.

Qwen3-Embedding-0.6B

PaddockvLLMSGLangmistral.rsllama.cpp
texts/s06251,2501,8752,5001 client32 per request1 client128 per request4 clients32 per request8 clients128 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 849.5 texts/sSGLang - 1 client · 32 texts per request: 521.4 texts/smistral.rs - 1 client · 32 texts per request: 219.1 texts/sllama.cpp - 1 client · 32 texts per request: 89.4 texts/sPaddock - 1 client · 32 texts per request: 1,200.7 texts/s1.4×vLLM - 1 client · 128 texts per request: 918.5 texts/sSGLang - 1 client · 128 texts per request: 632.0 texts/smistral.rs - 1 client · 128 texts per request: 307.2 texts/sllama.cpp - 1 client · 128 texts per request: 120.7 texts/sPaddock - 1 client · 128 texts per request: 1,915.7 texts/s2.1×vLLM - 4 clients · 32 texts per request: 1,331.3 texts/sSGLang - 4 clients · 32 texts per request: 1,096.4 texts/smistral.rs - 4 clients · 32 texts per request: 306.8 texts/sllama.cpp - 4 clients · 32 texts per request: 141.8 texts/sPaddock - 4 clients · 32 texts per request: 1,946.7 texts/s1.5×vLLM - 8 clients · 128 texts per request: 1,426.3 texts/sSGLang - 8 clients · 128 texts per request: 1,271.7 texts/smistral.rs - 8 clients · 128 texts per request: 345.7 texts/sllama.cpp - 8 clients · 128 texts per request: 144.6 texts/sPaddock - 8 clients · 128 texts per request: 2,044.1 texts/s1.4×vLLM - 16 clients · 8 texts per request: 1,347.3 texts/sSGLang - 16 clients · 8 texts per request: 1,129.1 texts/smistral.rs - 16 clients · 8 texts per request: 321.3 texts/sllama.cpp - 16 clients · 8 texts per request: 148.7 texts/sPaddock - 16 clients · 8 texts per request: 1,748.7 texts/s1.3×+997 texts/s

Qwen3-Embedding-4B

PaddockvLLMSGLangmistral.rsllama.cpp
texts/s01002003004001 client32 per request1 client128 per request4 clients32 per request8 clients128 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 207.2 texts/sSGLang - 1 client · 32 texts per request: 142.4 texts/smistral.rs - 1 client · 32 texts per request: 53.9 texts/sllama.cpp - 1 client · 32 texts per request: 54.9 texts/sPaddock - 1 client · 32 texts per request: 284.2 texts/s1.4×vLLM - 1 client · 128 texts per request: 249.2 texts/sSGLang - 1 client · 128 texts per request: 178.9 texts/smistral.rs - 1 client · 128 texts per request: 66.7 texts/sllama.cpp - 1 client · 128 texts per request: 69.5 texts/sPaddock - 1 client · 128 texts per request: 341.8 texts/s1.4×vLLM - 4 clients · 32 texts per request: 273.7 texts/sSGLang - 4 clients · 32 texts per request: 263.4 texts/smistral.rs - 4 clients · 32 texts per request: 68.5 texts/sllama.cpp - 4 clients · 32 texts per request: 86.0 texts/sPaddock - 4 clients · 32 texts per request: 346.9 texts/s1.3×vLLM - 8 clients · 128 texts per request: 279.7 texts/sSGLang - 8 clients · 128 texts per request: 277.5 texts/smistral.rs - 8 clients · 128 texts per request: 72.7 texts/sllama.cpp - 8 clients · 128 texts per request: 80.4 texts/sPaddock - 8 clients · 128 texts per request: 342.9 texts/s1.2×vLLM - 16 clients · 8 texts per request: 270.7 texts/sSGLang - 16 clients · 8 texts per request: 270.6 texts/smistral.rs - 16 clients · 8 texts per request: 70.5 texts/sllama.cpp - 16 clients · 8 texts per request: 85.0 texts/sPaddock - 16 clients · 8 texts per request: 338.2 texts/s1.2×+93 texts/s

Qwen3-Embedding-8B

PaddockvLLMSGLangmistral.rsllama.cpp
texts/s0631251882501 client32 per request1 client128 per request4 clients32 per request8 clients128 per request16 clients8 per requestvLLM - 1 client · 32 texts per request: 122.6 texts/sSGLang - 1 client · 32 texts per request: 92.5 texts/smistral.rs - 1 client · 32 texts per request: 33.5 texts/sllama.cpp - 1 client · 32 texts per request: 45.7 texts/sPaddock - 1 client · 32 texts per request: 191.6 texts/s1.6×vLLM - 1 client · 128 texts per request: 145.4 texts/sSGLang - 1 client · 128 texts per request: 102.9 texts/smistral.rs - 1 client · 128 texts per request: 40.3 texts/sllama.cpp - 1 client · 128 texts per request: 56.5 texts/sPaddock - 1 client · 128 texts per request: 240.2 texts/s1.7×vLLM - 4 clients · 32 texts per request: 152.0 texts/sSGLang - 4 clients · 32 texts per request: 148.7 texts/smistral.rs - 4 clients · 32 texts per request: 40.9 texts/sllama.cpp - 4 clients · 32 texts per request: 64.0 texts/sPaddock - 4 clients · 32 texts per request: 244.7 texts/s1.6×vLLM - 8 clients · 128 texts per request: 153.0 texts/sSGLang - 8 clients · 128 texts per request: 159.8 texts/smistral.rs - 8 clients · 128 texts per request: 44.3 texts/sllama.cpp - 8 clients · 128 texts per request: 60.0 texts/sPaddock - 8 clients · 128 texts per request: 249.4 texts/s1.6×vLLM - 16 clients · 8 texts per request: 151.5 texts/sSGLang - 16 clients · 8 texts per request: 150.2 texts/smistral.rs - 16 clients · 8 texts per request: 39.0 texts/sllama.cpp - 16 clients · 8 texts per request: 63.7 texts/sPaddock - 16 clients · 8 texts per request: 243.0 texts/s1.6×+95 texts/s

Test Environment

Hardware

GPU
RTX PRO 6000 Blackwell, 96 GB
Driver
610.43.02
CUDA
13.0
OS
Linux 6.8.0-124

Engine Versions

Paddock
pre-release build
vLLM
0.24.0 (latest on PyPI at benchmark time)
SGLang
0.5.15 (latest on PyPI at benchmark time)
mistral.rs
v0.9.0 (latest release, 2026-07-07)
llama.cpp
b9967 (latest tag at benchmark time)

Each engine ran the best setup it supports on this hardware, with everything on the GPU.

Accuracy Checks

Paddock tunes how each model is served on this GPU. We only let it use a faster serving mode when the measured answer quality stays at or above its reference setup. For models where no faster mode passed that check, the benchmark ran on the reference setup. The per-model results are in the raw data.

How We Tested

  • Every engine ran on the same machine and GPU, never at the same time.
  • The same test program sent identical requests to each engine.
  • Every input was unique, so nothing could be served from a cache.
  • Each number is the best timed round, after a warmup.
  • We measure end to end: what an application connecting over HTTP actually gets.
  • Where the engines differ in ways that could affect the comparison, we note it alongside the results.

Check our numbers

Paddock ships a benchmark harness, so you can run the same comparison on your own hardware.