← All benchmarks

Qwen3.5 9B

Paddock vs vLLM vs SGLang vs llama.cpp · Text Generation · RTX 5090 (32 GB) · 2026-07-20

Qwen3.5-series model - a Gated DeltaNet + gated-attention hybrid focused on architectural efficiency and long context.

By AlibabaOpen model card ↗

RTX 5090 · 2026-07-20

Up to 1.7x

vs vLLM

Faster in 5 of 6 tests

Up to 2x

vs SGLang

Faster in 5 of 6 tests

Up to 3.5x

vs llama.cpp

Faster in 6 of 6 tests

We measure throughput end to end and take the best timed round after a warmup. Each model is tested several ways, from one client sending large requests to many clients hitting the server at once. Every input is unique, so caching can't inflate anyone's numbers.

Measured Throughput

Each engine ran the same tests, in tokens/s. Longer bars are better; the number next to a competitor is how many times faster Paddock was.

PaddockvLLMSGLangllama.cpp

Qwen3.5-9B

vs vLLM: 0.8-1.7xvs SGLang: 0.9-2xvs llama.cpp: 1.1-3.5x
1 client · 1k-token prompts · 256 tokens out
Paddock
155.3 tokens/s
vLLM
91.41.7x
SGLang
93.91.65x
llama.cpp
136.21.14x
4 clients · 256-token prompts · 1k tokens out
Paddock
575.7 tokens/s
vLLM
337.21.71x
SGLang
350.01.64x
llama.cpp
427.31.35x
8 clients · 1k-token prompts · 256 tokens out
Paddock
851.8 tokens/s
vLLM
595.21.43x
SGLang
588.51.45x
llama.cpp
403.82.11x
8 clients · 4k-token prompts · 128 tokens out
Paddock
376.2 tokens/s
vLLM
303.81.24x
SGLang
287.51.31x
llama.cpp
193.21.95x
16 clients · agentic-token prompts · 256 tokens out
Paddock
894.1 tokens/s
vLLM
1,098.80.81x
SGLang
970.00.92x
llama.cpp
458.31.95x
32 clients · 1k-token prompts · 256 tokens out
Paddock
1,901.0 tokens/s
vLLM
1,573.81.21x
SGLang
934.32.03x
llama.cpp
543.23.5x

Time to First Token

How long a client waits for the first token (median, best round), in milliseconds. Lower is better. The multiplier is how many times lower Paddock's latency is than that engine's on the same test.

TestPaddockvLLMSGLangllama.cpp
1 client · 1k-token prompts · 256 tokens out75 ms95 ms1.3x75 ms1.0x191 ms2.6x
4 clients · 256-token prompts · 1k tokens out74 ms91 ms1.2x184 ms2.5x485 ms6.6x
8 clients · 1k-token prompts · 256 tokens out282 ms384 ms1.4x410 ms1.5x1,474 ms5.2x
8 clients · 4k-token prompts · 128 tokens out516 ms611 ms1.2x1,400 ms2.7x1,802 ms3.5x
16 clients · agentic-token prompts · 256 tokens out529 ms356 ms0.7x467 ms0.9x2,509 ms4.7x
32 clients · 1k-token prompts · 256 tokens out404 ms495 ms1.2x4,671 ms11.6x1,901 ms4.7x

Throughput Under Load

The same results drawn as lines, from a single client on the left to the most concurrent test on the right. The shaded area is the gap between Paddock and the strongest competitor at each point, and the small numbers are Paddock's speedup there.

Qwen3.5-9B

PaddockvLLMSGLangllama.cpp
tokens/s05001,0001,5002,0001 client1k in · 256 out4 clients256 in · 1k out8 clients1k in · 256 out8 clients4k in · 128 out16 clientsagentic in · 256 out32 clients1k in · 256 outvLLM - 1 client · 1k-token prompts · 256 tokens out: 91.4 tokens/sSGLang - 1 client · 1k-token prompts · 256 tokens out: 93.9 tokens/sllama.cpp - 1 client · 1k-token prompts · 256 tokens out: 136.2 tokens/sPaddock - 1 client · 1k-token prompts · 256 tokens out: 155.3 tokens/s1.1×vLLM - 4 clients · 256-token prompts · 1k tokens out: 337.2 tokens/sSGLang - 4 clients · 256-token prompts · 1k tokens out: 350.0 tokens/sllama.cpp - 4 clients · 256-token prompts · 1k tokens out: 427.3 tokens/sPaddock - 4 clients · 256-token prompts · 1k tokens out: 575.7 tokens/s1.3×vLLM - 8 clients · 1k-token prompts · 256 tokens out: 595.2 tokens/sSGLang - 8 clients · 1k-token prompts · 256 tokens out: 588.5 tokens/sllama.cpp - 8 clients · 1k-token prompts · 256 tokens out: 403.8 tokens/sPaddock - 8 clients · 1k-token prompts · 256 tokens out: 851.8 tokens/s1.4×vLLM - 8 clients · 4k-token prompts · 128 tokens out: 303.8 tokens/sSGLang - 8 clients · 4k-token prompts · 128 tokens out: 287.5 tokens/sllama.cpp - 8 clients · 4k-token prompts · 128 tokens out: 193.2 tokens/sPaddock - 8 clients · 4k-token prompts · 128 tokens out: 376.2 tokens/s1.2×vLLM - 16 clients · agentic-token prompts · 256 tokens out: 1,098.8 tokens/sSGLang - 16 clients · agentic-token prompts · 256 tokens out: 970.0 tokens/sllama.cpp - 16 clients · agentic-token prompts · 256 tokens out: 458.3 tokens/sPaddock - 16 clients · agentic-token prompts · 256 tokens out: 894.1 tokens/s0.8×vLLM - 32 clients · 1k-token prompts · 256 tokens out: 1,573.8 tokens/sSGLang - 32 clients · 1k-token prompts · 256 tokens out: 934.3 tokens/sllama.cpp - 32 clients · 1k-token prompts · 256 tokens out: 543.2 tokens/sPaddock - 32 clients · 1k-token prompts · 256 tokens out: 1,901.0 tokens/s1.2×+327 tokens/s

Test Environment

Hardware

GPU
RTX 5090, 32 GB

Engine Versions

Paddock
pre-release build
vLLM
0.25.1
SGLang
0.5.15.post1
llama.cpp
latest master @ 178a6c44 (~b10069)

Each engine ran the best setup it supports on this hardware, with everything on the GPU.

RTX PRO 6000 Blackwell Max-Q · 2026-07-17

Up to 1.7x

vs vLLM

Faster in 4 of 5 tests

Up to 1.7x

vs SGLang

Faster in 4 of 5 tests

We measure throughput end to end and take the best timed round after a warmup. Each model is tested several ways, from one client sending large requests to many clients hitting the server at once. Every input is unique, so caching can't inflate anyone's numbers.

Measured Throughput

Each engine ran the same tests, in tokens/s. Longer bars are better; the number next to a competitor is how many times faster Paddock was.

PaddockvLLMSGLang

Qwen3.5-9B

vs vLLM: 1-1.7xvs SGLang: 1-1.7x
1 client · 1k-token prompts · 256 tokens out
Paddock
138.6 tokens/s
vLLM
82.31.68x
SGLang
83.91.65x
4 clients · 256-token prompts · 1k tokens out
Paddock
475.8 tokens/s
vLLM
309.11.54x
SGLang
334.21.42x
8 clients · 1k-token prompts · 256 tokens out
Paddock
718.7 tokens/s
vLLM
566.61.27x
SGLang
566.21.27x
8 clients · 4k-token prompts · 128 tokens out
Paddock
296.6 tokens/s
vLLM
300.20.99x
SGLang
299.90.99x
32 clients · 1k-token prompts · 256 tokens out
Paddock
1,487.0 tokens/s
vLLM
1,461.01.02x
SGLang
1,338.71.11x

Time to First Token

How long a client waits for the first token (median, best round), in milliseconds. Lower is better. The multiplier is how many times lower Paddock's latency is than that engine's on the same test.

TestPaddockvLLMSGLang
1 client · 1k-token prompts · 256 tokens out98 ms100 ms1.0x77 ms0.8x
4 clients · 256-token prompts · 1k tokens out94 ms103 ms1.1x92 ms1.0x
8 clients · 1k-token prompts · 256 tokens out360 ms448 ms1.2x389 ms1.1x
8 clients · 4k-token prompts · 128 tokens out681 ms1,725 ms2.5x1,232 ms1.8x
32 clients · 1k-token prompts · 256 tokens out521 ms1,377 ms2.6x1,331 ms2.6x

Throughput Under Load

The same results drawn as lines, from a single client on the left to the most concurrent test on the right. The shaded area is the gap between Paddock and the strongest competitor at each point, and the small numbers are Paddock's speedup there.

Qwen3.5-9B

PaddockvLLMSGLang
tokens/s03757501,1251,5001 client1k in · 256 out4 clients256 in · 1k out8 clients1k in · 256 out8 clients4k in · 128 out32 clients1k in · 256 outvLLM - 1 client · 1k-token prompts · 256 tokens out: 82.3 tokens/sSGLang - 1 client · 1k-token prompts · 256 tokens out: 83.9 tokens/sPaddock - 1 client · 1k-token prompts · 256 tokens out: 138.6 tokens/s1.7×vLLM - 4 clients · 256-token prompts · 1k tokens out: 309.1 tokens/sSGLang - 4 clients · 256-token prompts · 1k tokens out: 334.2 tokens/sPaddock - 4 clients · 256-token prompts · 1k tokens out: 475.8 tokens/s1.4×vLLM - 8 clients · 1k-token prompts · 256 tokens out: 566.6 tokens/sSGLang - 8 clients · 1k-token prompts · 256 tokens out: 566.2 tokens/sPaddock - 8 clients · 1k-token prompts · 256 tokens out: 718.7 tokens/s1.3×vLLM - 8 clients · 4k-token prompts · 128 tokens out: 300.2 tokens/sSGLang - 8 clients · 4k-token prompts · 128 tokens out: 299.9 tokens/sPaddock - 8 clients · 4k-token prompts · 128 tokens out: 296.6 tokens/s1.0×vLLM - 32 clients · 1k-token prompts · 256 tokens out: 1,461.0 tokens/sSGLang - 32 clients · 1k-token prompts · 256 tokens out: 1,338.7 tokens/sPaddock - 32 clients · 1k-token prompts · 256 tokens out: 1,487.0 tokens/s1.0×+152 tokens/s

Test Environment

Hardware

GPU
RTX PRO 6000 Blackwell Max-Q, 96 GB
Driver
595.71.05
CUDA
13.0
OS
Ubuntu 24.04, Linux 6.8.0-117

Engine Versions

Paddock
pre-release build
vLLM
0.25.0 (torch cu130)
SGLang
0.5.15 (torch cu130)

Each engine ran the best setup it supports on this hardware, with everything on the GPU.

How We Tested

  • Every engine ran on the same machine and GPU, never at the same time.
  • The same test program sent identical requests to each engine.
  • Every input was unique, so nothing could be served from a cache.
  • Each number is the best timed round, after a warmup.
  • We measure end to end: what an application connecting over HTTP actually gets.
  • Where the engines differ in ways that could affect the comparison, we note it alongside the results.

Check our numbers

Paddock ships a benchmark harness, so you can run the same comparison on your own hardware.