GPU Support

Paddock serves a GPU generation only after that generation has been brought up, tuned and benchmarked on the real hardware. Cards outside that list are refused at startup with a message naming what is supported. 37 of the 85 cards below are validated. Some of the rest have kernels in the build and can be forced to run; others have none at all, and the table says which is which.

Requirements

Paddock requires an NVIDIA driver that supports CUDA 13; the engine is built against the CUDA 13 API. You do not need the CUDA toolkit installed, only the driver. The server runs a model on a single GPU today.

The driver is the whole dependency. Paddock ships no NVIDIA maths libraries and fetches none at runtime: the executables import no NVIDIA library, the kernel pack imports only the operating system, and the binaries do not contain the names of cuBLAS or the CUDA runtime anywhere, so they could not ask for them. The one NVIDIA binary the engine loads is the display driver.

Find Your Card

This is the same table the Studio shows on its own GPU page, from the same list the engine checks at startup, so the answer here and the answer on your machine cannot disagree.

Graphics cardGenerationTypeCompute capabilityStatus
NVIDIA GB10 (DGX Spark)Blackwell (Jetson)Jetsonsm_121Unvalidated
NVIDIA RTX PRO 6000 Blackwell Server EditionBlackwellData centersm_120Validated
NVIDIA RTX PRO 4500 Blackwell Server EditionBlackwellData centersm_120Validated
NVIDIA RTX PRO 6000 Blackwell Workstation EditionBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 5000 BlackwellBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 4500 BlackwellBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 4000 BlackwellBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 4000 Blackwell SFF EditionBlackwellWorkstationsm_120Validated
NVIDIA RTX PRO 2000 BlackwellBlackwellWorkstationsm_120Validated
GeForce RTX 5090BlackwellWorkstationsm_120Validated
GeForce RTX 5080BlackwellWorkstationsm_120Validated
GeForce RTX 5070 TiBlackwellWorkstationsm_120Validated
GeForce RTX 5070BlackwellWorkstationsm_120Validated
GeForce RTX 5060 TiBlackwellWorkstationsm_120Validated
GeForce RTX 5060BlackwellWorkstationsm_120Validated
GeForce RTX 5050BlackwellWorkstationsm_120Validated
Jetson T5000Blackwell (Jetson)Jetsonsm_110Not in this build
Jetson T4000Blackwell (Jetson)Jetsonsm_110Not in this build
NVIDIA GB300Blackwell UltraData centersm_103Not in this build
NVIDIA B300Blackwell UltraData centersm_103Not in this build
NVIDIA GB300 (DGX Station)Blackwell UltraWorkstationsm_103Not in this build
NVIDIA GB200Blackwell (data center)Data centersm_100Validated
NVIDIA B200Blackwell (data center)Data centersm_100Validated
NVIDIA GH200HopperData centersm_90Unvalidated
NVIDIA H200HopperData centersm_90Unvalidated
NVIDIA H100HopperData centersm_90Unvalidated
NVIDIA L4Ada LovelaceData centersm_89Unvalidated
NVIDIA L40Ada LovelaceData centersm_89Unvalidated
NVIDIA L40SAda LovelaceData centersm_89Unvalidated
NVIDIA RTX 6000 AdaAda LovelaceWorkstationsm_89Unvalidated
NVIDIA RTX 5000 AdaAda LovelaceWorkstationsm_89Unvalidated
NVIDIA RTX 4500 AdaAda LovelaceWorkstationsm_89Unvalidated
NVIDIA RTX 4000 AdaAda LovelaceWorkstationsm_89Unvalidated
NVIDIA RTX 4000 SFF AdaAda LovelaceWorkstationsm_89Unvalidated
NVIDIA RTX 2000 AdaAda LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4090Ada LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4080Ada LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4070 TiAda LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4070Ada LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4060 TiAda LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4060Ada LovelaceWorkstationsm_89Unvalidated
GeForce RTX 4050Ada LovelaceWorkstationsm_89Unvalidated
Jetson AGX OrinAmpere (Jetson)Jetsonsm_87Too old
Jetson Orin NXAmpere (Jetson)Jetsonsm_87Too old
Jetson Orin NanoAmpere (Jetson)Jetsonsm_87Too old
NVIDIA A40AmpereData centersm_86Validated
NVIDIA A10AmpereData centersm_86Validated
NVIDIA A16AmpereData centersm_86Validated
NVIDIA A2AmpereData centersm_86Validated
NVIDIA RTX A6000AmpereWorkstationsm_86Validated
NVIDIA RTX A5000AmpereWorkstationsm_86Validated
NVIDIA RTX A4000AmpereWorkstationsm_86Validated
NVIDIA RTX A3000AmpereWorkstationsm_86Validated
NVIDIA RTX A2000AmpereWorkstationsm_86Validated
GeForce RTX 3090 TiAmpereWorkstationsm_86Validated
GeForce RTX 3090AmpereWorkstationsm_86Validated
GeForce RTX 3080 TiAmpereWorkstationsm_86Validated
GeForce RTX 3080AmpereWorkstationsm_86Validated
GeForce RTX 3070 TiAmpereWorkstationsm_86Validated
GeForce RTX 3070AmpereWorkstationsm_86Validated
GeForce RTX 3060 TiAmpereWorkstationsm_86Validated
GeForce RTX 3060AmpereWorkstationsm_86Validated
GeForce RTX 3050 TiAmpereWorkstationsm_86Validated
GeForce RTX 3050AmpereWorkstationsm_86Validated
NVIDIA A100Ampere (data center)Data centersm_80Not in this build
NVIDIA A30Ampere (data center)Data centersm_80Not in this build
NVIDIA T4TuringData centersm_75Too old
QUADRO RTX 8000TuringWorkstationsm_75Too old
QUADRO RTX 6000TuringWorkstationsm_75Too old
QUADRO RTX 5000TuringWorkstationsm_75Too old
QUADRO RTX 4000TuringWorkstationsm_75Too old
QUADRO RTX 3000TuringWorkstationsm_75Too old
QUADRO T2000TuringWorkstationsm_75Too old
NVIDIA T1200TuringWorkstationsm_75Too old
NVIDIA T1000TuringWorkstationsm_75Too old
NVIDIA T600TuringWorkstationsm_75Too old
NVIDIA T500TuringWorkstationsm_75Too old
NVIDIA T400TuringWorkstationsm_75Too old
GeForce GTX 1650 TiTuringWorkstationsm_75Too old
NVIDIA TITAN RTXTuringWorkstationsm_75Too old
GeForce RTX 2080 TiTuringWorkstationsm_75Too old
GeForce RTX 2080TuringWorkstationsm_75Too old
GeForce RTX 2070TuringWorkstationsm_75Too old
GeForce RTX 2060TuringWorkstationsm_75Too old

What The Badges Mean

BadgeMeaning
ValidatedTesting finished: kernels tuned on that die, correctness gates green, results measured against the other engines. This is the only badge that is a promise.
In bring-upTesting is under way. Refused by default, runs under the override, every log line stamped.
UnvalidatedKernels for it are compiled into the runner and it will very likely work, but nobody has measured it, so the engine refuses rather than serve numbers no one has checked. The override runs it.
Not in this buildNo kernels ship for it. The override does not help, because there is nothing to load.
Too oldBelow Ampere. The Q8_0 serving path is built on integer tensor-core instructions that arrive with Ampere, so there is nothing here to bring up.

Both middle badges are refused at startup and the message looks similar, but only one of them can be overridden into working. Ada and Hopper have kernels in this build; the A100, Blackwell Ultra and the Jetson Thor boards do not.

Why The List Is Exact

Paddock matches your card's compute capability exactly, major and minor. It does not accept a card because the number is close, and this is deliberate: CUDA's own forward compatibility will happily load sm_120 code onto any 12.x device, so a GB10 appears to work right up until the first tensor-core kernel built for sm_120a specifically is launched, and then the request dies mid-flight. A startup check that refuses by name is the only version of this that does not waste your afternoon.

The same rule is why a generation moves onto the supported list only when its testing finishes. Serving on an untested die means either unknown performance or a failure partway through a request, and a clear "not yet" is better than either.

Running On An Unsupported GPU

Setting PADDOCK_UNVALIDATED_ARCH=1 serves anyway, for bring-up and testing. Every log line from such a run is stamped UNVALIDATED, so a number measured this way can never be mistaken for a supported result. Do not run production traffic this way.

It only helps a card marked Unvalidated. On one marked Not in this build there are no kernels to load, so the override changes nothing.

Not Supported At All

  • AMD GPUs (ROCm) and Apple hardware (Metal).
  • CPU inference. Paddock is a GPU engine; the CPU never runs the math.

GPU Kernels

Paddock's CUDA kernels are compiled into the runner, and the engine reads the device's capability at startup to pick the right one. The build carries native code for the generations the table marks Validated or Unvalidated, and for nothing else. Kernels are written per architecture rather than shared, which is where the performance comes from; see the published benchmarks for what that is worth on each board.

GPU Telemetry

The server samples GPU telemetry through NVML on a background thread, decoupled from inference. A snapshot is available at GET /api/gpu and a live stream over WebSocket at GET /api/gpu/stream; the embedded Studio uses the same feed. When NVML is not present the endpoint reports that telemetry is unavailable.