Run your AI stack
on your own hardware
Truespar builds the engines AI products and agents need - model inference, databases, search, cache, and object storage. Run each one as a single file on your own machines, with an API compatible with the tool it replaces. They average 7-10x the speed of those enterprise platforms.
What We Do
Truespar is a research arm of The Intelligence Company AB (publ). We started with a question: how much performance can modern hardware actually give?
We write narrow, deep engines in low-level languages, tune them to the CPU and GPUs they run on, and aim for at least a 10x average performance increase. Once we have used them in production we open-source them here.
This is Truespar
The shortest path is the one you own.
Inference, databases and search, running on your hardware, close to the CPU and GPU.
The Stack
Each one runs on your own machines and works on its own.
Paddock
Production inference on your own NVIDIA GPUs, for organisations that cannot send their data anywhere else. Datacenter-grade serving without a datacenter team, on OpenAI and Anthropic compatible APIs.
Explore PaddockTraverse
Works with your existing Neo4j drivers and tools, averaging 7.3x Neo4j's throughput. Runs as a server, embedded in your app, or entirely in a browser tab.
Explore TraverseTorque
A high-performance search engine with a Typesense-compatible API - 9.5x Typesense's search throughput. Hybrid text and vector search.
Explore TorqueSentio
Gives every AI agent its own real email address. Inbound mail arrives as a webhook, already authenticated and scored, and agents reply in thread over REST. Open source under MIT or Apache-2.0.
Explore SentioPaddock
For companies and public bodies running open-source models on their own hardware: continuous batching, chunked prefill, speculative decoding, and custom CUDA kernels tuned per GPU generation. OpenAI and Anthropic compatible, benchmarked against vLLM, SGLang and llama.cpp on the same weights and the same machine, and it feeds the model the document structure, metadata and forensic detail that raw text leaves out.
Rust, C and C++
Many of today's platforms were not built for the modern hardware stack and leave CPU and GPU capabilities unused. That gap is why some workloads run over 100x faster on our engines.
Low-Level, Few Dependencies
Rust, C and C++ with dependencies kept to a minimum - nothing between the code and the hardware.
Profiled at the Hardware Level
AMD uProf and Intel VTune for CPU work, NVIDIA Nsight for GPU - tuned against what the hardware actually does.
Single Binaries, Every Platform
Each engine is one binary, simple to run and maintain - Windows, Linux and macOS, and in some cases the browser through WASM.
Your hardware
Built close to the metal. Run close to home.
If you're building sovereign AI - systems on hardware you control - we'd like to hear what you're running.