DS4

A from-scratch inference engine for a handful of frontier open-weight models on hardware you own

No video published yet. The write-up below covers what the tool does and how to try it.

Category
AI models & inference
Audience
Data & ML

Published Updated

DS4 — also called DwarfStar, or just ds4.c — is a from-scratch local inference engine for running very large open-weight language models on a single machine, written by Salvatore Sanfilippo (antirez), the creator of Redis. It is aimed at people who already own serious hardware — a high-memory Apple Silicon Mac, an NVIDIA GPU box, an AMD mini-PC — and want a frontier-scale model answering prompts offline instead of calling a cloud API. It is not a general framework you build products on; it is a narrow, fast runtime for a short list of specific models.

What it does

DS4 loads and serves one family of open-weight models as quickly as it can on consumer-reachable hardware. According to the README it targets:

  • DeepSeek V4 Flash and V4 PRO, and DeepSeek V4.1 Flash
  • GLM 5.2 and 5.3
  • Qwen 3.8 Flash Next

Around that core it ships a server, a built-in coding agent called ds4-agent, vision support, and SSD streaming so a model larger than your RAM can still run by paging weights off disk. The project is MIT-licensed.

The number the video quotes — 126 tokens per second, counted as a total across 16 concurrent sessions — is what the project's own README reports. It is the author's figure, not an independent benchmark, and the channel has not measured it.

How it works

The interesting part is the quantization. antirez describes an "extremely asymmetric quants recipe of 2/8 bit": the experts the model barely touches are squashed down hard, while core layers are kept at higher precision. That is what lets a 284-billion-parameter model fit into 96–128 GB of RAM at the lightest setting instead of needing a server rack. Heavier quants trade memory for quality and, per the video script, push you to 256 GB and up.

Underneath there are three backends: Metal for Apple Silicon, CUDA for NVIDIA (DGX Spark and Ada Lovelace cards are named), and ROCm for AMD Strix Halo mini-PCs. antirez is explicit that DS4 does not link against GGML or llama.cpp code, but says the project "exists thanks to the path opened by the llama.cpp project," reusing its GGUF quantization ideas and kernel designs under MIT terms. He also says he had the first working version running in about a week — again, his own account.

Getting started

Requirements are the whole story here, so read them before cloning. You need one of the supported platforms: a high-RAM Apple Silicon Mac, an NVIDIA DGX Spark or Ada-Lovelace-class GPU, or an AMD Strix Halo machine. For the DeepSeek V4 Flash target, the lightest quant wants 96–128 GB of memory; antirez's companion GGUF conversion of that 284B model is 80.8 GB at Q2, which he notes fits a 128 GB Mac. A heavier Q4 conversion also exists and is considerably bigger.

Cost: the engine is free and MIT-licensed, and the weights it runs are open. There is no API key and no paid service in the loop — the intelligence runs on your own machine, which means you pay for it in RAM, GPU and disk rather than per token. The repository is github.com/antirez/ds4; the code is a C project built from source rather than a packaged installer, and the saved README excerpt here does not include the exact build command, so take the build and run instructions from the repository itself rather than from this page.

When to use it / when not

The catch is blunt and the author does not hide it: DS4 only runs a handful of models, on a short list of hardware, and that hardware is expensive. A 16 GB laptop is not in the conversation. It is also young software written fast by one person, which is a reasonable thing to accept from antirez and an unreasonable thing to assume about reliability in production.

Use it instead of llama.cpp or Ollama when the model you specifically want is one of the supported ones and you want it tuned rather than generically supported. Don't use it if you need wide model coverage, a stable plugin ecosystem, or anything running on modest hardware.

Alternatives

The obvious comparison is llama.cpp, and the friendlier tools built on it — Ollama and LM Studio — which support far more models across far more machines and are what most people already use to run an LLM locally. The other comparison is simply paying for a cloud API: ChatGPT- or Claude-style endpoints need no hardware at all, but your prompts leave your machine and you pay per token. DS4 sits deliberately at the opposite end: own the hardware, own the weights, run offline.

Take DS4 seriously if you have a 128 GB-class Mac or a comparable GPU machine sitting idle and you want a frontier-scale model on it, or if you want to read a compact, readable inference engine written by someone with a long record of writing clear C. If you are shopping for a local-LLM setup that works on normal consumer hardware across many models, this is a fascinating thing to read and the wrong thing to install.

More in AI models & inference

All of AI models & inference →