Strata

A local inference engine for anyone who wants a 125B-parameter model on one gaming GPU

Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

No video published yet. The write-up below covers what the tool does and how to try it.

Category
AI models & inference
Audience
Developers
Language
C++
Licence
MIT

Published Updated

Strata is an open-source C++ inference engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter language model, on a single consumer gaming PC rather than on a server. It is for people who already own a recent NVIDIA or AMD graphics card and want a large model running locally for chat, code and image questions, with nothing sent to a cloud provider. Developers who point coding agents at a local endpoint are the clearest audience, but the project's framing is broader: a one-click install for Windows and Linux, aimed at anyone with the hardware.

What it does

Strata loads the Qwen3.8-Flash-Next weights in quantised form and serves them from your own machine. The README describes a model that chats, writes code, reads pictures and works with your apps and coding agents; image input is listed as optional rather than guaranteed.

What makes it usable with tools you already have is the API: Strata exposes OpenAI- and Anthropic-compatible endpoints on localhost, so clients that already speak those protocols connect without modification. The channel's own summary points at Claude Code as the obvious example.

The project's demo is a voxel pagoda garden the model wrote from a single one-shot prompt, recorded on an RTX 5070 with the IQ3_S build at 128K context — a 49-second video published as a release asset alongside v0.1.10. That is the project showing its own result, not an independent test.

How it works

Everything hangs on quantisation. The README lists five builds — Q2_0, IQ2_XS, IQ3_XXS, IQ3_S and a Coder variant — which is how a model of this size is squeezed onto a 12 GB or 16 GB card plus system RAM.

The speed table is the authors' own measurement on two machines they describe. On an RTX 5070 (12 GB) with a Ryzen 5 7600 and 64 GB of RAM, they report 94 tokens per second of generated text at Q2_0, 62 at IQ3_XXS and 53 at IQ3_S, with prompt processing between roughly 1,620 and 2,650 tokens per second. On an RX 9070 XT (16 GB) with a Ryzen 9 3900X and 47 GB of RAM, they report 60 tokens per second at Q2_0, 52 at IQ2_XS and 44 for the Coder build.

"Reads your prompt" means taking in a 32K-token document or chat history, and the README notes that about 60 tokens per second is faster than you can read. These are the project's numbers on its own hardware; the channel does not run the tools it covers, and no third-party benchmark appears in the materials. The README adds that more VRAM is faster, but the saved text is cut off before the figure.

Getting started

You need Windows or Linux and an NVIDIA or AMD graphics card with 12 GB of VRAM or more. The video script is blunt about the second requirement: a full 32 GB of system RAM. The two benchmark machines had 64 GB and 47 GB, so headroom helps.

Nothing here charges money. Strata is MIT-licensed, the weights are the public Qwen3.8-Flash-Next release on Hugging Face, and there is no API key to buy — the intelligence runs on hardware you already own, so your costs are the GPU and the electricity rather than per-token fees.

The repository is github.com/Niko1221/Strata. The project advertises a one-click install for Windows and Linux, but the saved README excerpt does not print the command or the installer name, so take the build from the repository's Releases page — where the v0.1.10 assets live — and follow the install section there rather than assuming a plain clone and compile.

When to use it / when not

The catch to know first: image input is not yet available on AMD Windows, so on a Radeon card under Windows plan on text and code only. The second catch is age — the repository was created on 24 September 2026 and sits at v0.1.10, a few weeks of public history, so expect rough edges.

Use it instead of a paid cloud API such as Claude or ChatGPT when the work must stay on your machine, when you want the model offline, or when per-token billing hurts at your volume. Don't use it if your machine is under the bar — 12 GB of VRAM and 32 GB of RAM are floors, not suggestions — or if you are on a Mac, because the materials name Windows and Linux only and say nothing about macOS. Note too that 2-bit and 3-bit builds are a quality trade, and the README publishes speed tables but no accuracy comparison.

Take Strata seriously if you own a 12 GB or better gaming GPU and have been waiting for a model this size you can point Claude Code at without a subscription or an outbound connection. With 6,881 stars and 617 forks inside roughly ten days of its first commit, it has found that audience fast. What you are buying into is an early release of a single-model engine with a known platform gap, judged by its own benchmarks — worth an evening if the hardware matches, worth skipping until the version number grows if it doesn't.

More in AI models & inference

All of AI models & inference →