AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Strata GitHub project says its software can run a 125-billion-parameter Qwen model locally on supported gaming PCs, using compressed model versions and system RAM alongside GPU memory. Its published RTX 5070 benchmark reports 53–94 tokens per second for answer generation; the supplied report does not provide an RTX 4090 result or support the headline claim of 100T/s.

The open-source Strata project says users can run Qwen3.8-Flash-Next, a 125-billion-parameter model, on supported consumer PCs, but the supplied benchmark does not show it running on an RTX 4090 or at “100T/s.” Strata’s posted results use an RTX 5070 and report answer-generation speeds of 53 to 94 tokens per second, depending on the compressed model version.

Strata’s GitHub page describes software for running the model locally on Windows or Linux with supported NVIDIA or AMD graphics cards. The project says a GPU with at least 12 GB of VRAM, at least 32 GB of system RAM and about 80 GB of free disk space are needed. It also says the model download is about 70 GB and that the system loads roughly 35–55 GB into RAM, reserving some memory for the GPU. These are project-provided requirements, not independently verified results.

The published NVIDIA tests use an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600 processor and 64 GB of RAM. For answer generation, Strata reports 94 tokens per second with Q2_0 compression, 79 with IQ2_XS, 62 with IQ3_XXS and 53 with IQ3_S. Its Coder model is listed at 55 tokens per second. The project says the measurements used four-thousand-token answers and 32-thousand-token prompts; the Q2_0 result used engine version 0.1.36, while the other NVIDIA rows used version 0.1.26.

Strata also reports results from an AMD RX 9070 XT system: 44–60 tokens per second across the listed model variants. Those figures, like the NVIDIA benchmarks, are project-reported. The page distinguishes answer generation from prompt processing: for example, the RTX 5070 Q2_0 result is 2,650 tokens per second when reading a prompt, a different measurement from generating an answer.

At a glance
reportWhen: Current project documentation; publicat…
The developmentThe Strata project published instructions and benchmark results for running a compressed 125-billion-parameter Qwen model locally on consumer PCs, with measured speeds reported in tokens per second.

Local AI on Gaming PCs

The project’s results point to a practical distinction: a large model can be made to run on a consumer computer without fitting entirely in graphics-card memory, but this setup relies on substantial system RAM and model compression. That could give users a way to experiment with local inference rather than sending prompts to a hosted service. Strata says its software keeps activity on the PC; that is a project description, not an independent privacy audit.

The headline’s “100T/s” wording is not borne out by the supplied figures. The benchmarks use tokens per second, not trillions of tokens per second, and the highest listed answer-generation rate is 94 tokens per second on an RTX 5070. The project separately estimates that an RTX 3090 with 24 GB of VRAM may generate about 100–140 tokens per second. Neither statement supplies an RTX 4090 measurement.

Amazon

NVIDIA RTX 4090 graphics card

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Strata’s Benchmarks Measure

Strata presents several compressed versions of the model, with smaller versions generally trading some model capacity or quality for lower memory use and faster generation. Its documentation recommends different choices based on available RAM: for example, it lists IQ2_XS or Q2_0 for a 48 GB system, and says a 64 GB system can use larger options such as IQ3_S. The Coder variant is described as having half of the model’s experts removed and as being intended for coding; the project says it fits within 32 GB of RAM.

The benchmark figures refer to distinct workloads. Generation speed measures how quickly the model produces an answer, while prompt-processing speed measures how quickly it reads submitted text. Strata’s comparison uses a 32K-token prompt and a 4K-token answer, but the source does not give a test date, independent replication, or a result for the RTX 4090 named in the requested headline.

“A token is about ¾ of a word.”

— Strata project documentation

Amazon

high performance gaming PC RAM 32GB

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The RTX 4090 Claim Is Unverified

The material supplied for this report does not establish who originated the phrase “100T/s,” what unit it intends, or whether it is a typo. The project’s benchmark tables use tokens per second and contain no RTX 4090 test. It is not clear whether the requested headline refers to a separate demonstration, another measurement, or the RTX 3090 estimate on Strata’s page.

Strata’s reported figures also lack independent testing details in the supplied material. The project does not provide a benchmark for every listed GPU, and performance can vary with model compression, prompt length, system memory, software version and other hardware. The page says longer-chat and additional-card results are available from community users, but those results are not included here.

Amazon

AI model training GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Hardware Tests Needed

A substantiated RTX 4090 claim would require a benchmark that names the exact model variant, software and engine versions, prompt and output lengths, hardware configuration, and whether the rate measures prompt processing or answer generation. It should also clarify what “100T/s” means and provide reproducible results. Until then, the clearest available comparison is Strata’s own RTX 5070 and RX 9070 XT table, alongside its stated RTX 3090 estimate.

Users considering Strata can check the project’s installation documentation for supported hardware and memory needs. The repository says its installer selects a model based on available hardware and can download and start the local service. Whether the project adds an RTX 4090 benchmark or clarifies the headline claim remains unknown.

Amazon

large capacity SSD for AI workloads

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the supplied report confirm 100 trillion tokens per second?

No. The published figures are in tokens per second, with the RTX 5070 listed at 53–94 tokens per second for answer generation. The source does not define or substantiate “100T/s.”

Was Qwen3.8-Flash-Next tested on an RTX 4090?

No RTX 4090 result appears in the supplied Strata material. Its detailed NVIDIA benchmark is for an RTX 5070; the project separately gives an estimate for an RTX 3090.

What does Strata say a PC needs?

The project lists a supported NVIDIA or AMD GPU with at least 12 GB of VRAM, at least 32 GB of system RAM and about 80 GB of free disk space. Actual performance depends on the configuration and model version.

Does running the model locally mean prompts stay private?

Strata says its local setup keeps activity on the PC. The material provided does not include an independent privacy or network-traffic audit, so that claim has not been independently verified here.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Gaming In 2026: The Best Motherboards Featuring AI Technology

Explore the best gaming motherboards in 2026 featuring advanced AI technology, high performance, and future-proof connectivity options.

AI-Powered Marketing Automation Tools: A Halloween Guide

Discover how AI-powered marketing automation tools transform campaigns with content, personalization, and predictive insights. Learn what matters most for your business.

The Role Of Automation-Flow Rebuilders In Smooth Email Platform Shifts

A new automation-flow rebuilding tool streamlines email platform migrations for agencies, reducing manual work and errors.

DraftKings Is Using AI To Behaviorally Target Chronic Gamblers

The EFF says DraftKings uses betting records to identify customers likely to lose and sends them targeted promotions to encourage more betting.