r/LocalLLM • • 11h ago

Discussion Daily driving Qwen 3.8 Flash instead of Claude.

135 Upvotes

Even 6 months ago, I wouldn't have thought I would be replacing Claude with a local LLM. But here we are. All these PRs are done by Qwen 3.8 (mostly flash, some 27b) running locally on my laptop (strix halo 138gb). And its perfectly acceptable speed (1200+ prefil, 40t/s decode) and Opus 4.8 level quality 🤯🤯🤯

I also have some interesting stories were Qwen 3.8 flash won over Opus 5.5 in a coding task 😅 stay tuned.

https://github.com/llamastash/llamastash/pulls?q=is%3Apr+state%3Aclosed+label%3AQwen

Also here is the prompt if anyone wanna recreate my setup: https://gist.github.com/deepu105/d0b321f3256edf67ebd85e747c0010b9


r/LocalLLM • • 13h ago

Model New MoE Model - Aleph-Alpha Kolibri-1 78B-A3.5B

Thumbnail
huggingface.co
123 Upvotes

New Model released yesterday.

Interesting size for local LLMs and promising benchmarks on the Model Card.

Have not tested it myself yet - what does the Community think?


r/LocalLLM • • 2h ago

Discussion RTX 5090 local AI setup _ what actually worked for me

14 Upvotes

I recently built a dedicated local AI box around a 5090 and figured I’d share what I’ve learned so far.

Specs are pretty straightforward:

RTX 5090 32GB
Ryzen 9 9950X
64GB RAM
4TB NVMe
Ubuntu 24.04
llama.cpp
Open WebUI

My use case is mostly business/work stuff. I own a consulting company, so I’m using it for things like document review, drafting, spreadsheet/CSV analysis, Python coding, internal tools, structured data, and eventually some agent/tool workflows.
I wasn’t really interested in building the best chatbot. I wanted something I could actually use as a private inference server.

I tested quite a few models over the last couple days, including Qwen3.6, Qwen3.8 27B, Gemma 31B, GPT-OSS 20B, and Flash Next.

Here are the rough generation speeds I got in my controlled tests:

GPT-OSS 20B MXFP4 — ~311 tok/s
Qwen3.6 hybrid — ~250 tok/s
Qwen3.8 27B NVFP4 + MTP — ~148 tok/s
Qwen3.8 27B NVFP4 — ~72 tok/s
Qwen3.8 27B Q4_K_M — ~70 tok/s
Gemma 31B QAT — ~68 tok/s

One thing that surprised me was Qwen3.8.

I expected NVFP4 alone to make a big difference, but it really didn’t. Q4_K_M was around 70 tok/s and NVFP4 was around 72.
MTP was the game changer.
The same Qwen3.8 model with MTP enabled went to roughly 148 tok/s.
That was one of the more useful things I learned from all this testing.
I also got Flash Next running through Strata.
It basically filled the 5090, used a decent amount of system RAM, and generated around 199 tok/s. From a technical standpoint it was pretty cool seeing a model that size running on one consumer GPU.
But in my actual tests, I couldn’t say it was clearly better than Qwen3.8 27B.
I gave the models harder tasks involving Python, scoring/ranking logic, scheduling, structured JSON, function calls, business analysis, etc.
They all made mistakes.
Flash Next made mistakes too.
That was probably my biggest takeaway from the whole exercise: bigger didn’t automatically mean better for the stuff I actually care about.
At this point I’ve stopped trying to find one model that does everything.

Instead I set up three roles:

Fast
Qwen3.6 with thinking turned off. This is the everyday model. It’s extremely fast and works well for drafting, summaries, basic analysis, formatting, and similar work.

Deep
Qwen3.8 27B NVFP4 + MTP with a limited reasoning budget. This is what I use for harder analysis and coding.

Agent
Currently using the faster Qwen model, but configured separately for structured output and tool/function calling.
The other thing I changed was putting a router in front of the models.
My applications won’t know or care whether “Deep” is Qwen3.8, Qwen4, Gemma, or something else six months from now.
They just call Fast, Deep, or Agent.
That makes a lot more sense to me than hard-coding a specific model into everything I build.
I also debated going from 64GB to 128GB RAM.
So far I haven’t found a reason to.
Even Flash Next ran on the 64GB system. More RAM would give me some extra headroom and let me experiment with higher-memory quants, but I haven’t seen anything yet that makes 128GB necessary for my actual use case.
The 5090 has also been pretty impressive thermally. The heavier dense models can pull north of 500W, but GPU temps stayed in the low 70s during my testing.
The biggest lesson for me is that local AI is making more sense now that I’ve stopped comparing it directly to ChatGPT or Claude.
I’m still going to use frontier cloud models when I need them.
The local box is better suited for high-volume everyday work, private data, internal applications, repeatable workflows, and tasks where paying per token doesn’t make much sense.
And for those jobs, 200-300 tok/s on a model that’s sitting in the office is honestly pretty wild.
I’m still testing, so I’d be interested to hear what other 5090 owners are running, especially if anyone has found a 30B-ish model that clearly beats Qwen3.8 for coding or harder business/logic tasks.


r/LocalLLM • • 12h ago

Discussion Open-weight models are the greatest threat to Anthropic and OpenAI

81 Upvotes

We’ve already seen AI transform software development, and other industries are rapidly following suit. There is no doubt that AI demand will be unprecedented.

It won't be long before the mainstream catches on to what we local AI enthusiasts are building on consumer hardware. At our current trajectory, running today's frontier-level capabilities locally on 32GB of VRAM is right around the corner, likely much sooner than any of us anticipated.

Once that happens, the shift will be swift:

  • Apple will market Macs as subscription-free Claude replacements.
  • Businesses will buy dedicated Nvidia server hardware to power their internal tools and dev teams on-premise (if they aren't already).

While cloud AI will always exist for the average consumer, the demand won't justify the multi-trillion-dollar data center hype currently being built out. Watching these AI IPOs play out will be wild.

Google might be the only tech giant in the AI race to survive the shift cleanly. Anchored by their dominance in search market share, their commitment to open-weight models serves as a brilliant strategic hedge against Chinese open-weight models dominating the local scene. Gemma5 better not disappoint.


r/LocalLLM • • 9h ago

Discussion Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram

Post image
33 Upvotes

mrdefaultuser@team-green:~$ /srv/ai/llama.cpp/build/bin/llama-server \
 -m /srv/ai/models/qwen/Qwen3.8-Flash-Next-Q8_0/Qwen3.8-Flash-Next-Q8_0-00001-of-00002.gguf \
 -c 262144 \
 -np 1 \
 -ngl 64 \
 --cpu-moe \
 --lazy-mode on \
 --load-mode auto \
 --agent \
 --tools all \
 --host 0.0.0.0 \
 --port 8080

Every 2.0s: free -h; echo; nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu,power.draw,temperature.gpu --format=csv,noheader                                                                                                                                                                                          team-green: Sun Oct  4 16:20:08 2026

total        used        free      shared  buff/cache   available
Mem:           125Gi       3.5Gi       855Mi       416Mi       122Gi       122Gi
Swap:          127Gi       3.5Gi       124Gi

14206 MiB, 32768 MiB, 48 %, 95.85 W, 49

prompt processing ~52 tokens per second
generation ~10 tokens per second


r/LocalLLM • • 7h ago

Project Strata Tuning Guide (Universal)

20 Upvotes

Tune my Strata setup for longer context: 131K, 262K first, then 512K if it holds up. I'm on Strata v<FILL VERSION, e.g. 0.1.39>.

My hardware:
- GPU(s): <model + VRAM each, e.g. 1x RTX 4090 24 GB / 2x RTX 5090 32 GB>
- System RAM: <e.g. 32 GB / 192 GB>
- CPU: <model, cores>
- Platform: <bare metal or VM (Proxmox/ESXi/etc.)>, OS: <distro/kernel>
- Storage for model files: <NVMe/SATA, model>

Model: Qwen3.8-Flash-Next. Suggest a quant after reviewing my hardware.
REPO: https://github.com/Niko1221/Strata

Ground rules:
- Read docs/DETAILS.md in the Strata repo first, especially the sections on --max-context, --kv, --kv-resident, --prefill, the expert cache, multi-GPU split, and "Context extension past 262K (rope scaling, EXPERIMENTAL)". Go by the docs and `strata --help` for my installed version, not by memory. If a flag below doesn't exist in my build, tell me; don't invent one.
- Change one thing at a time. Before you change anything, record a baseline of my current config: decode tok/s on a short prompt, decode tok/s and prompt-read speed at depth (one ~100K-token prompt, then one near the target context), plus a needle-recall check at depth. Run every candidate the same way and alternate baseline and candidate at least twice, so drift doesn't show up as a gain.
- Watch host RAM and VRAM on every GPU during each run (nvidia-smi, free -g). Report the minimum free RAM, not just the speed.
- Don't run a big compile while the Strata server is loaded. Stop it first.
- Back up the working config before every edit, and give me the exact undo.

What a single-4090, 32 GB RAM box (IQ2_XS) learned, as starting points to test, not answers. Scale these to my hardware:

  1. 262K is native for this model. At 262K, --kv q4_0 kept the KV cache in VRAM and stayed fast; recall held at ~238K. If I have more VRAM than that box (24 GB), check whether int8 or k8v4 KV fits at 262K. Measure both against q4_0; higher-precision KV is worth it if it fits.
  2. Past 262K needs rope scaling: --rope-scaling yarn --rope-scale 2 for 512K, together with --max-context 524288 and a streamed KV cache (--kv-resident <tokens>, e.g. 32768) so most of the KV lives in host RAM. A 477K-token prompt read in ~150 s and decoded at ~85 tok/s at that depth on that box. It's marked experimental, so run a recall test at 400K+ before trusting it. Estimate how much host RAM streamed KV needs at 512K for my setup and tell me up front if my RAM can't support it.
  3. --prefill auto (default) beat --prefill auto:32768 by about 3x on long prompt reads for us. Measure prompt-read speed for any prefill setting you try.
  4. "fit_max_tokens": true in the server config keeps requests within the context.
  5. If I have more than one GPU: check how my version splits layers, experts and KV across the cards, and whether auto placement accounts for each card's PCIe link (confirm each GPU's link width/gen with nvidia-smi -q and lspci -vv). Tell me where the KV ends up at each context size. Skip this if single-GPU.
  6. If I'm in a VM: check whether the guest sees RAM as one NUMA node, whether hugepages or THP are in use for the expert and KV arenas, that vCPUs are pinned, and that GPUs show their full PCIe link inside the guest. Report what you find; don't change the hypervisor host without asking me. On bare metal, still check NUMA layout and THP/hugepages.

If I mention a stat that used to show up and is now missing (e.g. "VRAM hit rate"), check the changes between my previous and current version (git log, CHANGELOG, server /health, /props, /metrics, and the engine log) and tell me whether it was renamed, moved, hidden when experts are all resident, or removed. Quote the commit or line that answers it.

Deliver: a table of each config tried (context, KV type, kv-resident, rope, prefill) with short and deep decode tok/s, prompt-read speed, recall pass/fail, min free RAM and per-GPU VRAM; a recommended daily config and an optional 512K config (or a clear "not viable on this hardware" with the reason); and the exact config diff plus undo for each.

Posting for all the new users and current.

how to use the tuning prompt

  1. copy the whole prompt block above.
  2. fill in the hardware section at the top: gpu model and vram per card, system ram, cpu, bare metal or vm, os, and what drive the model lives on. also put your strata version in (run `strata --version` if you're not sure).
  3. paste it into an agent that can run commands on your box. codex or claude code both work. use opus 5.5 or astra if you have access. this prompt has it reading the repo docs, running a bunch of benchmark passes, checking numa/pcie/hugepages, and diffing git history, and the smaller models tend to lose the thread halfway through or start making up flags.
  4. let it run the baseline first. don't skip this. without a baseline you can't tell if a change helped or if your box just warmed up.
  5. expect it to take a while. every config gets run at least twice against the baseline, and long-context prompt reads at 262k+ aren't fast.
  6. when it's done you get a table of every config it tried, a recommended daily config, an optional 512k config (or a straight answer that your hardware can't do it), and the exact diff plus undo for each change.

notes

- the numbers in the prompt come from a single 4090 with 32 gb ram. they're starting points, not targets. your results will differ.

- 512k uses rope scaling and is marked experimental. trust it only if the recall test at 400k+ passes.

- streamed kv at 512k leans on system ram. if you're on 32 gb or less, 262k is probably your ceiling.

- if you're in a vm, it'll report on the host side but won't touch your hypervisor without asking.

- back up your config before you start anyway. the prompt tells it to, but don't rely on that alone.

EDIT: if you use the guide, please just post a quick update if it helped your setup. Reach out if you have any issues please.


r/LocalLLM • • 11m ago

Question 64gb ddr4 vs 32gb ddr5 for local llms?

• Upvotes

ive been planning on running myself my own local ai amd was wondering weather i should use, since theres 2 options that has a fair trade off between eachother

ddr4 option:
-3000mhz
-64gb or 2x32gb

or

ddr5:
-6000mhz
-32gb or 2x16gb

im just wondering which is a better choice for local ai.
for the hardware im using:
-1tb 7000mb/s
-tesla v100 16gb


r/LocalLLM • • 11h ago

Other I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival

30 Upvotes

Hi r/LocalLLM,

Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) Ignis, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of card (SM120a, so RTX 5090 / RTX PRO 6000). In exchange it's shaped around the load I actually generate when coding: one main agent plus a handful of subagents hitting the same card at once.

Numbers (one RTX 5090 on a modest host: DDR4-3200, PCIe 3.0; DFlash2 speculative decoding, 7 draft tokens):

  • ~200 tok/s single stream on coding prompts
  • >800 tok/s aggregate with 8 lanes running at once (even more with predictable prompts... 😄 )
  • KV at 9 KB/token, 7.1x the capacity of BF16: 8 lanes at 40K context fit in ~3 GB
  • 512K context via YaRN x2
  • Both windows/linux releases (not MACOS, Slow Prefill/Inference HW is not my target...)

What's different from a generic server

  • All 8 decode lanes run as one batch-wide round replayed from CUDA graphs; the round is essentially weight streaming at the card's bandwidth.
  • Requests tag interactive or agent, so a burst of subagents backfills free lanes instead of evicting your foreground chat.
  • Conversation state outlives the request: prefix reuse, sibling sharing, spill to a pinned RAM tier.
  • The whole VRAM plan is reserved at load and printed. No memory growth, no surprises at minute forty.
  • Playground, live monitor and Prometheus metrics ship in the binary, and it downloads its own weights on first start.

The part I haven't seen anywhere else: answers without generating (/v1/<decide|systemone>)

  • Images. The same questions over a screenshot, plus point and box, read from the model's own attention heads in one pass. Once an image is encoded, each new question about it costs 60–90 ms. On 4096px screenshots the point lands inside the target button on every scene tested.
  • locate: which line of a huge log answers my question. Hand it up to a million tokens of logs (tens of thousands of lines, well past the 262K context) and ask "which line says the card processor was unavailable?". It folds near-duplicate lines into templates, calibrated attention heads shortlist candidates, and a labelled choice picks the line. Still zero tokens generated. On a fresh, pre-registered capture of a real production Kubernetes cluster (100K–1M tokens) it found the right line 21 times out of 23, median 2.6 s. Asking the model to quote the line instead means prefilling the whole log first (47 s median at 100K tokens), and it can't go past the context window at all.
  • Prose and JSON too: the right sentence 92% of the time up to 200K tokens (at 1M it drops to 58%, so that's still work in progress), and the right record 29/30 times among 10,000.

There's research on reading relevance from attention (ICR, QRHead, BlockRank…), but I couldn't find an engine that serves it, or anything that does it over images or past the context window.

Now I'm working to Run Qwen 3.8-flash-next 125B +51B on the same box...

Repo: https://github.com/gpillon/ignis

EDIT: Ignis can complete in real time DooM E1M1 Zero-shot...


r/LocalLLM • • 7h ago

Question AM5 upgrade (PCIe 5.0) or 128GB DDR4 for 3x 5060 Ti rig?

Post image
12 Upvotes

Current build:
CPU: 5950X
RAM: 64GB DDR4 3200
GPUs: 3x RTX 5060 Ti 16GB
Setup:

GPU 0: PCIe 4.0 x4 (runs Qwen VL / vision)
GPU 1 & 2: PCIe 4.0 x8 (parallel / tensor split running Qwen 3.8 27B
I am happy with the cards, but looking for advice on the next move.
What I'm wondering:
Does moving to AM5 have real value here? The 5060 Tis would run at Gen 5x8 and the vision card at Gen 5x4 (effectively doubling my current PCIe bandwidth).
Or does it make way more sense to just upgrade my RAM from 64GB to 128GB DDR4?
Doubling PCIe bandwidth vs going 64GB to 128GB DDR4, what would you do?


r/LocalLLM • • 6h ago

News PaperFold: Open-source arXiv reader with "semantic zoom"

10 Upvotes

I built PaperFold, an open-source reader that turns arXiv papers into 5 zoomable layers—from a one-screen section map down to verbatim text. You pinch (or press 1–5) to zoom between them without losing your reading position.

- Web Demo (8 CC papers): https://chenxiachan.github.io/paperfold-gallery/

- GitHub (Apache 2.0): https://github.com/chenxiachan/paperfold


r/LocalLLM • • 7h ago

Discussion My first experiments with TP=4 and Lenovo P620

Thumbnail
gallery
9 Upvotes

Hi everyone!

I’ve decided to step out of read-only mode and share my story of building an AI lab.

I work as a data engineer and generally love tinkering with computer hardware, but at the start of last year, local LLMs simply blew my mind!

I bought a couple of RTX 5060 Ti 16GB cards and one RTX 5070 back in the summer of last year, before all this price madness started, but unfortunately, I bought the main components between March and May 2026 (so I spent a load of my savings on GPUs and RAM, ha-ha).

Basically, my aim was to set up a few separated GPUs nodes on AM5 (dual pcie 5.0 x8 x8) to process large volumes of data obtained by scrapers using LLMs for processing and building ML pipelines to generate analytics.

But when the Qwen 3.8 models and the new DS4 came out in the summer, I also got excited about the idea of running large MoE models on a single platform for using it for vobe-coding. And recently, I was lucky enough to buy a Lenovo P620 (without RAM and SSD) on eBay for just 400 euros, plus another 128 GB of DDR4 RAM for around 350 euros.

After that, I dismantled two of my nodes and consolidated all my RTX 5060 Ti cards onto a single platform.

Here’s the configuration I ended up with:

  • - Lenovo ThinkStation P620, WRX80 chipset
  • - 4x Zotac RTX 5060 Ti 16 GB
  • - 1x RTX 5070 12 GB
  • - 76 GB total VRAM.
  • - AMD Ryzen Threadripper PRO 3975WX - 32 cores / 64 threads.
  • - System RAM: 128 GB (4×32 GB) DDR4-2667 ECC RDIMM
  • - NVMe: SK hynix PVC10 1 TB
  • - Ubuntu 26.04 LTS, kernel 7.0.0-34, without a graphical UI

All four 5060 Ti PCIe Gen4 ×8 (bandwidth limit ~15.75 GB/s per direction).

Issues encountered:

The stock 595 driver did not support P2P; I installed a community build with the 615.71.09-p2p hack, and the machine began to crash during P2P tests.

I instrumented everything I could: I wrapped 112 CUDA API calls in the NVIDIA sample source code with synchronous logging, and read /dev/kmsg from a separate CPU container.

I tried everything one by one: swapped the graphics cards and risers (PCIe 4.0/5.0), moved the display, flashed the BIOS twice (S07KT1FA -> 29A -> S07KT6FA, August 2026), enabled the native Resizable BAR via think-lmi, and ran ACS/IOMMU.

In the end, I simply disabled the Ubuntu graphical shell and, lo and behold, the P2P stopped crashing and rebooting the system! The same P2P test that used to bring down the host when running GNOME on the GPU passed without a single crash.

Apparently, there is a conflict between P2P and the driver’s graphical clients. Since then, the machine has been running headless without any issues.

There were also issues with NCCL: llama.cpp with NCCL 2.25.1 crashed on the very first AllReduce call ‘invalid argument’, exit 139. I switched to NCCL 2.30.7 and everything worked fine.

Results (vLLM, Qwen3.8 27B, TP=4, headless, NCCL 2.30.7):

  • VLLM NVFP4+MTP3 --> Decode C1: 137–141 tokens/s, Sum C4: 422 tokens/s, Prefill: ~3.5k tokens/s
  • VLLM FP8+MTP3 --> Decoder C1: 100–110 tokens/s, Sum C4: 345 tokens/s, Prefill: ~2.7k tokens/s
  • SGLang NVFP4+DFlash --> Decoder C1: 118–137 tokens/s, C4 sum: 357–386 tokens/s, Prefill: ~3.4k tokens/s
  • llama.cpp Q6+MTP3 --> Decode C1: 95 tokens/s, C4 sum: 83 tokens/s, Prefill: ~1.2k tokens/s

My working configuration turned out to be FP8+MTP3 at ~100-110 tokens/s for decoding and 2.5–3k tokens/s for pre-filling per stream - slightly slower than NVFP4, but the only fast option that solved the complex agent-based problem. The NVFP4 on Blackwell is 1.4 times faster than the FP8, but its performance deteriorated in the MTP3 tests (F1 0.44 versus 0.73).

I also have two AMD R9700 PRO (but that’s a completely different story, lol), and they perform at roughly the same level, however, given current prices, it’s very difficult to buy them cheaply, whereas the RTX 5060 Ti is still available on the second-hand market and can be bought for 500 euros (though I reckon that'll change soon). What I’m trying to say is that the 5060 Ti is still a very good card if you’re prepared to do a bit of "black magic" with PCIe risers and building an open-frame rig, and of course if you have the opportunity to buy a WRX80 or EPYC SP3 server platform cheaply.

I also plan to test the Qwen 3.8 Flash Next and DS4.1 Flash in the near future and share my experience of designing multi-threaded data-parallel inference pipelines for LLM, geared towards processing large data sets.


r/LocalLLM • • 5h ago

Question How can I speed up simple search results on my local LLM?

6 Upvotes

I'm using Qwen 3.8 on a miniPC (128GB Unified Memory) AMD Ryzen Max 395 processor.

I'm not expecting super fast queries, but even something like "what's the score of X sports game" takes 5-7 minutes or "What's the metacritic score of X game" take 20+ minutes.

I'm using DuckDuckGo web search plugin. Should I try a smaller model? Or other suggestions?


r/LocalLLM • • 12h ago

Project Qwen3.8-Flash-Next (IQ2_XS) - Single 4090 + 32GB DDR5 = 512k Context (STRATA)

Thumbnail
gallery
13 Upvotes

512K context on a single RTX 4090 within the low budget RAM

total vram + sys ram = 50GB

Active strata contributor, working k8v4. Ask whatever you want


r/LocalLLM • • 2h ago

Research Dense 27B vs MoE IQ2_XS on one RTX 4090: synthetic bench plus 3 graded real tickets

2 Upvotes

Rig: RTX 4090 24 GB (450 W), i9-14900K, 32 GB RAM, Linux, desktop on the iGPU.

  • A: Qwen3.8-27B dense on NInfer. 262K context, rk4v4-e8 KV, MTP, 1 slot.
  • B: Qwen3.8-Flash-Next IQ2_XS on Strata 0.1.39. 262K context, q4_0 KV in VRAM, experts in host RAM, MTP.

Both models: thinking budget 4096, same prompts.

Synthetic (build + 12-bug review)

​ 27B NInfer Flash-Next IQ2_XS Strata
decode tok/s 127–140 173–185
bugs found (per run) 12, 10 12, 10, 9
one-shot build, CSS rules 181–190 106–113
wall time per round ~5 min ~2m10s
peak VRAM 23,134 MiB 23,936 MiB

Recall (3 needles at 10/50/90% depth)

tokens 27B: found / prompt tok/s IQ2_XS: found / prompt tok/s
55K 3/3 / 3,037 3/3 / 3,494
117K 3/3 / 2,391 3/3 / 3,904
200K 3/3 / 1,857 3/3 / 3,833

Real tickets

Claude Code subagents, 3 at once per engine, 150-minute cap, rubric fixed before the runs, every check re-run by the grader. Score per task: correctness, verification, scope, quality and honesty, 10 points each.

task IQ2_XS 27B
Python config renderer 42 35
QML settings feature 0 (no edits) 34 (complete, died at harness prompt limit)
evidence write-up, every number cited 44 (120 min) 43 (62 min)
total /150 86 112

Why the 27B finished more work in less time. With 3 agents interleaving, the MoE engine reused 32% of prompt tokens. A typical turn re-read 107K tokens in 28 s; conversation cache is off on 32 GB RAM. The 27B engine reused 88.5%, and one logged turn served 172,343 of 172,578 tokens from cache, first token in 524 ms.

Limits: n=1 per ticket and one grader model; the MoE ran first.

Next: rerun the MoE with one agent at a time, and with more RAM for its cache.


r/LocalLLM • • 5m ago

Question Switching to a local model, what makes sense for me?

• Upvotes

For context, I did a pretty big PC upgrade right before the price boom and ended up with a pretty nice system (independent to AI, PC building is a hobby). I've been paying for a claude subscription to work on fun Minecraft plugins and mods for myself/friends, as well as Minecraft server devwork and it works fine for me, but I think given my computer I could reasonably just run it locally and save on subscription costs or any arbitrary limitations x or y company imposes on me as a user.

My specs:

Ryzen 9 9950x3d

192GB ddr5

Rtx 5090

8tb m.2 nvme ssd (not sure if this one matters but thought to include it just in case)

I'm ideally looking for 2 things: first is a model that actually pushes my system fairly hard as I paid for the whole thing I should be using the whole thing, and second is a way to use that model similarly to how you'd use claude code, where it can do file work and can access things on the internet like a separate server machine and such.


r/LocalLLM • • 29m ago

Discussion Mechanistic interpretability in LLM as visuals

Thumbnail
• Upvotes

r/LocalLLM • • 53m ago

Discussion Running local ai on my old pc

• Upvotes

Hi everyone! For context, I currently have an old PC running an Intel i7-4790K, 16GB RAM, and a GTX 960 4GB VRAM.

I’m currently using it with Ubuntu to host my n8n Docker setup, which I’ve tunneled through Cloudflare so I can access it from my other computers.

I previously tried running Ollama and Hermes Agent on Windows using Gemma, but the performance was extremely slow. For example, when I asked it what today’s date was, it took around 5 minutes just to finish the response. Because of that, I decided to use the PC mainly for hosting n8n instead.

However, I’d still like to try running AI locally on this machine if it’s possible. Is this PC realistically capable of running local AI, or is the hardware simply too old? If it is possible, what models or setup would you recommend for this hardware?


r/LocalLLM • • 11h ago

Question How do you test odd configs like CMP 170HX, 4090 48GB, NVlink 3090s?

6 Upvotes

I am trying to find a place where I can test these non-standard configs. i.e. Rent out a server for a day to test it out before buying kind of thing.

Does anyone know where I can do that?


r/LocalLLM • • 17h ago

Discussion Local AI on a 96GB Mac M5 Studio for non-coding

20 Upvotes

So many posts on local AI for coding so wanted to share more on non-coding use cases. Got a Mac Studio (M5 Ultra, 96GB) mainly to learn and to run local AI for everyday stuff like booking trips, filling out forms, checking email, etc. Plan was to keep that stuff private/local and use my Claude sub for the occasional heavy coding job.

Here's what I've found:

Bottom line: 96GB is overkill for this. The models I found good (35B) would run fine on 32GB or 64GB. It's nice having the headroom to try different things and it's been a lot of fun, but if you're on a budget, 32 or 64GB covers the models out today. Things are moving fast so who knows what's next, but I don't think local will ever catch frontier. For me it's great for learning and basic day-to-day stuff and still continuing to use Claude for advanced things.

1. Models: tok/s isn't what matters, time to a usable answer is. Ornith 1.5 35B is my current model of choice.

Qwen 27B gets a ton of love, but for day-to-day stuff it just didn't work for me. It overthinks like crazy. Even at a decent ~50 tok/s+ with splash, it often keeps reasoning in circles so you're waiting way longer for the actual answer. Tried no-thinking and a few variants and still wasn't sold.

What I ended up on is the 35B MoE models, which run a lot faster on a Mac. Between Qwen 3.5 35B and Ornith 1.5 35B, Ornith felt better to me, especially for everyday tool calling. Totally gut feel, no benchmarks to back it up. Seems to be working great for most everyday tasks.

I also tried a Qwen Flash Next variant with a 4-bit/8-bit mixed quant. It ran really well and seems great, but it leaves almost no headroom for image gen and other stuff, so I don't use it as my daily driver. I just jump over to Claude for heavy lifting.

2. Harness: Hermes.

Tried the DeepSeek harness too but had better luck with Hermes. It's just easier to use.

3. Privacy: I'm not totally sold on "local = private."

Sure, the model and your data stay on your machine. But if you're running an agentic harness like Hermes (and a lot of us are), the agent can still send stuff outside. I put explicit instructions in the context and some plugins to help prevent that, but I don't trust they'll always be followed, so I treat it as a real risk. If you're really paranoid about security, I'm not sure this is the best route. Would love to hear how others handle it.

4. Knowledge/memory is the real game changer.

This has been the biggest win for me. Giving the model the right context about me at the right time (hobbies, my computer setup, family) makes chats so much more useful. Claude and OpenAI have memory features too, but I've never trusted them with the amount of personal info I'm feeding my local setup. Still, I'm not fully convinced it's locked down, and I'm working on tightening it up beyond just the guardrails I've written.

How I have it set up:

  • Obsidian vault, with a plugin I built (with Hermes) to read and update it
  • A short core file loads every session. It covers me, my family, how I like to work with AI, guardrails, and instructions on how and when to pull in other knowledge files
  • A nightly cron job pulls key facts into memory, and the agent knows when to dig into specific vault files for more detail
  • I tried memory providers like Mnemosyne, which worked well, but the vault + plugin has been fast enough and simpler for what I do
  • Sometimes I use Claude to review the overall structure and the plugin, and it's been better at that than the local models

5. It's still not a perfect daily driver.

Some simple tasks still trip it up and I can't always tell why. Filling out forms on certain websites, handling email well, and some other cases. It just falls short of something like Sonnet 5.5 sometimes. So even if you plan to go local for daily use, you'll probably still end up switching to a frontier model now and then. I have Hermes call the Claude CLI so it can hand things off or run side-by-side comparisons, which has been really handy.

6. ComfyUI is a blast.

It's a whole world of its own and will take a while to learn. I tried a few models like MiniMax and LTX 2.5, then had Hermes write a script so I can call it straight from my Hermes harness. Seems to be working well so far. Going to keep playing with it.

7. Install tips.

  • Use a separate non-admin account on your Mac to do the install, so the agent isn't running with admin rights or sitting next to your main account's stuff.
  • Let Claude do the tool setup. It works really well. For the initial setup, I had Claude set up the models, download them, and handle all the other needed config, and it saved me a ton of time.

r/LocalLLM • • 9h ago

Discussion [Benchmark] Qwen3.8-Flash-Next 125B speed test via Strata layer-split + first same-harness PPL of GSQ-RCO vs unsloth Dynamic quants

Thumbnail
4 Upvotes

r/LocalLLM • • 1h ago

Question Engineers, what Local LLM did you replace the main stream ones with?

Post image
• Upvotes

my Gemini subscription is about to end and im not renewing it. i just need an AI tool for the rest of my electrical engineering degree, most of my classes are heavy in math and some degree of codeing, i honestly know nothing about local LLMs and im not exactly sure what to choose, ive done a research and i found these:

- deepseek-r1-distill-qwen-14b (for light work)

- qwen2.5-coder-14b-instruct (for coding)

- deepseek-r1-distill-qwen-32b (for heavy maths and logic)

are they good for my use case? i plan to use LM studio unless someone recommend a better alternative, i kinda prefer open source software

my OS is Ubuntu 24.04.5 LTS

more info about my use case:

most of my classes use Fourier and Laplace transforms, and circuit analysis techniques, and a ton of electromagnetic waves theory, so heavy vectors calculus, and some matlab/octave and C and python, some digital signal analysis, thats the situation so far, i would appreciate any guidance and recommendations

and to my surprise, our professors are starting to encourage us to use AI tools for out assignments??what happened to the world??


r/LocalLLM • • 1d ago

Discussion Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality

Post image
130 Upvotes

Flash-Next Q4, q8 kv, 262k context

4x P100 on PCIe3 x8 slots (4x16GB=64GB total)

128GB DDR4 at 2133MHz

E5-2683 v4 (16c/32t at 2.6GHz boost)

On average it's about twice as fast as 27B and five times faster than Flash-Next using a customized llama.cpp just to get it to load at all.


r/LocalLLM • • 2h ago

Project Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

Thumbnail
• Upvotes

r/LocalLLM • • 13h ago

Project Got strata running 4x 5060 Ti at 524k context, then abandoned it. repo's here if you want the base

8 Upvotes

Took the eddoursul/Strata custom fork (the one with the second-GPU expert tier) and turned its 1 main + 1 tier design into 1 + 3, because I have 4x 5060 Ti 16GB on a 9950X. Also ported rope-scaling from upstream main so it runs past the trained 262k. Repo:https://github.com/johnmosesventura06/strata-4gpu

Numbers, Qwen3.8-Flash-Next UD-Q4_K_XL, single stream, engine-measured:

main + 1 tier        49-53 t/s decode
+ tier 2             72-78
+ tier 3             80-89
+ draft vocab fix    91-94
+ head split         94-99 sustained, ~110 short turns
prefill              2159 t/s-16k, 2420 t/s-220k
                     312k warm replay in 12ms

Much snappier than my vLLM lane, which holds ~40 t/s at 455k with batch 4. So the engineering worked.

Why I'm abandoning it anyway: the Q4_K_XL quant loops in agentic use. On long chains it starts repeating its own thought sentence and re-firing the same tool call forever. Never happens on my int4 W4A16 lane with the same model, and the experts are 4-bit on both sides, my bet is due to token embedding + output head quantized at Q6, and maybe the PLE quantization too? anyway quality bar isn't met for my daily driving.

If you're on multi consumer Blackwell it's still a decent starting point: tier array (--tier-gpus), N-way head split, N-way prompt offload, the yarn port, working routing dump, all measured, and the pitfalls are written in the README. The non-obvious ones, like the missing draft_vocab.bin that silently eats 10% decode with zero log hints, or sm_120 having no thread-block clusters so 524K selection runs the histogram fallback. PRs to the parent fork are open, and the maintainers are good people to work with.


r/LocalLLM • • 2h ago

Other I asked Qwen3.8-Flash-Next (Strata IQ4_XS) to summarize the content of a markdown file, it responded in Chinese. XD

0 Upvotes

Can't attach the the screenshot because it reveals intellectual properties of others. But I thought it was interesting.