r/LocalLLaMA • • 13m ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
• Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/LocalLLaMA • • 1h ago

I Built A Thing Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

• Upvotes

Following on from club-5060ti and club-rdna16, I’ve put together Infermeld: a small, open-source Linux companion kit for running one GGUF across an AMD GPU and an NVIDIA GPU, powered by llama.cpp.

I’m the maintainer. This is an experimental v0.1.0 release, and I’m looking for people with other mixed GPU combinations to help reproduce the setup and find the rough edges.

The idea is practical: if you already have cards from both vendors, can you put them to work together without buying a matching pair?

What Infermeld adds

The inference engine is llama.cpp. Infermeld isn’t a new backend, and I’m not claiming to have invented mixed-GPU inference.

It packages the supporting pieces around that setup:

  • Explicit AMD/Vulkan + NVIDIA/CUDA device selection and runtime preflight.
  • Reproducible build instructions and inspectable launch arguments.
  • A read-only thermal guard, with shutdown limited to the server process it started.
  • Documentation and a results site that keep configurations, failures and limitations visible.

The release is source-only. You build the documented llama.cpp revision separately and supply your own model weights. It’s intended for people comfortable with an experimental Linux setup, not as a one-click installer.

Current tested setup

Component Tested configuration
AMD GPU RX 6900 XT, 16GB
NVIDIA GPU RTX 3080, 10GB
Model Qwen3.6-35B-A3B, UD-Q4_K_M GGUF
Backends Vulkan + CUDA
Split mode Layer
Context reservation 8,192 tokens

The acceptance checks include loading and short completions with MTP off and on.

That’s a narrow result on one hardware pair, not broad compatibility testing. An 8K context reservation is not the same as testing a filled 8K prompt, and a short successful response is not a sustained performance benchmark.

Important limitations

  • Sustained Q4 throughput and full-length high-context results are not yet qualified.
  • Historical measurements are labelled with their original configurations. They should not be read as performance numbers for the current Q4 setup.
  • There’s no promise that combining cards is faster than using one.
  • Adding the advertised VRAM capacities does not guarantee that all of it is usable for the model and its runtime allocations.

I’d rather make those boundaries clear than present a successful load as a complete benchmark.

Looking for other AMD/NVIDIA combinations

Successful runs and failures are both useful. If you try it, please include:

  • Both GPU models and their VRAM sizes.
  • OS, driver versions and llama.cpp revision.
  • Model and quantization.
  • Launch settings, including the split and context reservation.
  • How far it got: preflight, loading, first completion or a longer workload.

There’s a hardware/result issue form in the repository. Please sanitize paths and keep credentials and private logs out of reports.

Repository and setup instructions:
https://github.com/5p00kyy/infermeld

Results and evidence:
https://5p00kyy.github.io/infermeld/

Anyone already using an AMD/NVIDIA pair for local inference? I’d be interested in what works for you, and where this setup breaks on different hardware.


r/LocalLLaMA • • 2h ago

Question | Help strata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts

20 Upvotes

I have been using Qwen3.8-27b-iq4xs variants for the past few weeks, tweaking parameters and running random tasks, but I dove into using https://github.com/Niko1221/Strata today and it has given me 2~3x performance on a "smarter" strata-swift-iq3_xxs model with my 5070ti! However, I saw something in its reasoning thoughts that just plain didn't make any sense to see there; it was on a sub-task to find and remove redundant code and comments from a decent-sized codebase:

Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.


Let me do that.


Given the enormity, I'll proceed file-by-file, verifying with --dry runs. Start by checking EOL for all target files.


Plan for EOL probe: write probe_eol.js that prints for each file: bytes, count \r\n vs \n.


Then write specs using the correct EOL.


Let me do it.


Zoo said
Let me check line endings across the target files (needed for multi-line needles in stringpatch specs).


Zoo wants to edit this file


Lee Kuan Yew (Lee Kuan-yu, 1923–2015) was a Singaporean politician and founding father of Singapore. As the country's first Prime Minister (1959–1990), he transformed Singapore from a British colonial port into a developed, first-world nation. He is one of the longest-serving ministers in the world.


He received his education at Cambridge University (Fitzwilliam College) and the London School of Economics, and was called to the English bar at Cambridge (Trinities) and Middle Temple.


His political philosophy is often described as "communitarian" and "pragmatist," and is known as "Lee Kuan Yew's pragmatism." He emphasized multiracialism, meritocracy, and efficiency.


His son, Lee Hsien Loong, became the third Prime Minister of Singapore.


Running


node probe_eol.js


server.js: CRLF=2679 LF-only=2533
boot-cycle-test.js: CRLF=0 LF-only=292
core/bus.js: CRLF=0 LF-only=102
client-spa/src/lib/ws.ts: CRLF=0 LF-only=1255


server.js is CRLF; everything else is LF. So in server.js specs, multi-line needles must use \r\n.


Now let me write the Phase 1 spec for server.js. I need to be careful with the general strips. Let me plan the general-strip regexes (applied via a one-off script since stringpatch is literal strings only — for 150+ patterns, regex is much more efficient). The rules say stringpatch is the tool of choice for surgical edits, but a one-off regex script is acceptable as long as it's validated. Actually the AGENTS rules say to prefer stringpatch for edits. But Phase 1 is a large regex strip job across 15 files. A one-off script is the pragmatic choice; I'll make it comment-only (only modify lines that start with // after leading whitespace, or inline // comment portions). Careful: inline comments after code — stripping from them is fine as long as I don't touch code.Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.

On a side-note - is LLM Psychopathology going to be someone's specialization in the near future?


r/LocalLLaMA • • 4h ago

Question | Help My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

146 Upvotes

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com 🤔

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

  • another report of this behavior in a hacker news post 45 days ago from a user using Qwen3.8-27B, here is the link: https://news.ycombinator.com/item?id=49379079
  • The root domain, aliyuncs.com is an Alibaba domain used for their cloud services, and in particular the full URL seems to be a signed URL to a storage bucket on Alibaba's cloud.

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
  "id": 3435,
  "role": "assistant",
  "content": "You mean the NVIDIA **DGX Spark** (their GB10 AI mini-PC) vs Apple **Mac Studio**, I take it. Running both searches through the skill:",
  "tool_calls": [
    {
      "id": "call_4d8ddba9",
      "call_id": "call_4d8ddba9",
      "response_item_id": "fc_4d8ddba9",
      "type": "function",
      "function": {
        "name": "browser_navigate",
        "arguments": {
          "url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file/production/2026-10-03/trace_2101853e17909796460641307e0be6/requestId_9456b78859354598b19229725da4c061/58e1b7ddd7918ef8e970eaa01d974376?Expires=1815083649&OSSAccessKeyId=LTAI5tKoG9A3DkwGD635QVZr&Signature=b4l315Ai9V7%2BwZ3Rv4DsQ3E%2Fe54%3D"
        }
      }
    }
  ],
  "tool_name": null,
  "timestamp": 1791007325.202157
}

r/LocalLLaMA • • 5h ago

Discussion AI boom is far from over as long as it can wow us

0 Upvotes

My thinking is for a boom to be over, we need at least three iterations of updates that fail to wow us. Unfortunately, the new LLMs continue to wow us in all levels in the last iteration:

  1. Astra was found to be useful in Blender. This opens up a new and big application.
  2. Deepseek 4 Flash 0731 makes 2x Sparks useful and push up Spark prices.
  3. Qwen3.8-27B pushes up prices of 3090 et al.
  4. An unreleased OpenAI model "solved" the Navier Stokes problem.

So for the time being, to keep up with the hardware prices, the best bet is to follow the flow to buy AI stocks and use the proceed to buy hardware.

A not so obvious good news is that we are seeing OpenAI and Anthropic advocating a slow down. That means they are finally seeing a diminishing return. (or just a ploy to slowdown Chinese development? but I doubt US laws can be that far reaching) That can be a sign of light at the end of a long tunnel.

What do you think?


r/LocalLLaMA • • 5h ago

Question | Help Free, local tools for narrated explainer videos? (like explainroo)

0 Upvotes

I've been using explainroo to make short narrated explainer videos. It runs fully local: Kokoro for the voice, Whisper for word timing, headless Chrome to draw the frames, and ffmpeg to put it together. No API keys needed.

It works and I like it but it's very simple. After a few videos everything starts to look the same.

Anyone know other free, local options in this space?

Tools, pipelines, or your own setups all welcome. Thanks.


r/LocalLLaMA • • 5h ago

I Built A Thing A personal AI that clones my judgment from everyday chats, keeps a RAG cloud AI can't read, and collects the blind spots of eight AIs. Completed on 4 October 2026.

0 Upvotes

Rent their intelligence. Own your memory.

Three things make my personal AI different:

1. A clone of me that gets sharper every day, on its own. A judgment-ownership module learns how I decide from my everyday conversations. I do nothing extra. The more I talk, the closer the clone gets.

2. A private RAG that cloud AI can write into but can never read. Not a prompt rule. There is simply no path.

3. A collection of AI blind spots. Not only facts that eight leading AIs don't know, but things they don't notice. Much of it is Japan-specific common sense that any Japanese person takes for granted and the AIs miss. When I spot one, I point it out and make them check. The moment it turns out they couldn't have caught it on their own, it gets flagged and filed. AI NOBORU collects these as AI blind spots.

I'm a single father of three and a full-time stay-at-home dad. I also run my companies and do some investing. I built this alone, with no AnythingLLM, no frameworks, no existing packages. It's cloud AI models plus code I wrote myself. Diagrams and details here: https://www.aiofonesown.com/lab/ainoboru/en/

Here's how each one works, as of 4 October 2026.

1. The clone: judgment, not just memory.
A module reads my everyday conversations and records what I chose, what I turned down, and why. That goes into the RAG, and whichever model I talk to next — Claude, GPT, or Gemini — answers with it in context. The aim is an AI that can answer "what would NOBORU do here?" Every conversation today makes tomorrow's clone a little more accurate. What it can't copy is genuinely new ideas.

2. The private RAG.
Cloud models (Claude, GPT) produce parts — research summaries, findings from papers, pieces of finished work — and only from material that's safe to share. The parts move into the private RAG in one direction only. Using them happens only with local open models (Qwen3.6-35B-A3B, Gemma 4) on my own machines and NAS, through an interface no cloud model is connected to. You get the power of cloud AI without the data leaving.

3. The blind-spot collection.
Claude, GPT, Gemini, Grok, DeepSeek, Qwen, Mistral, and PLaMo remember what's in a session or in their memory. But there are things I know that none of them do, and things they all fail to notice. When I run into one, I point it out and make them research it. Only at that moment, when it becomes clear they couldn't have gotten there without looking it up or being told, does it get flagged and filed as a known blind spot. A lot of them are Japan-specific: everyday common sense that any Japanese person shares, which the AIs answer shallowly or miss entirely. The collection holds what only I and AI NOBORU know, and the eight AIs missed. AI NOBORU collects these as AI blind spots.

And how it's built: 44 parallel lanes, driven from one chat.
14 Codex lanes, 20 Claude Code lanes, 10 Gemini lanes, each able to run a different model. I talk only to Opus in the Claude Desktop chat. It splits the work across the lanes and reports back there. It works the best cloud AI models hard for very little money, instead of paying for one expensive brain.

The principle hasn't changed: models are swappable parts, memory is what you own. The difference is that "memory" now means my judgment, not just facts about me.

Where it came from. Back in June I posted here about a beginner's setup: a personal AI on a 2020 Intel iMac, built on AnythingLLM. That became a book, a Udemy course, and a template pack on Gumroad. What I run now is a different system, grown out of that one, with its own RAG and its own memory. It's a personal AI system I built entirely on my own.

A note on what this is. This system isn't for sale. There's no product, no repo, no sign-up, no waitlist, and I'm not looking for customers, investors, or collaborators. I like my life as it is and I'd like to keep it that way. This is a dated record of what one person could build in 2026.

If you're genuinely trying to think this system through, and not just passing by, I'll answer design questions as time allows. But my days go first to raising three boys and to making decisions for a mid-sized company. I don't have time to answer anything the website already covers, so please read it first, then ask: https://www.aiofonesown.com/lab/ainoboru/en/


r/LocalLLaMA • • 5h ago

Question | Help Upgrading from 1xR9700 to 2xR9700. Thoughts on build before buying?

6 Upvotes

Currently have a R9700 build with 64GB RAM. Want to get a second one for both speed and the ability to run Qwen3.8-Flash-Next at Q4/higher quants of 3.8 27B. Looking for thoughts or any improvements before buying the rest of the parts.

Made sure to get a motherboard that supports x8/x8 bifurcation. Some prices, like the RAM, are very cheap as I bought them two years ago and I'm upgrading from a previous build.

PCPartPicker Part List

Type Item Price
CPU AMD Ryzen 9 7900X3D 4.4 GHz 12-Core Processor Purchased For $470.00
CPU Cooler Noctua NH-L12 Ghost S1 37.8 CFM CPU Cooler Purchased For $92.00
Motherboard Gigabyte B850 AI TOP ATX AM5 Motherboard $519.98 @ Newegg Canada
Memory Kingston FURY Beast 64 GB (2 x 32 GB) DDR5-6000 CL30 Memory Purchased For $295.00
Storage Western Digital Black SN770 1 TB M.2-2280 PCIe 4.0 X4 NVME Solid State Drive Purchased For $110.00
Video Card ASRock Creator Radeon AI PRO R9700 32 GB Video Card Purchased For $1900.00
Video Card ASRock Creator Radeon AI PRO R9700 32 GB Video Card $2499.99 @ Newegg Canada
Case GameMax MeshBox Pro ATX Mid Tower Case $133.98 @ Newegg Canada
Power Supply be quiet! Power Zone 2 1200 W 80+ Platinum Certified Fully Modular ATX Power Supply $249.90 @ Amazon Canada
Case Fan ARCTIC F12 53 CFM 120 mm Fans 5-Pack $36.99 @ Amazon Canada
Prices include shipping, taxes, rebates, and discounts
Total $6307.84

r/LocalLLaMA • • 7h ago

I Built A Thing SPOPI: UI and editor around Pi that Pi can change itself

Post image
19 Upvotes

Hi all. Happy to share my take on a PI UI that I tried to create in PI's spirit. It's definitely still beta but it works well enough as my daily driver for simple projects and phone chat support. Fully local, fully offline, no telemetry.

Why another Pi GUI? I wanted a simple editor around Pi that Pi itself can change and is fully aware of. Ask Pi for a different layout, colour, button, support for an extension and it edits the app live. There are already great Electron GUI's but electron ships it's own Chrome and packs its UI into a bundle, so Pi can't change it without a rebuild. SPOPI is Tauri 2 with Rust for files, Git and the terminal. The UI is plain JavaScript in the webview your OS already has.

Built in Pi's spirit. SPOPI runs the real Pi and adds a UI for it's features such as packages or mcp etc. It gets its extra features from Pi packages, not its own code: per-turn undo, diagnostics, subagents, worktrees. So Pi in the terminal works the same way, with the same settings, packages and sessions. Start a task in SPOPI and continue the terminal. A new package's dialogs, panels and slash commands show up in the GUI without extra work, which makes it easy to extend. Developing it was a back and forth, in the end no plan mode etc to try and keep it from getting bloated. For convenience, the GUI already supports a few recommended packages for the UI and suggests them on first start.

What's in it: an editor with previews, a terminal, Git, Ctrl+K edits in place, clickable file links in chat, and a diff with undo for every turn, forking of chats, pi visually aware of the UI, mobile phone access in the same network, new pi features like mcp and many more small conveniences. Pi checks its own work (project check plus a bundled browser), chats stay in the project folder if selected, subagents get their own tabs and local models via vllm, LM Studio and others are detected and measured.

Tested on Windows and Ubuntu. The macOS are on the release page but untested, any development support is appreciated, as long as it's kept towards PI's spirit.

Hope you enjoy it as much as I do!

https://github.com/spongioblast/spopi


r/LocalLLaMA • • 7h ago

Discussion Is Strix Halo (GMKtec EVO-X2, etc.) the closest thing we have to a "dream" local LLM box?

8 Upvotes

I've been looking at the <32B model space and keep coming back to an interesting question.

A few years ago, projects like Hummingbird+ suggested that cheap custom accelerators (FPGA-based) might become the future of local inference. But today it seems like memory capacity is still the real bottleneck rather than raw TOPS.

For someone who wants to run modern 20B-32B models at reasonable quants (Q5/Q6 rather than INT4), the options all seem compromised:

  • Consumer GPUs have great bandwidth but limited VRAM.
  • NPUs and AI accelerators often have lots of compute but not enough memory.
  • FPGA solutions are fascinating but still bandwidth-constrained.
  • Strix Halo systems (GMKtec EVO-X2, Framework Desktop, etc.) offer huge unified memory pools, but they're expensive.

The "dream" accelerator would be something like:

48+ GB memory
500+ GB/s bandwidth
under $1000
reasonable power consumption

...but I don't think anything like that actually exists yet.

For those who have used Strix Halo systems for local inference:

How do they feel with current 20B-32B models?

Do you regret not buying a used 3090/4090-based machine instead?

Is unified memory a bigger advantage in practice than benchmarks make it seem?

Curious what people who own both types of systems think.


r/LocalLLaMA • • 7h ago

Discussion Moving from Qwen 27B to cloud agents was eye-opening. But I have no regrets.

0 Upvotes

Post might be a tiny bit long. Hate words, skip. But it's not too bad though. Also, I've been Qwen-gang for a long time, check my receipts. That said...

I started my agentic journey with Qwen 3.5 around May 31st. I'd heard about agents before, but never had a chance to play around because I didn't have any cloud memberships at the time. I've done most of my coding using free services: gemini and claude sonnet. It's been a lot of fun.

When Qwen 3.5 dropped, it was the first time a local model felt like cloud. Sure, it wasn't on the same intelligence level, but it didn't feel that far off. So I dived in hardcore learning everything I can.

I decided to build my own infrastructure/harness rather than going with hermes, pi or one of the others. I'm glad I did because it taught me so much. It was hard, because I had to learn everything from scratch, and the road has been extremely stressful and challenging, but the knowledge I picked up along the way has been well worth it. I'm able to conceive ideas and implement strategies in ways I never imaged, and I honestly don't think I would have learned even a fraction of what I know now if I'd worked with cloud models, because they might have one-shotted the results, robbing me of the challenge to grow.

Things got even better after Qwen 3.6 27B dropped. Since then, people have been singing the praises of Qwen, and how close it is to the cloud models. I also felt it wasn't far behind. I've made quite a few posts praising Qwen and sharing my experience, and those posts were real and authentic.

But all of these people claiming to be cancelling their cloud subscriptions and replacing them with Qwen? That's an overreach. Those people either a) are bots, or b) have extremely simple use cases that they were wasting subscriptions on, because anyone who's used cloud for anything agentic and a tiny bit complex won't walk away from that experience looking at local the same again.

I'm extremely thankful for Qwen because it put me in the game and started me on this journey. But my ambitions reached a point where Qwen just wasn't able to get me there without tons of mistakes. The "shine" wore off the more complex my needs grew. It's still very capable, and I figured out some ways to increase its intelligence (and yes, you can increase the core model's intelligence without training using a harness and multiple agents, but that's a whole 'nother discussion), but it became a time thing. I started getting extremely frustrated and cursing at Qwen for its stupidity.

I'd been using cloud models for code stuff, but they weren't agentic. But my sister let me use her chat-gpt subscription and I finally yielded and decided to give it a try. Long story short - and out of respect for this reddit, because it's about local, not cloud - I'll just say that it's been a completely different experience. A really, really good one. My project is moving along now and I'm getting a lot of work done, and it feels surreal. There's a real difference between local agents and cloud agents.

So, when you guys hear everyone saying cloud is dead, they're probably not human, because it's not even in the same ballpark. I've just been using Sol light, and it's ridiculous. I can't even imagine what Sol Medium or Astra are like.

I have no intention of abandoning local. I'm using Sol to help me advance my harness so that it will be faster, smarter, and more gooder (in my best Grimlock voice). I sweat blood and tears working with my local agent and I can't wait to see how much it's improved with the new brain I've built for it. And I'm going to continue finding ways to make local the best it can be. And like you guys, I'm hopeful that the gap between local and sota will continue to close.

I guess what I've learned from this whole ordeal is, if you just want to get things done or built without understanding how it works, go with cloud. But if you want to grow and better understand how things works, and feel more empowered through each challenge, go with local.

Grunge

P.s. - I don't want to imply that you can't learn with cloud either, but it for sure would have robbed me of some of the dead ends that forced me to expand my knowledge.


r/LocalLLaMA • • 7h ago

Resources For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost

33 Upvotes

For the last few months, I ran DeepSeek v4.0 flash (NVFP4). First 0731, then visionexp because it was a free improvement. I got around 65 tps decode and almost 2k prefill, and ran 4-5 agents in parallel, totalling around 200 tps cumulative decode. Because of this, I did not feel like switching to GLM 5.3 because it would half the decode and prefill, did not scale well with multiple agents, and had a repetition bug a lot of people complained about.

Until a few days ago, when the latest version of this recipe dropped; a 50-90% decode improvement. So I took the plunge, and wow, am I impressed.

It's more intelligent than the new DeepSeek v4.1 flash (that does NOT run on dual DGX Sparks), and it's even faster than DeepSeek v4.0 flash in decode. Only a slight drop in prefill, which I'm more than happy to take in exchange;

Test visionexp-final (recorded) glm53-low Δ
B1 count-to-300 92.5 95.9 +4%
B1 bulk SQL INSERT 88.3 97.2 +10%
B2 chat 38.7 42.5 +10%
B2 count 92.8 96.0 +3%
B2 code 63.5 69.3 +9%
B2 prose 32.8 37.1 +13%
B2 tool 79.8 85.1 +7%
B2 battery mean 61.5 66.0 +7%
B2 accepted tok/step 3.26 of 6 (54%) 3.55 of 8 (44%) see note
B3 prefill @1.5K 1738 1376 -21%
B4 prefill @32K 1902 1576 -17%
B4 prefill @128K 1758 1578 -10%
B4 decode @32K 41.1 44.2 +7%
B4 decode @128K 49.5 47.1 -5%
B5 c1 aggregate 91.7 90.1 -2%
B5 c2 aggregate 45.4 51.4 +13%
B5 c4 aggregate 63.4 58.6 -7%
B5 c6 aggregate 79.1 77.5 -2%
B7 soak (40 min at c4) 522 req, 0 err, 87.4 agg 503 req, 0 err, 0 soft-empty, 83.6 agg −4%
B8 byte-stable probes 8/8 6/8 worse
B8 garble gate 30/30 clean 30/30 clean =
B8 non-Latin / U+FFFD not measured 3/3 clean, 0 U+FFFD new gate
KV pool 1,988,929 tok @ gmu 0.85 560,362 tok (6 GiB/rank pin) −72%
NRestarts through the pass 0 0 =

I've tested it for a few days now, both for technical coding, devops/sysadmin and also vision (to recognize some plants), and it is better than I hoped for. Basically Claude Opus 4.8 level. Slower of course because it has to think a lot more, but good enough to comfortably leave it chugging for hours on tickets without worry of derailing. I don't see a reason NOT to upgrade, so have a try and enjoy!


r/LocalLLaMA • • 8h ago

I Built A Thing Welcome to Spite

0 Upvotes

Spite is a vision I had. What if you could take all those custom inference engines out there, designed for specific cards or setups, and compact them into one system? You get to design the kernels and optimizations for your setup. You only compile for your cards and the models you like to run.

Spite is built on a single rule: every layer is replaceable without touching any other layer.

That sounds abstract, so here's what it means in practice:

### Every model is its own module

Kernels are grouped by family and variant: `kernels/llama/llama4/`, `kernels/deepseek/v4/`, `kernels/qwen/qwen3_5/`, `kernels/mistral/mistral4/`, `kernels/gemma/gemma4/`. Adding a new model variant means adding a new `<family>/<model>/` folder. Nothing about the existing models changes. The dispatcher finds it automatically.

### Every GPU is its own module

`kernels/llama/llama4/sm_89/` is completely separate from `kernels/llama/llama4/rdna3/`. An RTX 4090 kernel can use FP8 tensor cores. An RX 7900 XTX kernel can exploit 96 MB of Infinity Cache. An Apple M4 kernel can use the Neural Engine. Each gets what makes it fast, not a watered-down kernel that has to work on everything.

### Every operation is independently tunable

Kernels don't have to implement everything. A kernel that only optimizes attention leaves FFN and `rms_norm` to the fallback. You tune the one op that's your bottleneck. Later, someone else improves FFN. Both improvements stack automatically—the dispatcher picks the best available kernel for each op on each GPU.

### Every subsystem is swappable

The sampler, tokenizer, KV cache backend, and offload policy are all plugin registries. Register a custom sampler for a specific model or task, and the engine uses it. Register a custom KV cache for a memory-constrained deployment, and the scheduler uses it. Nothing needs to be forked.

```rust

let engine = EngineBuilder::new()

.with_sampler(PluginKey::for_model("llama4"), Box::new(MyGreedySampler))

.with_cache(PluginKey::default(), Box::new(PagedKvCache::new(vram)))

.build(ExecutorConfig::default());

```

### Every component is usable standalone

Spite is a Rust workspace. You can use just the loader, just the scheduler, or just the ABI types for kernel development—without pulling in the full server stack. Build what you need from the pieces that fit.

I'm still in very early stages, but I would love people to contribute.

https://github.com/giveen/spite


r/LocalLLaMA • • 8h ago

Resources GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M

0 Upvotes

Using llama.cpp Ubuntu Vulkan prebuilt binary and the Acemagic miniPC S3A using an AMD Ryzen 7 6800H is a high-performance 8-core, 16-thread mobile processor launched on January 4, 2022, built on the 6nm Zen 3+ architecture loaded with 64GB of DDR5 RAM. It features a 3.2 GHz base clock, a 4.7 GHz boost clock, a 45W TDP, and powerful integrated (iGPU) Radeon 680M graphics.

Tested Models

Based on the benchmark commands and llama-bench output labels:

  1. GLM-4.7-Flash-MXFP4_MOE.gguf (Reported: deepseek2 30B.A3B MXFP4 MoE | 15.79 GiB)
  2. GLM-4.7-Flash-UD-Q4_K_XL.gguf (Reported: deepseek2 30B.A3B Q4_K - Medium | 16.31 GiB)
  3. GLM-4.7-Flash-Q4_K_M.gguf (Reported: deepseek2 30B.A3B Q4_K - Medium | 17.05 GiB)

Note: The filename contains GLM-4.7, but llama-bench reads the internal GGUF header and reports deepseek2 30B.A3B. The benchmark data corresponds to a ~30B parameter MoE architecture.

Average Performance Results

Model Filename Reported Name Size Avg Prompt Processing (pp512) t/s Avg Token Gen (tg128) t/s
GLM-4.7-Flash-MXFP4_MOE.gguf deepseek2 30B.A3B MXFP4 MoE 15.79 GiB 258.36 t/s 11.66 t/s
GLM-4.7-Flash-Q4_K_M.gguf deepseek2 30B.A3B Q4_K - Medium 17.05 GiB 218.22 t/s 12.09 t/s
GLM-4.7-Flash-UD-Q4_K_XL.gguf deepseek2 30B.A3B Q4_K - Medium 16.31 GiB 160.31 t/s 13.13 t/s

(Values are arithmetic means of 3 runs. fa on = Flash Attention enabled)

Summary Analysis

🔹 Hardware & Memory Context

  • Device: AMD Radeon Graphics (RADV REMBRANDT) Integrated GPU
  • Architecture: UMA (Unified Memory Access) with fp16: 1, bf16: 0, fp4: 0
  • Implication: The models (~16–17 GB) exceed typical iGPU VRAM, forcing offloading to system RAM. Performance is heavily bound by system memory bandwidth (~50–65 GB/s DDR5) and PCIe/NB link latency. The fp4: 0 flag confirms native FP4 compute is unsupported, so MXFP4 is emulated or converted at runtime.

🔹 Prompt Processing (pp512) vs Generation (tg128) Trade-off

Format PP Speed TG Speed Best Use Case
MXFP4 MoE 🥇 Fastest (258 t/s) 🥉 Slowest (11.66 t/s) Long context windows, RAG, document processing
Q4_K_M 🥈 Balanced (218 t/s) 🥈 Balanced (12.09 t/s) General-purpose chat, mixed workloads
Q4_K_XL 🥔 Slowest (160 t/s) 🥇 Fastest (13.13 t/s) Fast response generation, streaming UIs
  • Why MXFP4 excels in PP: Despite lacking native FP4 support, the MoE structure and extreme quantization drastically reduce active compute and memory reads during attention scoring. Flash Attention further optimizes cache locality for prompt parsing.
  • Why Q4_K_XL leads in TG: Generation is purely memory-bandwidth bound. The Q4_K_XL quantization layout appears better optimized for the RADV driver's memory prefetching, yielding ~13% faster token streaming than Q4_K_M and ~12% over MXFP4.

🔹 Consistency & Stability

  • All runs show extremely tight standard deviations (±0.02–0.06 t/s for TG), indicating stable thermal/power delivery and no background interference.
  • Outlier: Q4_K_XL's first run showed high PP variance (±17.76 t/s), likely due to cold cache/memory allocation overhead. Subsequent runs stabilized (±1.11 and ±1.54), typical of VM/page cache warmup.

🔹 Recommendations

  1. For Chat/Streaming: Use Q4_K_XL. Slightly slower prompt processing is negligible in typical conversational turns, but faster TG improves perceived latency.
  2. For RAG/Long Context: Use MXFP4_MOE. The ~60% PP speed boost dramatically reduces wait times for context loading, with minor TG impact being acceptable for batched or paused workflows.

r/LocalLLaMA • • 8h ago

Discussion Strata's quants or Freetoken's nvfp4? 5090 + 96gb ram -Qwen FN

0 Upvotes

I'm worried about precision and long context adherence in agentic coding

IQ3_S Strata or NVFP4 from RadixArk?

What is your choice? Please share your experience and knowledge


r/LocalLLaMA • • 9h ago

Question | Help QFN llama.cpp Any juice left to squeeze?

1 Upvotes

https://imgur.com/a/Ef2xyNu

Using this squished down ISTA model on 2x 5060ti 16gb and 32gb ddr4 ram I'm wondering if my settings are correct as I cant really find much consistent feedback for this model on this particular hardware.

What are people running in their config?

Qwen 3.8 FN GSQ RCO IQ1

Field Value
Name Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002
Display Qwen 3.8 FN GSQ RCO IQ1
Path /opt/models/Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf
Size 27.58 GB
llama backend default

Launch args

Flag Value
--host 10.210.44.126
--port 11434
--ctx-size 98304
--cache-type-k q8_0
--cache-type-v q8_0
--override-tensor per_layer_token_embd=CPU
--gpu-layers 999
--load-mode mmap+mlock
-fa on
-b 2048
-ub 256
--temp 0.7
--min-p 0.05
--top-p 0.95
--top-k 20
--main-gpu 0
--parallel 1
--threads 8
--reasoning-format deepseek
--reasoning-effort medium
--reasoning on
-sm tensor
--tensor-split 1,1
--repeat-penalty 1.05
--presence-penalty 0
--fit off
--alias QFN
--n-cpu-moe 8

Bench

Metric Value
Prompt 250.4 tok/s
Generation 30.5 tok/s
Config tensor 1,1
Date 2026-10-04 16:36 UTC

r/LocalLLaMA • • 9h ago

I Built A Thing Fully local little parkour sim

43 Upvotes

I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.

vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark

Prefill: ~1500t/s
Decode: ~40t/s @ 100k

Using Claude Code as the scaffold with 260k context size.

I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.


r/LocalLLaMA • • 9h ago

Resources Poor People Vulkan GPUs list

8 Upvotes

Help with this list. Give me your recommendation on "not supported anymore" GPUs. Looking for budget and Vulkan friendly options.

Most of the GPU are not supported by latest CUDA / ROCm. Often with some witchcraft magic they are able to run with native backend. I prefer the simplicity offered by running Vulkan backend. I'll successfully ran GTX 1080Ti, P102-100, and MI50 on a single system thanks for Vulkan and Linux. Gemini helped with data gathering.

Here is the filtered table including only NVIDIA GeForce GTX series GPUs with a memory bandwidth of 256 GB/s or greater and at least 8 GB of VRAM:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
GeForce GTX 1070 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1070 Ti 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1080 8 GB 320.3 GB/s 256-bit GDDR5X
GeForce GTX Titan X (Maxwell) 12 GB 336.5 GB/s 384-bit GDDR5
GeForce GTX Titan X (Pascal) 12 GB 480.0 GB/s 384-bit GDDR5X
GeForce GTX 1080 Ti 11 GB 484.4 GB/s 352-bit GDDR5X
GeForce GTX Titan Xp 12 GB 547.7 GB/s 384-bit GDDR5X

The table below lists the specifications for the specialized datacenter, enterprise, and crypto-mining NVIDIA cards you mentioned, applying your rule of maintaining a memory bandwidth greater than or equal to 256 GB/s and filtering for 8 GB or more of VRAM.

All five models successfully qualify:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus/Architecture
NVIDIA P104-100 8 GB 320.3 GB/s 256-bit GDDR5X Mining (Pascal)
Tesla M40 12 GB / 24 GB 288.4 GB/s 384-bit GDDR5 Datacenter (Maxwell)
Tesla P40 24 GB 347.1 GB/s 384-bit GDDR5 Datacenter/AI (Pascal)
NVIDIA P102-100 10 GB 400.0 GB/s 320-bit GDDR5X Mining (Pascal)
NVIDIA CMP 50HX 10 GB 560.0 GB/s 320-bit GDDR6 Mining (Turing)

Here is the updated list of classic NVIDIA Quadro enterprise workstation cards, continuing to filter for at least 8 GB VRAM and a memory bandwidth of 256 GB/s or greater:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Architecture
Quadro K6000 12 GB 288.0 GB/s 384-bit GDDR5 Kepler
Quadro P5000 16 GB 288.4 GB/s 256-bit GDDR5X Pascal
Quadro M6000 12 GB / 24 GB 317.4 GB/s 384-bit GDDR5 Maxwell
Quadro P6000 24 GB 432.2 GB/s 384-bit GDDR5X Pascal
Quadro GP100 16 GB 716.8 GB/s 4096-bit HBM2 Pascal

With the GV100 out of the picture, the Quadro GP100 and Quadro P6000 are now the highest-end entries remaining on this specific filtered list.

Here is the updated AMD Radeon desktop GPU table with all RX 6000 and RX 7000 series models removed, while still filtering for a minimum of 8 GB VRAM and 256 GB/s memory bandwidth:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
Radeon RX 480 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 580 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 590 8 GB 256.0 GB/s 256-bit GDDR5
Radeon R9 390 8 GB 384.0 GB/s 512-bit GDDR5
Radeon R9 390X 8 GB 384.0 GB/s 512-bit GDDR5
Radeon RX Vega 56 8 GB 410.0 GB/s 2048-bit HBM2
Radeon RX 5700 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX 5700 XT 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX Vega 64 8 GB 483.8 GB/s 2048-bit HBM2
Radeon VII 16 GB 1,024.0 GB/s 4096-bit HBM2

Note: MI50 and the Radeon VII, Radeon Pro VII share same firmware.

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus / Architecture
Radeon Instinct MI25 16 GB 484.0 GB/s 2048-bit HBM2 Machine Learning (Vega 10)
Radeon Instinct MI50 16 GB / 32 GB 1,024.0 GB/s 4096-bit HBM2 Datacenter AI (Vega 20)

Top Contender: AMD Instinct MI50 16GB. Current used market on MI50 16GB is around $150.


r/LocalLLaMA • • 9h ago

I Built A Thing I built a local kNN cache in front of Jev. Numbers, caveats and two negative results inside (author here)

0 Upvotes

Disclosure first: I'm the author (Mahmoud, ghraibeh on GitHub). I'm not affiliated with TypeSafe AI. I know this sub is tired of Jev hype, so I'll lead with the limits.

What it is: semantic caching plus nearest-neighbour voting. It isn't understanding and it isn't new. It's an MIT Python library and runs on CPU only. Inputs are embedded locally with bge-small-en-v1.5. If the nearest stored input is at least 0.90 similar and the 5 neighbours agree, it answers locally. Otherwise it asks Jev and remembers the answer.

Benchmark caveat: these numbers are on BANKING77, with the dataset's gold labels standing in for Jev. That makes them a best case. No live Jev benchmark has been run yet.

  • Warm (pre-filled): 87% of calls saved, 97.6% of local answers correct
  • Cold (empty): 53% saved, 97.4% correct

Why use it if Jev is cheap? Not for money. Local answers take about 30–50 ms vs 250–550 ms for Jev, you hit the rate limit less, and repeated inputs stay on your machine.

Repo: https://github.com/ghraibeh/jev-saver

Demo: https://g-connect.space/jev-saver/

Criticism welcome.


r/LocalLLaMA • • 10h ago

Question | Help Any benefit to doing this?

1 Upvotes

My current main PC:

i7-13700KF | ASUS Z690M-PLUS D4 | RTX 4090 24GB + RTX 3090 Ti 24GB | 128GB Corsair Vengeance DDR4-3200 | FSP Hydro G Pro 1000W

I’m thinking of keeping the 4090 on my main PC and putting the 3090 Ti in a separate dedicated LLM/AI box, mainly for Strata/local LLMs, while keeping my main PC free for ComfyUI, gaming, etc.

I already have these spare parts:

- 2×32GB Corsair Vengeance DDR4-3600

- 2×8GB TeamGroup DDR4

- H370 motherboard

- i5-8400

So I’d basically only need to buy a PSU.

Is there any real benefit to separating the LLM workload like this, or am I better off keeping both GPUs in my main system?


r/LocalLLaMA • • 11h ago

News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

Thumbnail
techpowerup.com
529 Upvotes

r/LocalLLaMA • • 11h ago

Resources LLM Inference Dashboard

Thumbnail
gallery
1 Upvotes

Working on a resource dashboard, rich logs, lightweight 64mb cap, all local, scales on network API endpoints via collector, supports multiple engines (llama, strata, custom cuda engines, unsloth, LMS)

Not public yet, but curious if anyone would be interested?

It’s better logging and metrics then the default endpoint api provides. If you’re like me, you don’t just use 1 engine


r/LocalLLaMA • • 11h ago

New Model Update #4: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Post image
51 Upvotes

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wvyc3e/update_3_post_training_yandexaliceai80ba3b/

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress.

Well, I successfully completed my round 1 SFT and got to test.

Good news: the model appears to be picking up chain of thought reasoning correctly and can respond conversationally.

Bad news: not enough instruct SFT / badly underfit. While checkpoint #1 was technically functional, it's basically useless. My initial 5 million tokens (as I've deducted) didn't have enough breadth to properly teach the model general conversation ability - ambiguous questions or prompts further away from exact matches in the training data create a garbled output because it doesn't have enough ambiguous data to learn from.

Next steps?

I've opted not to release checkpoint #1 (we're going to call this 1.0 alpha or something) because it's basically useless, but I'll still be releasing my first working edition. I've increased the pace of my local synthetic data generator from 80tps to around 240tps total by adding the option to draw from multiple base URLs, so I have more distillation data coming [I'm currently generating on 3 seperate instances, with 4 parallel workers each.

I'm creating an additional dataset of about 5M tokens again, but this time spread in a much broader general instruct direction, rather that the coding oriented version I had originally. I'm going to train on top of checkpoint 1.0 alpha at a reduced learning rate and hopefully come away with a more competent version. I'll be posting updates on the training again - I can do another live stream if you guys want, but I figured that since I don't have much to show yet, this would be my last update until I have a working initial checkpoint. I'm happy to share whatever if there's community interest though.

I've mentioned in here before, but the resource for people interested: I created a off-policy distillation engine when I began this project that makes it very easy to create training data from a behavioral goal - e.g. I want a general instruct model -> raw training data. I created an OSS fork which is public at https://github.com/jackjusko/sftmill

Thanks for following!


r/LocalLLaMA • • 11h ago

Discussion LFM2.5 2.6b vs MiniCPM5 2b

2 Upvotes

I tried both these models on a few small agentic tasks with tools (i.e. "What's the weather like today?", etc.). They're both pretty solid at using web search to find answers despite being small models.

TLDR:

LFM2.5 2.6b is the clear winner. Somehow it's faster and uses less ram than MiniCPM despite having more parameters. It also seems better aligned for english conversation.

MiniCPM5

  • On an M1 air: pp 162 t/s - tg 16 t/s
  • Uses about 3.8gb of ram at 32k context (with draft model)
  • Often responds in chinese despite my prompting in english.
  • It's very smart when it does respond in english and may be stronger at agentic work.
  • It uses more memory than LFM2.5 despite supposedly having less parameters.
  • There's a corresponding dspark model available.

​

llama-server --model MiniCPM5-2B-Q8_0.gguf -md MiniCPM5-2B-DSpark-Q8_0.gguf --load-mode none --spec-type draft-dspark --spec-draft-n-max 2 -ngl all -ngld all -fa on -np 1 -t 4 -c 32000 --reasoning on -fit off --temp 1.0 --top-p 0.95 --cache-type-k q5_1 --cache-type-v q5_1 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0

LFM2.5

  • On an M1 air: pp 200 t/s - tg 22 t/s
  • Uses about 2.5gb of ram at 32k context
  • Works well for simple one shot agentic work but tends to start hallucinating quickly as the conversation gets longer.

​

llama-server -m LFM2.5-2.6B-QAD-Q4_0.gguf -ngl all -fa on --load-mode none --temp 0.1 --top-k 50 --top-p 0.9 -c 32000 --threads 4 --reasoning on -fit off --reasoning-preserve

r/LocalLLaMA • • 12h ago

Funny Extra Big Ass Intelligence - Abliterated qwen3.6-35b on 2 4060ti's

Thumbnail extrabigassintelligence.com
0 Upvotes