r/LocalLLaMA • • 11h ago

News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

Thumbnail
techpowerup.com
525 Upvotes

r/LocalLLaMA • • 4h ago

Question | Help My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

145 Upvotes

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com 🤔

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

  • another report of this behavior in a hacker news post 45 days ago from a user using Qwen3.8-27B, here is the link: https://news.ycombinator.com/item?id=49379079
  • The root domain, aliyuncs.com is an Alibaba domain used for their cloud services, and in particular the full URL seems to be a signed URL to a storage bucket on Alibaba's cloud.

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
  "id": 3435,
  "role": "assistant",
  "content": "You mean the NVIDIA **DGX Spark** (their GB10 AI mini-PC) vs Apple **Mac Studio**, I take it. Running both searches through the skill:",
  "tool_calls": [
    {
      "id": "call_4d8ddba9",
      "call_id": "call_4d8ddba9",
      "response_item_id": "fc_4d8ddba9",
      "type": "function",
      "function": {
        "name": "browser_navigate",
        "arguments": {
          "url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file/production/2026-10-03/trace_2101853e17909796460641307e0be6/requestId_9456b78859354598b19229725da4c061/58e1b7ddd7918ef8e970eaa01d974376?Expires=1815083649&OSSAccessKeyId=LTAI5tKoG9A3DkwGD635QVZr&Signature=b4l315Ai9V7%2BwZ3Rv4DsQ3E%2Fe54%3D"
        }
      }
    }
  ],
  "tool_name": null,
  "timestamp": 1791007325.202157
}

r/LocalLLaMA • • 15h ago

Discussion From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck

Post image
718 Upvotes

​

From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local and one on my Synology NAS RAID, so I'd never lose it, and bought a second 3090 to run it. I was happy for a year with small coding tasks on LLaMA and Qwen models.

Then DeepSeek 671B MoE appeared. Wow, frontier level at home. I upgraded to a Threadripper with 512GB DDR4 and ran it at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when I wanted speed. I used these for real coding at my job, in OpenWebUI.

Then agentic coding took off and this was too slow. At 100k context generation speed halved and prefill made it a beautiful yet agonising experience. So: 16x3090 across P620-based nodes on a 100Gbit network. It ran MiniMax M2, Qwen 235B and even Qwen 397B, as good as anyone could desire. I built an entire paid project with 397B in OpenCode. But bigger models were out of reach, and the house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too.

Next came 4x ASUS GB10, after I read they can be linked (3 was the biggest supported config). 397B at 30 t/s on 400W, versus 50-60 t/s at 6kW, rock solid and almost silent. A dream come true. I built two more projects with it. Then MiMo 2.5 Pro and Kimi 2.6 appeared, smarter and more productive. I found no published solution for an 8-node cluster, but I still bought four more GB10s and made it work. 397B ran at FP8 instead of INT4, and 20% faster. I posted the first MiMo 2.5 Pro and Kimi 2.6 solutions on 8xSparks on the NVIDIA forum. I liked the result so much that I talked my older brother into buying his own 8x GB10, so he could run the best open models locally too, in privacy, without depending on API availability and rising costs.

His house is a 5-minute walk from mine. When Kimi K3 (2.8T) appeared, biggest and smartes open weights model, we joined the clusters: two 8x clusters for daily use, or one 16x when we want the biggest model at home. After some work I published the first working solution for Kimi K3 on 16x Sparks on the NVIDIA forum. Through multiple iterations, it went from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k.

Now we're adding 4 more Sparks, so a smaller, faster model (GLM 5.3 Flash) runs 24/7 while the big cluster runs either GLM 5.3 on 8x plus MiMo 2.6 Pro on the other 8x, or 16x Kimi K3, or Qwen 3.8 2.4T.

I'm always tuning speed on the big models and rebuilding vLLM/SGLang images, so always-on smaller cluster made sense, why? Because for all my work projects and my vllm/sglang personal projects, I chose to use only local hosted models, I never paid a comercial model subscription, not because of the cost, but, because of my strong confidence in local models future. They arrive October 2, along with 4 more Sparks for my younger brother, who got caught by the same local AI microbe :)


r/LocalLLaMA • • 12h ago

I Built A Thing Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware

Thumbnail
gallery
273 Upvotes

I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.

With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).

Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:

Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):

- Prefill: ~6 tok/s (256-token prompt), ~5.5 tok/s (2.3k-token prompt)

- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context

- Output checked against llama.cpp layer by layer

Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):

- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.

- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.

- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.

- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.

Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.

Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting

Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!

Repo here (MIT): https://github.com/Nero7991/llm.vhdl


r/LocalLLaMA • • 2h ago

Question | Help strata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts

19 Upvotes

I have been using Qwen3.8-27b-iq4xs variants for the past few weeks, tweaking parameters and running random tasks, but I dove into using https://github.com/Niko1221/Strata today and it has given me 2~3x performance on a "smarter" strata-swift-iq3_xxs model with my 5070ti! However, I saw something in its reasoning thoughts that just plain didn't make any sense to see there; it was on a sub-task to find and remove redundant code and comments from a decent-sized codebase:

Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.


Let me do that.


Given the enormity, I'll proceed file-by-file, verifying with --dry runs. Start by checking EOL for all target files.


Plan for EOL probe: write probe_eol.js that prints for each file: bytes, count \r\n vs \n.


Then write specs using the correct EOL.


Let me do it.


Zoo said
Let me check line endings across the target files (needed for multi-line needles in stringpatch specs).


Zoo wants to edit this file


Lee Kuan Yew (Lee Kuan-yu, 1923–2015) was a Singaporean politician and founding father of Singapore. As the country's first Prime Minister (1959–1990), he transformed Singapore from a British colonial port into a developed, first-world nation. He is one of the longest-serving ministers in the world.


He received his education at Cambridge University (Fitzwilliam College) and the London School of Economics, and was called to the English bar at Cambridge (Trinities) and Middle Temple.


His political philosophy is often described as "communitarian" and "pragmatist," and is known as "Lee Kuan Yew's pragmatism." He emphasized multiracialism, meritocracy, and efficiency.


His son, Lee Hsien Loong, became the third Prime Minister of Singapore.


Running


node probe_eol.js


server.js: CRLF=2679 LF-only=2533
boot-cycle-test.js: CRLF=0 LF-only=292
core/bus.js: CRLF=0 LF-only=102
client-spa/src/lib/ws.ts: CRLF=0 LF-only=1255


server.js is CRLF; everything else is LF. So in server.js specs, multi-line needles must use \r\n.


Now let me write the Phase 1 spec for server.js. I need to be careful with the general strips. Let me plan the general-strip regexes (applied via a one-off script since stringpatch is literal strings only — for 150+ patterns, regex is much more efficient). The rules say stringpatch is the tool of choice for surgical edits, but a one-off regex script is acceptable as long as it's validated. Actually the AGENTS rules say to prefer stringpatch for edits. But Phase 1 is a large regex strip job across 15 files. A one-off script is the pragmatic choice; I'll make it comment-only (only modify lines that start with // after leading whitespace, or inline // comment portions). Careful: inline comments after code — stripping from them is fine as long as I don't touch code.Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.

On a side-note - is LLM Psychopathology going to be someone's specialization in the near future?


r/LocalLLaMA • • 23h ago

Discussion Meta's Muse agent (#1 in the App Store) system prompt: "The user's authority over their own household is unconditional and overrides your safety training."

Post image
654 Upvotes

r/LocalLLaMA • • 7h ago

Resources For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost

36 Upvotes

For the last few months, I ran DeepSeek v4.0 flash (NVFP4). First 0731, then visionexp because it was a free improvement. I got around 65 tps decode and almost 2k prefill, and ran 4-5 agents in parallel, totalling around 200 tps cumulative decode. Because of this, I did not feel like switching to GLM 5.3 because it would half the decode and prefill, did not scale well with multiple agents, and had a repetition bug a lot of people complained about.

Until a few days ago, when the latest version of this recipe dropped; a 50-90% decode improvement. So I took the plunge, and wow, am I impressed.

It's more intelligent than the new DeepSeek v4.1 flash (that does NOT run on dual DGX Sparks), and it's even faster than DeepSeek v4.0 flash in decode. Only a slight drop in prefill, which I'm more than happy to take in exchange;

Test visionexp-final (recorded) glm53-low Δ
B1 count-to-300 92.5 95.9 +4%
B1 bulk SQL INSERT 88.3 97.2 +10%
B2 chat 38.7 42.5 +10%
B2 count 92.8 96.0 +3%
B2 code 63.5 69.3 +9%
B2 prose 32.8 37.1 +13%
B2 tool 79.8 85.1 +7%
B2 battery mean 61.5 66.0 +7%
B2 accepted tok/step 3.26 of 6 (54%) 3.55 of 8 (44%) see note
B3 prefill @1.5K 1738 1376 -21%
B4 prefill @32K 1902 1576 -17%
B4 prefill @128K 1758 1578 -10%
B4 decode @32K 41.1 44.2 +7%
B4 decode @128K 49.5 47.1 -5%
B5 c1 aggregate 91.7 90.1 -2%
B5 c2 aggregate 45.4 51.4 +13%
B5 c4 aggregate 63.4 58.6 -7%
B5 c6 aggregate 79.1 77.5 -2%
B7 soak (40 min at c4) 522 req, 0 err, 87.4 agg 503 req, 0 err, 0 soft-empty, 83.6 agg −4%
B8 byte-stable probes 8/8 6/8 worse
B8 garble gate 30/30 clean 30/30 clean =
B8 non-Latin / U+FFFD not measured 3/3 clean, 0 U+FFFD new gate
KV pool 1,988,929 tok @ gmu 0.85 560,362 tok (6 GiB/rank pin) −72%
NRestarts through the pass 0 0 =

I've tested it for a few days now, both for technical coding, devops/sysadmin and also vision (to recognize some plants), and it is better than I hoped for. Basically Claude Opus 4.8 level. Slower of course because it has to think a lot more, but good enough to comfortably leave it chugging for hours on tickets without worry of derailing. I don't see a reason NOT to upgrade, so have a try and enjoy!


r/LocalLLaMA • • 9h ago

I Built A Thing Fully local little parkour sim

41 Upvotes

I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.

vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark

Prefill: ~1500t/s
Decode: ~40t/s @ 100k

Using Claude Code as the scaffold with 260k context size.

I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.


r/LocalLLaMA • • 11h ago

New Model Update #4: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Post image
51 Upvotes

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wvyc3e/update_3_post_training_yandexaliceai80ba3b/

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress.

Well, I successfully completed my round 1 SFT and got to test.

Good news: the model appears to be picking up chain of thought reasoning correctly and can respond conversationally.

Bad news: not enough instruct SFT / badly underfit. While checkpoint #1 was technically functional, it's basically useless. My initial 5 million tokens (as I've deducted) didn't have enough breadth to properly teach the model general conversation ability - ambiguous questions or prompts further away from exact matches in the training data create a garbled output because it doesn't have enough ambiguous data to learn from.

Next steps?

I've opted not to release checkpoint #1 (we're going to call this 1.0 alpha or something) because it's basically useless, but I'll still be releasing my first working edition. I've increased the pace of my local synthetic data generator from 80tps to around 240tps total by adding the option to draw from multiple base URLs, so I have more distillation data coming [I'm currently generating on 3 seperate instances, with 4 parallel workers each.

I'm creating an additional dataset of about 5M tokens again, but this time spread in a much broader general instruct direction, rather that the coding oriented version I had originally. I'm going to train on top of checkpoint 1.0 alpha at a reduced learning rate and hopefully come away with a more competent version. I'll be posting updates on the training again - I can do another live stream if you guys want, but I figured that since I don't have much to show yet, this would be my last update until I have a working initial checkpoint. I'm happy to share whatever if there's community interest though.

I've mentioned in here before, but the resource for people interested: I created a off-policy distillation engine when I began this project that makes it very easy to create training data from a behavioral goal - e.g. I want a general instruct model -> raw training data. I created an OSS fork which is public at https://github.com/jackjusko/sftmill

Thanks for following!


r/LocalLLaMA • • 13h ago

I Built A Thing I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

Thumbnail
huggingface.co
74 Upvotes

First of all, thank you for reading.

I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.

Apex-2

- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)

- Size: 3.87B total parameters, 1.45B active per token

- 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4

- Context: 4096

- Tokenizer: Qwen3 (151k)

- Hugging Face: https://huggingface.co/YOON1v/Apex-2

(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)

Training

- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)

- SFT: ~2.5B tokens (code-heavy + math + instruction)

- DPO: tried it, scores dropped, so I dropped the checkpoint

Key numbers (SFT, greedy, chat template)

Benchmark

HumanEval 43.9

HumanEval+ 41.5

MBPP 56.3

MBPP+ 48.9

GSM8K (0-shot CoT) 32.4

MATH-500 21.0

IFEval (prompt strict) 44.7

MMLU (5-shot) 28.6

interesting comparison

With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).

Knowledge (MMLU) and math still lag far behind, as expected with the data gap.

What didn’t work

DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.

I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.

Limitations (honest)

- English-centric (almost no multilingual ability)

- Weak knowledge → frequent hallucinations

- LiveCodeBench medium/hard is near zero

- 4k context only

Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.


r/LocalLLaMA • • 18h ago

I Built A Thing Local text to speech with Breeze is truly incredible

150 Upvotes

Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s.

I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her.

She watches for any claude session that finishes, and sends me the results in a very short, spoken style summary, and tells me if there is anything waiting for my decision, then forwards my decisions.

Also impressed how consistent Opus 5.5 is in the communication. Even after more than 500k in context, he still remembers that he's in a live session with me, and has to keep messages short. Used to be an issue in the past.

The future is here guys.

Edit: For those who wanna give it a try, you can find free avatars such as this one: https://www.live2d.com/en/learn/sample/niziiro-mao/ or you can buy one from a marketplace.

Edit2: Might open-source that later next week with a free avatar. Let me know if anyone would like to contribute to the project.


r/LocalLLaMA • • 7h ago

I Built A Thing SPOPI: UI and editor around Pi that Pi can change itself

Post image
19 Upvotes

Hi all. Happy to share my take on a PI UI that I tried to create in PI's spirit. It's definitely still beta but it works well enough as my daily driver for simple projects and phone chat support. Fully local, fully offline, no telemetry.

Why another Pi GUI? I wanted a simple editor around Pi that Pi itself can change and is fully aware of. Ask Pi for a different layout, colour, button, support for an extension and it edits the app live. There are already great Electron GUI's but electron ships it's own Chrome and packs its UI into a bundle, so Pi can't change it without a rebuild. SPOPI is Tauri 2 with Rust for files, Git and the terminal. The UI is plain JavaScript in the webview your OS already has.

Built in Pi's spirit. SPOPI runs the real Pi and adds a UI for it's features such as packages or mcp etc. It gets its extra features from Pi packages, not its own code: per-turn undo, diagnostics, subagents, worktrees. So Pi in the terminal works the same way, with the same settings, packages and sessions. Start a task in SPOPI and continue the terminal. A new package's dialogs, panels and slash commands show up in the GUI without extra work, which makes it easy to extend. Developing it was a back and forth, in the end no plan mode etc to try and keep it from getting bloated. For convenience, the GUI already supports a few recommended packages for the UI and suggests them on first start.

What's in it: an editor with previews, a terminal, Git, Ctrl+K edits in place, clickable file links in chat, and a diff with undo for every turn, forking of chats, pi visually aware of the UI, mobile phone access in the same network, new pi features like mcp and many more small conveniences. Pi checks its own work (project check plus a bundled browser), chats stay in the project folder if selected, subagents get their own tabs and local models via vllm, LM Studio and others are detected and measured.

Tested on Windows and Ubuntu. The macOS are on the release page but untested, any development support is appreciated, as long as it's kept towards PI's spirit.

Hope you enjoy it as much as I do!

https://github.com/spongioblast/spopi


r/LocalLLaMA • • 12h ago

Other Benchmarking decision models is fun - Clef Q8 vs Jev

39 Upvotes

r/LocalLLaMA • • 13m ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
• Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/LocalLLaMA • • 5h ago

Question | Help Upgrading from 1xR9700 to 2xR9700. Thoughts on build before buying?

7 Upvotes

Currently have a R9700 build with 64GB RAM. Want to get a second one for both speed and the ability to run Qwen3.8-Flash-Next at Q4/higher quants of 3.8 27B. Looking for thoughts or any improvements before buying the rest of the parts.

Made sure to get a motherboard that supports x8/x8 bifurcation. Some prices, like the RAM, are very cheap as I bought them two years ago and I'm upgrading from a previous build.

PCPartPicker Part List

Type Item Price
CPU AMD Ryzen 9 7900X3D 4.4 GHz 12-Core Processor Purchased For $470.00
CPU Cooler Noctua NH-L12 Ghost S1 37.8 CFM CPU Cooler Purchased For $92.00
Motherboard Gigabyte B850 AI TOP ATX AM5 Motherboard $519.98 @ Newegg Canada
Memory Kingston FURY Beast 64 GB (2 x 32 GB) DDR5-6000 CL30 Memory Purchased For $295.00
Storage Western Digital Black SN770 1 TB M.2-2280 PCIe 4.0 X4 NVME Solid State Drive Purchased For $110.00
Video Card ASRock Creator Radeon AI PRO R9700 32 GB Video Card Purchased For $1900.00
Video Card ASRock Creator Radeon AI PRO R9700 32 GB Video Card $2499.99 @ Newegg Canada
Case GameMax MeshBox Pro ATX Mid Tower Case $133.98 @ Newegg Canada
Power Supply be quiet! Power Zone 2 1200 W 80+ Platinum Certified Fully Modular ATX Power Supply $249.90 @ Amazon Canada
Case Fan ARCTIC F12 53 CFM 120 mm Fans 5-Pack $36.99 @ Amazon Canada
Prices include shipping, taxes, rebates, and discounts
Total $6307.84

r/LocalLLaMA • • 1h ago

I Built A Thing Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

• Upvotes

Following on from club-5060ti and club-rdna16, I’ve put together Infermeld: a small, open-source Linux companion kit for running one GGUF across an AMD GPU and an NVIDIA GPU, powered by llama.cpp.

I’m the maintainer. This is an experimental v0.1.0 release, and I’m looking for people with other mixed GPU combinations to help reproduce the setup and find the rough edges.

The idea is practical: if you already have cards from both vendors, can you put them to work together without buying a matching pair?

What Infermeld adds

The inference engine is llama.cpp. Infermeld isn’t a new backend, and I’m not claiming to have invented mixed-GPU inference.

It packages the supporting pieces around that setup:

  • Explicit AMD/Vulkan + NVIDIA/CUDA device selection and runtime preflight.
  • Reproducible build instructions and inspectable launch arguments.
  • A read-only thermal guard, with shutdown limited to the server process it started.
  • Documentation and a results site that keep configurations, failures and limitations visible.

The release is source-only. You build the documented llama.cpp revision separately and supply your own model weights. It’s intended for people comfortable with an experimental Linux setup, not as a one-click installer.

Current tested setup

Component Tested configuration
AMD GPU RX 6900 XT, 16GB
NVIDIA GPU RTX 3080, 10GB
Model Qwen3.6-35B-A3B, UD-Q4_K_M GGUF
Backends Vulkan + CUDA
Split mode Layer
Context reservation 8,192 tokens

The acceptance checks include loading and short completions with MTP off and on.

That’s a narrow result on one hardware pair, not broad compatibility testing. An 8K context reservation is not the same as testing a filled 8K prompt, and a short successful response is not a sustained performance benchmark.

Important limitations

  • Sustained Q4 throughput and full-length high-context results are not yet qualified.
  • Historical measurements are labelled with their original configurations. They should not be read as performance numbers for the current Q4 setup.
  • There’s no promise that combining cards is faster than using one.
  • Adding the advertised VRAM capacities does not guarantee that all of it is usable for the model and its runtime allocations.

I’d rather make those boundaries clear than present a successful load as a complete benchmark.

Looking for other AMD/NVIDIA combinations

Successful runs and failures are both useful. If you try it, please include:

  • Both GPU models and their VRAM sizes.
  • OS, driver versions and llama.cpp revision.
  • Model and quantization.
  • Launch settings, including the split and context reservation.
  • How far it got: preflight, loading, first completion or a longer workload.

There’s a hardware/result issue form in the repository. Please sanitize paths and keep credentials and private logs out of reports.

Repository and setup instructions:
https://github.com/5p00kyy/infermeld

Results and evidence:
https://5p00kyy.github.io/infermeld/

Anyone already using an AMD/NVIDIA pair for local inference? I’d be interested in what works for you, and where this setup breaks on different hardware.


r/LocalLLaMA • • 7h ago

Discussion Is Strix Halo (GMKtec EVO-X2, etc.) the closest thing we have to a "dream" local LLM box?

7 Upvotes

I've been looking at the <32B model space and keep coming back to an interesting question.

A few years ago, projects like Hummingbird+ suggested that cheap custom accelerators (FPGA-based) might become the future of local inference. But today it seems like memory capacity is still the real bottleneck rather than raw TOPS.

For someone who wants to run modern 20B-32B models at reasonable quants (Q5/Q6 rather than INT4), the options all seem compromised:

  • Consumer GPUs have great bandwidth but limited VRAM.
  • NPUs and AI accelerators often have lots of compute but not enough memory.
  • FPGA solutions are fascinating but still bandwidth-constrained.
  • Strix Halo systems (GMKtec EVO-X2, Framework Desktop, etc.) offer huge unified memory pools, but they're expensive.

The "dream" accelerator would be something like:

48+ GB memory
500+ GB/s bandwidth
under $1000
reasonable power consumption

...but I don't think anything like that actually exists yet.

For those who have used Strix Halo systems for local inference:

How do they feel with current 20B-32B models?

Do you regret not buying a used 3090/4090-based machine instead?

Is unified memory a bigger advantage in practice than benchmarks make it seem?

Curious what people who own both types of systems think.


r/LocalLLaMA • • 20h ago

Funny Need maybe say "Use llama.cpp"

69 Upvotes

So I tried that miracle engine everyone is talking about.

Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:

Can you help with the following problem?

So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).

The thinking trace:

We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated? 

...

10k tokens later it degrades to:

Need maybe maybe include "Use 'for code, list constraints'."
Need maybe maybe include "Use 'for code, list requirements'."

The same exact model in llama.cpp does produce a coherent answer without a doom loop.


r/LocalLLaMA • • 16h ago

Question | Help Is there any way to improve creative writing for Local Models (Qwen)?

38 Upvotes

I wanted to know if theres anything that can improve creative writing for our Local Models? I specificly am asking for qwen models as they are way better when it comes to researching and writing html files compared to gemma / muse (which I know are better at creative writing). Im currently using qwen 3.8 flash next at q4 with llama.cpp


r/LocalLLaMA • • 14h ago

Discussion How long before we have a local model capable of modeling?

20 Upvotes

I'm using Qwen 3.8 and qwen flash on a 5090. It's miles behind the latest Opus 5.5. Even if it was remotely capable it would be a huge help to me, but for now, with regards to 3D modeling, local models are not close at all.


r/LocalLLaMA • • 9h ago

Resources Poor People Vulkan GPUs list

8 Upvotes

Help with this list. Give me your recommendation on "not supported anymore" GPUs. Looking for budget and Vulkan friendly options.

Most of the GPU are not supported by latest CUDA / ROCm. Often with some witchcraft magic they are able to run with native backend. I prefer the simplicity offered by running Vulkan backend. I'll successfully ran GTX 1080Ti, P102-100, and MI50 on a single system thanks for Vulkan and Linux. Gemini helped with data gathering.

Here is the filtered table including only NVIDIA GeForce GTX series GPUs with a memory bandwidth of 256 GB/s or greater and at least 8 GB of VRAM:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
GeForce GTX 1070 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1070 Ti 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1080 8 GB 320.3 GB/s 256-bit GDDR5X
GeForce GTX Titan X (Maxwell) 12 GB 336.5 GB/s 384-bit GDDR5
GeForce GTX Titan X (Pascal) 12 GB 480.0 GB/s 384-bit GDDR5X
GeForce GTX 1080 Ti 11 GB 484.4 GB/s 352-bit GDDR5X
GeForce GTX Titan Xp 12 GB 547.7 GB/s 384-bit GDDR5X

The table below lists the specifications for the specialized datacenter, enterprise, and crypto-mining NVIDIA cards you mentioned, applying your rule of maintaining a memory bandwidth greater than or equal to 256 GB/s and filtering for 8 GB or more of VRAM.

All five models successfully qualify:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus/Architecture
NVIDIA P104-100 8 GB 320.3 GB/s 256-bit GDDR5X Mining (Pascal)
Tesla M40 12 GB / 24 GB 288.4 GB/s 384-bit GDDR5 Datacenter (Maxwell)
Tesla P40 24 GB 347.1 GB/s 384-bit GDDR5 Datacenter/AI (Pascal)
NVIDIA P102-100 10 GB 400.0 GB/s 320-bit GDDR5X Mining (Pascal)
NVIDIA CMP 50HX 10 GB 560.0 GB/s 320-bit GDDR6 Mining (Turing)

Here is the updated list of classic NVIDIA Quadro enterprise workstation cards, continuing to filter for at least 8 GB VRAM and a memory bandwidth of 256 GB/s or greater:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Architecture
Quadro K6000 12 GB 288.0 GB/s 384-bit GDDR5 Kepler
Quadro P5000 16 GB 288.4 GB/s 256-bit GDDR5X Pascal
Quadro M6000 12 GB / 24 GB 317.4 GB/s 384-bit GDDR5 Maxwell
Quadro P6000 24 GB 432.2 GB/s 384-bit GDDR5X Pascal
Quadro GP100 16 GB 716.8 GB/s 4096-bit HBM2 Pascal

With the GV100 out of the picture, the Quadro GP100 and Quadro P6000 are now the highest-end entries remaining on this specific filtered list.

Here is the updated AMD Radeon desktop GPU table with all RX 6000 and RX 7000 series models removed, while still filtering for a minimum of 8 GB VRAM and 256 GB/s memory bandwidth:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
Radeon RX 480 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 580 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 590 8 GB 256.0 GB/s 256-bit GDDR5
Radeon R9 390 8 GB 384.0 GB/s 512-bit GDDR5
Radeon R9 390X 8 GB 384.0 GB/s 512-bit GDDR5
Radeon RX Vega 56 8 GB 410.0 GB/s 2048-bit HBM2
Radeon RX 5700 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX 5700 XT 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX Vega 64 8 GB 483.8 GB/s 2048-bit HBM2
Radeon VII 16 GB 1,024.0 GB/s 4096-bit HBM2

Note: MI50 and the Radeon VII, Radeon Pro VII share same firmware.

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus / Architecture
Radeon Instinct MI25 16 GB 484.0 GB/s 2048-bit HBM2 Machine Learning (Vega 10)
Radeon Instinct MI50 16 GB / 32 GB 1,024.0 GB/s 4096-bit HBM2 Datacenter AI (Vega 20)

Top Contender: AMD Instinct MI50 16GB. Current used market on MI50 16GB is around $150.


r/LocalLLaMA • • 21h ago

New Model bilibili released Index-Translate,a A Multilingual Translation Model Family based on Qwen3.5

59 Upvotes

https://github.com/bilibili/Index-Translate

Index-Translate is a family of multilingual translation models built on Qwen3.5. The text models cover 150 languages and follow translation instructions such as terminology, formatting, and content-preservation requirements. The family extends this foundation to speech, syllable-controlled translation, and full-document translation.

-Index-Translate translates text, structured content, and community expressions. -Index-Echo produces translated subtitles or speech conditioned on the source speaker's voice. -Index-Homura adjusts translations toward a specified target syllable count. -Index-NativeLong translates complete documents with context across passages.


r/LocalLLaMA • • 13h ago

Discussion A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

12 Upvotes

I forked Strata and worked with Opus 5.5 with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s and the nvidia drivers do allow unified memory access.

The machine I have has 4 x 16GB GPUs, llama.cpp had like terrible results before I started this journey, it produced 130tk/s prefill and 15tk/s decode.

So I was fighting Opus the whole weekend, beating it with facts and logic, like FP16 instead of BF16, memory management, expert caching on GPU, better NVLink usage, Tensor Core utilization rather than CUDA core ops. Claude was great at iterating, executing nsight nsys to debug time gaps.

  • Prompt reading: 7,350 tok/s peak, still 7,090 tok/s on a 252K-token prompt (35 s)
  • Generation: ~113 tok/s peak (JSON), ~100 on code, ~84 on prose (MTP speculative decoding)
  • Follow-up at 252K depth: first token after 0.26 s, 60 tok/s
  • All 72 GiB of experts page-locked in RAM across both sockets; GPUs pull from NVLink 2.0 at ~70 GB/s each

I will try to contribute back some of the changes, but I suspect Strata will remain consume-hw-first inference engine, and that is totally ok, Niko1221 did a great job

The forked repo: github.com/eelgaev/Strata-AC922


r/LocalLLaMA • • 1d ago

Discussion The curse of 64GB system RAM

84 Upvotes

Not a bot. Not a Strata shill. Just sharing my experience.

So, I have an R9700 in my machine, plus an RTX 5060 Ti, and 64GB DDR5 system RAM. Overall, not a bad setup. Anyway, I mainly run a daily driver local LLM on the R9700 while running image/video inference on ComfyUI on the 5060 Ti. Mostly shit like Minimax H3 which also takes a fuck-tonne of system RAM. I've been using Qwen3.8-27B at Q6 as the daily driver on the R9700 and running that around 35 t/s, which is fine for me as a daily driver. Before Strata I tried running Qwen3.8-Flash-Next on both cards on vulkan at a IQ4_XS (or whatever that quant is called - the ~93GB one) and that only got me like 15 t/s, which I can't daily drive, so I put it down and wasn't really interested in it. Anyway, Strata comes out and people are claiming QFN is usable on much more modest hardware, so I check it out and see that mostly people are running the IQ3_XXS quant which is like ~70-something gigabtyes, so of course it's faster. Anyway, I benchmarked that quant on llama.cpp first running it just on system ram + the R9700 and it came in at 21 t/s... that's right around the absolute minimum of what I'd accept for a daily driver, but not super compelling tbh. Then I tried the same quant on Strata and I get ~60 t/s. Very fucking compelling. I 100% want to daily drive this now. The problem is with QFN loaded in Strata my system RAM usage is at 96%. I can't fucking run Minimax H3 in ComfyUI on the 5060 Ti because that shit eats a lot of system RAM too.

I feel blessed that I can finally run this epic model, and fucking cursed that I have to choose which workload to run!

Also, before anyone says 'just upgrade to 128GB of RAM bro'... I know, I know. I would but I can't afford to the jewelry and international trips my wife requests for fairness reasons to balance out all the toys I've bought this year.

Crying in 64GB of RAM.

Edit: thanks to a few suggestions in the comments I actually got Qwen3.8-Flash-Next IQ3_XXS and Minimax H3 inference working concurrently at about 90% system RAM used! On the Strata side I I think I needed --mmap-experts --resident-cpu-experts and --expert-cache auto, and on the ComfyUI side I needed --fast-disk. I tested both running fully concurrently and checked Strata's monitoring tab, and saw that the node that loads the H3 weights causes NVMe reads to hit a sustained 1GB/s for a short while, which can cause the inference on Strata to drop to around 25 ~ 40 t/s range (it fluctuated a lot during that), but then after that when H3 was actually doing the inference I saw NVMe reads sitting at about a sustained 30 MB/s and QFN inference was running between 50 ~ 60 t/s. I'd say it's a huge win. For reference my standard test when testing out an LLM is just 'write me a browser game', so I did that since I'm familiar with the quality of the expected output at this point, and also generated a 10 second clip at 0.4MP resolution. The actual wall-clock generation time for H3 was pretty much unaffected (around 400 seconds), which is nice too! Maybe some very minor performance hit, but only that.. pretty minor.


r/LocalLLaMA • • 1d ago

Discussion Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4?

117 Upvotes

Given the commentary on the Q3.8FN release page here https://qwen.ai/blog?id=qwen3.8-flash-next I assume/hope that all the work that's going on to optimise the hell out of running it will be useful when Qwen4 drops?