r/LocalLLaMA • u/Cautious_Chicken_604 • 1d ago
Discussion The curse of 64GB system RAM
Not a bot. Not a Strata shill. Just sharing my experience.
So, I have an R9700 in my machine, plus an RTX 5060 Ti, and 64GB DDR5 system RAM. Overall, not a bad setup. Anyway, I mainly run a daily driver local LLM on the R9700 while running image/video inference on ComfyUI on the 5060 Ti. Mostly shit like Minimax H3 which also takes a fuck-tonne of system RAM. I've been using Qwen3.8-27B at Q6 as the daily driver on the R9700 and running that around 35 t/s, which is fine for me as a daily driver. Before Strata I tried running Qwen3.8-Flash-Next on both cards on vulkan at a IQ4_XS (or whatever that quant is called - the ~93GB one) and that only got me like 15 t/s, which I can't daily drive, so I put it down and wasn't really interested in it. Anyway, Strata comes out and people are claiming QFN is usable on much more modest hardware, so I check it out and see that mostly people are running the IQ3_XXS quant which is like ~70-something gigabtyes, so of course it's faster. Anyway, I benchmarked that quant on llama.cpp first running it just on system ram + the R9700 and it came in at 21 t/s... that's right around the absolute minimum of what I'd accept for a daily driver, but not super compelling tbh. Then I tried the same quant on Strata and I get ~60 t/s. Very fucking compelling. I 100% want to daily drive this now. The problem is with QFN loaded in Strata my system RAM usage is at 96%. I can't fucking run Minimax H3 in ComfyUI on the 5060 Ti because that shit eats a lot of system RAM too.
I feel blessed that I can finally run this epic model, and fucking cursed that I have to choose which workload to run!
Also, before anyone says 'just upgrade to 128GB of RAM bro'... I know, I know. I would but I can't afford to the jewelry and international trips my wife requests for fairness reasons to balance out all the toys I've bought this year.
Crying in 64GB of RAM.
Edit: thanks to a few suggestions in the comments I actually got Qwen3.8-Flash-Next IQ3_XXS and Minimax H3 inference working concurrently at about 90% system RAM used! On the Strata side I I think I needed --mmap-experts --resident-cpu-experts and --expert-cache auto, and on the ComfyUI side I needed --fast-disk. I tested both running fully concurrently and checked Strata's monitoring tab, and saw that the node that loads the H3 weights causes NVMe reads to hit a sustained 1GB/s for a short while, which can cause the inference on Strata to drop to around 25 ~ 40 t/s range (it fluctuated a lot during that), but then after that when H3 was actually doing the inference I saw NVMe reads sitting at about a sustained 30 MB/s and QFN inference was running between 50 ~ 60 t/s. I'd say it's a huge win. For reference my standard test when testing out an LLM is just 'write me a browser game', so I did that since I'm familiar with the quality of the expected output at this point, and also generated a 10 second clip at 0.4MP resolution. The actual wall-clock generation time for H3 was pretty much unaffected (around 400 seconds), which is nice too! Maybe some very minor performance hit, but only that.. pretty minor.
152
u/PraiseThePidgey 1d ago
Crying in 32GB of RAM
51
u/Illustrious_Ant_9242 1d ago edited 1d ago
I just got some extra DDR4 Sticks already... I might skip ddr5 entirely.
in june 2026 when local AI turned out to be actually useful, it was still possible to get 2x32G DDR4-3200 for around 250 bucks total. The idea is to survive the next 2-3 years until better hardware comes along at reasonable pricing.
22
u/No-Yak4416 1d ago
June is when I built my PC, blissfully ignorant of what was to come in the months that lay ahead. I built with 32GB DDR5 and a 3090, but looking back I kick myself for not doing at least 64GB, even 128, 256… heck, even dumping my life savings into DDR5 RAM.
Edit: I misread your comment. I built in June of 2025, so even DDR5 was cheap back in those good ole days
3
u/Thistlemanizzle 1d ago
I was thinking June!?!! of 2026? Did memory prices wildly spike yet again? (They have a bit).
3
u/Illustrious_Ant_9242 1d ago
yes absolutely, the same RAM would cost 2x now (again)
i get all that second hand stuff on ebay obviously. Also added another RTX3060 12G, they're still somewhat reasonably priced for how they perform.
3
u/Thistlemanizzle 1d ago
I mean, why not just invest in RAM sticks...
5
u/Illustrious_Ant_9242 1d ago edited 1d ago
the house of cards will collapse eventually, since shortages in markets typically don't last forever.
It's literally 5+ year old hardware, there will be newer releases and some competition eventually.
I actually made money by driving Teslas during the 2022 car shortage and literally sold it used at a 11000€ markup. needless to say, the same car could be had for 15 grand less nowadays. I also sold my used A/C during the 2026 heat wave at a 50% markup. The craze has calmed down by now
5
u/Thistlemanizzle 1d ago
I know, but it would be hilarious to double your money buying and then selling the very thing you so desperately desire.
2
u/NineThreeTilNow 23h ago
Yeah, I decided on a whim that I needed to upgrade my system in early October of 25.
I was like "These SSD and RAM prices are good."
I only got 64GB of DDR5 and a 4TB NVMe, but looking at that now... It would cost a fortune. Then? It was like 450 with tax or something.
Then again, I paid a pretty ridiculous amount of money for my first 1gb of memory in a system... So it will sort out eventually.
2
u/FatheredPuma81 19h ago
I built in Nov 2024 and man let me tell you. 48GB 8000MHz for under $200. I regret not buying 96GB even if 8000MHz would have been off the board.
3
u/OddUnderstanding2309 1d ago
I bought 64GB DDR4 cl19 3800MT second hand for 290€.
Thats 100% fine with meDDR5 is just hype, especially if you run 4800MT cl 38 or some shit
4
u/Illustrious_Ant_9242 23h ago
yes the main reason why the price hiked is because stuff actually got useful. therefore I'm willing to pay for it as long as I don't feel completely ripped off and the product does what it is supposed to do - no need for the cutting edge stuff with marginal improvements
2
u/575_Inverse 8h ago
my thought as well. But right now I see the 5060ti 16GB running at ~800bx. Any more than that and I'd nope out. DRR5: since yhe la hike it's too much of a rip off. I'll stick with the current 128GB, altough adding 64 more to reach 192GB would be awesome... but guess I'll give up. This is turning into a circus.
8
2
u/inthesearchof 1d ago
Cry tears of joy one more 32gb stick might be the lowest price for sota entry
2
u/Cautious_Chicken_604 1d ago
There's a lot to cry about it. Truly.
1
u/Material-Database-24 19h ago
Just run 27B. It's only 1-15% in most benchmarks from FN. And at IQ3_XSS FN already hangs on the edge of lobotomization, so the difference is even smaller against Q4.
And you should run it at Q4_K_XL, you get more speed with quality drop that fades into noise between test runs.
3
u/Cautious_Chicken_604 18h ago
I've been running 27B already. I find in agentic coding there is definitely a big difference in output between 27B and QFN. I want to use QFN, since my subjective experience of using it for my tasks it gives me better results.
2
u/Material-Database-24 18h ago
I have very different experience. The heavily quantized models of FN I can run on my 64gb UMA often fail compared to 27B.
For example, 27B one shotted well working tetris on python. No bugs, nothing. FN coder failed. A lot of simple bugs, like wrong case on pygame's globals etc. I have very hard times to understand this hype that heavily quantized FNs would be "far better" than 27B.
3
u/Cautious_Chicken_604 17h ago
I don't use the FN Coder one. That has a huge amount of the experts stripped out.
1
u/Material-Database-24 16h ago
Yeah, but here's many folks who tell it's great. It isnt.
But I'll give another chance for IQ2_XS.. it's pretty much the biggest I can comfortably get in use. I could try IQ3_XSS but that will take the whole system memory
1
u/Cautious_Chicken_604 16h ago
I don't think I'd go lower than IQ3_XXS. Tbh I really want to run at least one of the Q4 ones, but so long as what I can run doesn't suffer major problems and performs better on my usecase, then I'm not going to be super picky about it. If you're legit experiencing problems with a particular quant though, yeah that doesn't sound like a good one to use.
2
u/Material-Database-24 14h ago
Got IQ2_XS running, and it was clearly better than FN-coder.
However, it is not quite up to 27B either. At least with short testing. I can see that it has more "world knowledge" than 27B, but it's even worse overthinker.
That said, with this experience, I can see that IQ3_XSS and above are better than 27B at Q4, especially if general purpose use is desired.
1
u/Material-Database-24 15h ago
I believe the IQ3_XSS is the lowest one should go. I am about 5min from trying IQ2_XS, so we'll see.
1
u/DjebbZ 12h ago
You can configure Strata to keep some experts out of RAM to save it. I've seen someone with the same issue somewhere on Reddit that used the 2/3 flags required , which allowed him to run both Strata and ComfyUI at the same time since he had enough RAM after. Sorry on my phone, can't tell you the exact flags, but look for something like
--experts-keep-moe.1
u/Material-Database-24 12h ago
I do not have Nvidia, so no strata for me. But I'd assume we see similar tricks from llama etc in future.
It'd be great feature to drop MoEs off RAM that rarely get action.
1
1
u/Cautious_Chicken_604 11h ago
I'm literally running Strata on my R9700 getting 60 t/s out of it now. It supports AMD. Doesn't support Intel cards yet, though, and doesn't support Vulkan, so you can't do AMD+NVIDIA setups, but AMD support is decent!
→ More replies (0)1
1
u/Cautious_Chicken_604 14h ago
You might not need to cry all that badly if you have a really fast NVMe. Fucking glad as fuck I upgraded mine to a properly performant PCIe Gen5 one recently. That's what made this actually work in the end. Digging through things it looks like there are some options that can run things with lower RAM requirements relying more on the NVMe.
→ More replies (4)1
16
u/brainchillzZ 1d ago
Put all your ai shiz in an older computer with ddr4 ram that you can still afford to do 128 in …
9
u/Lurksome-Lurker 1d ago
lol DDR4 is now relegated to older computers? Way I see it, RAM bandwidth hasn’t been an issue in ages and DDR5’s main gain is higher capacity. Which ironically is unobtainable at a reasonable price.
If it wasn’t for DDR3’s cap at 32GB (Yes, I am aware of enterprise grade and very last gen CPUs doing 64GB) I would bet a good chunk of us would be reviving 4th or 6th gen processors and pairing them GTX10xx or RTX20xx cards.
2
u/DerpityHerpington 22h ago
It wasn’t just Xeons and final-edition models that supported 64 GB. Sandy Bridge-E also did, for example. Yes, they were relabeled Xeons, but they were sold as HEDT/gaming chips, so hardware availability on eBay shouldn’t be terrible, which is the point I’m guessing you’re getting at by mentioning enterprise CPUs.
1
u/Minimum_Chemical_428 7h ago
But for the strata it matters right? Cause people are just using ram and the speed difference is probably important?
I'm still trying to find out if I should buy some additional ddr4 ram and try strata over buying a 3090 for 27B.
1
u/Lurksome-Lurker 5h ago
So what I am about to advise is probably going to be polarizing
I don’t think its gonna matter as much as people think it will. If you assess your overall system the slow paths are going to be NVme/SATA speeds to your storage and PCIE lane speed and width. Basically pushing data to and from various places around the MOBO. In that context, RAM bandwidth is only one segment in that pipe.
Without doing some digging, I would wager your RAM could be really slow and your bottle neck is still going to be the PCIE bus on its x16 lanes or bifurcated x8/x8 split between GPUS.
So now comes the cost/risk analysis between DDR4 RAM and RTX3090s . Given current economics, these purchases are currently stable or appreciating assets. (i.e. you should be able to resell them for what you bought them for or more) So your risk is basically 0 provided the value doesn’t crash any time soon.
So, your real consideration is investment costs. You can at present resell and make back the money spent in the HW assuming you don’t break anything. RAM is retailing at 16GB sticks $100 by 4 of them of the same brand in the same purchase. (This minimizes risk of timing miss match of not buying a set. I personally never have had an issue doing this)
An RTX3090 is retailing North of $1000 new and used.
so I would buy four 1x16gb sticks for ~$400 and test it out, if it works either keep the sticks or sell them and get a $600 4x16GB matched timing kit. If it doesn’t work then sell it and get the RTX 3090
1
u/Cautious_Chicken_604 1d ago
That plan is itself probably enough of a capital outlay that just buying the RAM upgrade is easier.
14
u/BAL-BADOS 1d ago
Crying in 16GB of RAM
1
u/Cautious_Chicken_604 1d ago
It must really hurt to have to cry 4x harder than this.
3
u/BAL-BADOS 23h ago
Never expected I needed more than 16Gb of RAM. Too late now due to RAM crisis
2
u/Cautious_Chicken_604 23h ago
I've run 64GB of ram for quite a long time now, and wanted 128GB for quite a long time too. I figure we're just always going to need more now. I remember back in around... 2021 I think... hearing a colleague of mine mention his PC had 128GB of RAM. Bet he's smiling now, even if it is older, slower ram.
1
u/BAL-BADOS 14h ago
Before AI, what do you do with 64GB? Most demanding for me is games but it all worked with 16GB
1
64
u/johndeuff 1d ago
omfg did mods get paid to let this bot marketing go on?
11
u/3rdPoliceman 20h ago
My "not a bot" preface is raising questions answered by my "not a bot" preface
3
u/chensium 23h ago
lol seriously. randos with no history spelling "fuck-tonne". haha it's hilarious.
25
u/GowsenBerry 18h ago
OP has been posting on here for months, can you guys stop talking out your ass.
"Tonne" - oh no is he from the UK! Your point literally makes no sense, who is upvoting this shit.
→ More replies (8)14
u/IntQuant 23h ago edited 20h ago
Is it that hard to believe that people would spread the word of the software they like?
Strata really is multiple times faster, too bad it consistently hangs on a second request on my machine and I'm yet to figure out why.
Edit: of course I'm not gonna get any response an would simply get downvoted instead.
10
6
u/Useful_Disaster_7606 1d ago
I'm kinda in the same boat.
I always had to remind Qwen not to do any tests because it just might kill itself accidentally.
4
u/Xonzo 17h ago
Are their any benchmarks showing QFN IQ3_XXS clearly much better than 27B at Q6? I’ve got a 5070ti + 5060ti / 64GB DDR5, but holy hell adding another 64GB is EXPENSIVE. With Qwen 27B I’m getting 1500+ PP (not optimized, only P2P patched) and 50-70 TPS.
9
u/No_Drag_5205 15h ago
I used to run Qwen 27b at Q6, now I run QFN at Q3 xxs it's way more impressive. Just do the test.
4
u/Selfmade_Watchmaker 14h ago
People focus on quant over the architecture.
QFN just simply packs so much more intelligence that even at lower quants it’s more capable than 27b
2
u/Cautious_Chicken_604 16h ago
Somewhat. The best comparison is just to try compare it against 27B at Q8 or BF16 as an upper-bound. Q6 isn't drastically far off Q8, so it's a decent proxy to it. I had this question too and the answer did appear to be that it comes out slightly ahead. I'd rather run a Q4 quant of QFN, but I'm banking on being able to either use Android phones or my 5060 Ti to help out with inference in some capacity to boost it all the way there for when I just want to run it at max quality.
19
u/KnownAd4832 19h ago
Owner of Strata here - I’m working on two modes currently where you choose to leave yourself 10-20% breathing room (pc usage) or full speed (for inference). Will be in Monday-Tuesday as scheduled.
1
u/Cautious_Chicken_604 14h ago
Yo - I updated the post body with the options I used to actually get QFN inference and Minimax H3 inference working concurrently on my system. Fucking goated that this works! I dunno if it's going to wreck my disk, but I don't think I care that bad, since that's cheaper than GPUs and RAM to replace. Not but much actually... but still.
1
u/KnownAd4832 13h ago
Disk has very little to none usage, usage goes higher only if it spills (you dont have enough VRAM+RAM) combo.
1
u/Intrepid_Yak_5776 18h ago
Ciao complimenti x il lavoro. Ho un rog x399 con 3 3090 1 5070ti da 16gb 256 ram e con qwen q4 xxl faccio dai 40 ai 50 t sec. E minsembra allucinante che gente con 1vga faccia meglio delle mia configurazione. Detto questo bisogna provare se con configurazione migliore si ottiene ancora di piu
1
u/KnownAd4832 17h ago
You can run Strata with 1 gpu (5070ti) and use 3 3090 to run other model in parallel. You will see heaven :)
→ More replies (2)1
u/lukaskel 16h ago
Out of curiosity as you seem active.
I wanna dabble in local llm models and have 64gb ddr4 ram still, so this sounds very promising.
Now I’m just trying to decide between a 5060 Ti 16GB or a 4070 Ti with 12Gb.
What would you recommend if I wanna experiment mostly, as the rig is for casual “gaming” at the end of the day?1
3
u/grabber4321 1d ago
get 32GB, its still 96GB and gonna be enough to cover the good quant (if you have 2 sticks in your rig now)
3
u/k8-bit 21h ago edited 19h ago
Similar setup to you. R9700 attached to a miniPC with 64gb ddr5 has become the brains orchestrating my main server (3xRTX 3090, 128GB DDR4, Lots of storage) used for minimax h3, LTX2.5, music lora creation, image gen, and comic book generation processes tasked between three GPUs, o e of which is prioritised for a subwave radio station server)
I hit a ram ceiling a lot with the 3090 machine because of video tasks.
The r9700 machine has been running q38-27b vllm-sly-radiance for a stream 262K context with vision about 50-70ts. I'm super jealous of strata, which i have installed but is limited to 1 stream, no good for my deepseek harness + Hermes workflows. In an ideal world I'd run a 2nd r9700 with another chunk of ram, but those days are gone. I bought all my hardware before costs went mental last year.
2
1
u/alex_bit_ 18h ago
How do you run minimax h3 on this?
1
u/k8-bit 17h ago
Sorry, probably easier for me to let it explain at length 😅
How I run AI video and music from an agent box on an AMD card, with the rendering done on 3090sHardware
Bubblina (AMD): a mini-PC with a Radeon AI PRO R9700 (32 GB) in an OCuLink dock, running ROCm. It only does the thinking. It serves the agent LLM (Qwen 27B-class, 200K+ context, vision) behind one OpenAI-style endpoint. The agent harnesses I drive it from (DSH, Hermes) use it for planning, scripting, prompt writing and checking results.Trogdor (NVIDIA): an Unraid box with 3× RTX 3090 (24 GB each) and 125 GB RAM. It does all the heavy generation. The three cards are not equal:
GPU2 and GPU1 sit on PCIe 4.0 x8. They each run a ComfyUI instance. GPU2 is the main one and GPU1 is the overflow. H3 video goes to GPU2 by default, and LTX goes to GPU1.
GPU0 sits on a PCIe 3.0 x4 slot behind the chipset, so it's the slowest. It runs the radio station and TTS (always on). It also has a third ComfyUI instance for image jobs, but that's only used when both main lanes are busy and GPU0 has at least 18 GB free.
There's no P2P between the cards, so each job stays on one GPU.The pipeline
In DSH or Hermes (running on Bubblina, the Mac or Trogdor), I describe the piece in plain language.
The agent loads a skill, which is a spec-file CLI (h3-longform-video or ltx-longform-video). It writes the shot spec, generates reference images and keyframes, then submits everything to ComfyUI over HTTP.
Clips are chained into one long piece (5–15 s each, up to several minutes), then assembled and mixed. Music, titles and the score come from the same skill family.
A lock file allows only one video render at a time, because two overlapping renders ran Trogdor out of RAM and forced a power-off (twice).H3 vs LTX
MiniMax H3: native synced audio and voice cloning from reference clips. It uses an int8 VAE swap (about 2× faster encode) and fp16 accumulation. FastH3 V2 renders a 5 s, 8-step clip in about 87 s. There are optional post passes: VOSR2 upscale, and a face-refine pass for warped faces (it costs about 2.4× the base render per subject).
LTX-2.5 (22B): it generates audio and picture together. Identity locking via reference subjects holds up to 4 characters across clips, including close-ups. It only takes 2 voice references per clip, and the rest get invented voices.
Where other work runs: image generation (Qwen-Image, Krea2) and music (YuE2, MiniMax Music 3, ACE-Step) go through the same ComfyUI lanes. Because of the one-render rule, video waits for those or the other way round. TTS and radio stay on GPU0 and don't compete.
What I learned
The agent's own LLM doesn't need the big GPUs. Keeping it on the AMD card leaves every 3090 free for generation.
The main cost is serialising work, not raw speed. Without the lock I crashed the box twice.1
3
u/Most-Trainer-8876 21h ago
I can feel ya, what I do is run ComfyUI in one GPU, and QFN in another. And enable low ram mode (`--mmap-experts`) in Strata, so that if comfyui runs, windows will automatically remove strata from ram on demand to make room for comfyui. So during generation, if strata model request is triggered, speed drops to about 15tk/s, which is fine.
In hermes, my model randomly decides to check on ComfyUI for current status, since comfyui is currently running, model speed drops from 35tk/s to 15tk/s.
1
u/Fuk_The_Rich 15h ago
So I seen you mentioned Hermes and thought I'd ask you a couple questions. I have Hermes agent on the computer that my boss owns and I'm using it to learn how to do trading using unusual whales and at some point I'm going to hook it up to interactive brokers. My question is how did you set it up to begin with if you use it as your main AI source? The computer we have is a Nvidia RTX pro 6000 Blackwell workstation with a thread ripper build it's pretty expensive but my boss wanted to get the best for just a regular desktop computer not server level. I use Hermes agent on a 120b model and it works really fucking fast. I tried a few things to begin with and things didn't work out well so I started from basics and I had Hermes learn about itself first so it downloaded libraries full of information about how to use its own tools and skills and I used grok bot to install everything. I'm not a computer programmer I'm literally just learning how to use computers in the last few months and any help would be greatly appreciated
3
u/Max-_-Power 21h ago edited 21h ago
This is an easy fix:
bash
sudo bash -c "systemctl set-default multi-user.target && reboot"
Then use the AI rig remotely and have all the systems' resources ready for AI.
If you want to use it as gaming rig again, run
bash
sudo bash -c "systemctl set-default graphical.target && reboot"
and enjoy.
That's how I do it.
edit: Before anyone follows my advice, obviously make sure that you can SSH into the rig
1
u/Xonzo 14h ago
Another option (for some) is to used the integrated GPU on your motherboard. Also called Reverse PRIME or PRIME Display Offload. So for example my monitor is plugged into my 5070ti, but the integrated AMD GPU is being used. Which leaves the entire 5070ti + 5060ti for AI work, and it's trivial to select the 5070ti for rendering if needed. This works great for me.
3
u/espece-de-bon 7h ago
I'm very intrigued by this. What I learned from this post:
- I should keep the 5060 Ti even if I want an R9700 one day
- People are skeptical of this post as a
bot - I've got 96GB DDR5 RAM and now I'm curious what Strata can do. (I think I got lucky when I found a used PC a few weeks ago with a lot of system RAM).
I'm assuming we're speaking of this Strata?
1
u/Cautious_Chicken_604 5h ago
I'm pretty sure people think their own mothers are bots at this point. They can go to hell. And never, ever, ever give up a GPU. I regret every day I sold my 4090.
1
u/espece-de-bon 3h ago
That's good advice; I was thinking about it simply to partially pay for the R9700 but I "see the light" now. On that note, there are people who are willing to sell decent 8GB cards for cheap (locally, like for $100) and that would be a nice addition to my setup or a future setup.
I can only run a second PCIe lane at 4x:
CPU:
Chipset:
- 1 x PCIe 5.0 x16 Slot (PCIE1), supports x16 mode*
- 1 x PCIe 4.0 x16 Slot (PCIE2), supports x4 mode
but for $100, I might as well! Either for a smaller model running simultaneously or to split the workload.
I'm gonna read more about Strata and hopefully have some time later in the week or next weekend to set it up locally.
1
u/Cautious_Chicken_604 2h ago
Depending on your motherboard you can sometimes bifurcate the PCIe 5.0 x16 slot with risers that turn it into two x8/x8 slots. PCIe mostly affects model load and layer offloading, and you only need the max performance it can give you if going any lower than that makes a task you care about actually bottleneck. Case in point - I was running Minimax H3 on my RTX 5060 Ti on PCIe 5.0 16x assuming that running it at PCIe 3.0 x4 would be too slow due to needing to stream layers over PCIe since the weights don't all fit in VRAM. It turns out PCIe 3.0 x4 was basically enough bandwidth to handle the task, since there was a slowdown but pretty minor. Hence something like PCIe 5.0 x8 is actually quite a lot of bandwidth, and model loading only takes seconds anyway, so that isn't an issue even in a x1 slot.
I'd only buy minimum 16GB cards though.
1
u/espece-de-bon 2h ago
Ah, I'll check. For the money (riser cables and the splitter card), it might make more sense to just get a Taichi; unfortunately, I'm so close, I have an x870 motherboard but it isn't a Taichi (with 2 5.0 PCI3 x8/x8 slots).
2
u/laerciosantana 23h ago
I have the course of 16gb vram and 96gb, because I fell that mostly of models that I like target 32gb vram or 128gb ram
2
u/mrgreatheart 23h ago
I have 64Gb RAM too. Have your agent fiddle with Strata. There is a low memory mode that uses the ram but allows other processes to borrow it if needed. It will slow strata down if both are active at once, but it works well. Claude also managed to move more of the experts into my VRAM and pin less memory so I now have 14Gb RAM completely free and I’m even running the larger Q6_S quant. I have given 40Gb VRAM for Strata to use. I’m getting crazy speeds and my system has plenty of headroom.
2
u/jeremyohara450 22h ago
So know how you fee. i want more ram, but to even get more ram is like buying a new graphics card!!!!
2
u/mechkbfan 22h ago
Try this to get your RAM back but decent speeds?
1
u/Cautious_Chicken_604 20h ago
I tried this on Qwen3.8-27B but kept llama.cpp in favour of it even though llama.cpp was slower because I wanted to run it at Q6 instead.
Does it run Qwen3.8-Flash-Next? The appeal of Strata right now is finally getting usable speeds on QFN for the first time.
1
u/mechkbfan 19h ago
Int5 Paro maybe close enough to Q6 while still going faster.
I'm aware now the devs who are worked on it have seen Strata's approach, they're doing their own take specific for R9700, but I think they're starting with 2x first.
https://codeberg.org/StillDeadcode/radiance
Just came out 15mins before I made this comment :P
1
1
2
u/Affectionate-Hat-536 21h ago
You can at least cry and then upgrade.. Weeping from 64 GB M4 MacBook Pro .
3
2
u/beigepccase 21h ago
My solution a while back was to make my LLM box a dedicated LLM box so it doesn't have to do anything but LLM's, and then I just access it over my local network. Right now it's running 64GB DDR5 w/ an RTX 4090, and IQ3_S runs great on it. The system ram hits 55-57GB usage, but it doesn't matter because the box doesn't have to do anything else (that said, I am considering adding a 2x16GB kit to it so maybe I can try UD_Q4_K_XL).
My daily driver desktop is more modest hardware and still on DDR4. And that's fine because it doesn't have to run LLM's at all.
1
u/Ammoryyy 10h ago
Any benefit of doing this, currently i have main pc i7-13700KF | ASUS Z690M-PLUS D4 | RTX 4090 24GB + RTX 3090 Ti 24GB | 128GB Corsair Vengeance DDR4-3200 | FSP Hydro G Pro 1000W
I gave spare 2x32 ddr4 Corsair vengeance 3600, + team group 2*8 ddr4 + H370 mb + i5 8400 lying around, any advice?
2
u/sleight42 20h ago
Even if you could run both at once, both would run slower for hammering the PCIe bus hard to move data to and from RAM, right?
1
u/Cautious_Chicken_604 14h ago
So, I got them running concurrently and perf takes a very temporary hit during model load of H3, and other than that is likely most unaffected for both QFN and H3, which is wild. Basically need to enable the options which make things rely more on disk and less on ram on both strata and comfy, but it works!
1
u/Cautious_Chicken_604 20h ago
I used to have the 5060 Ti in the PCIe Gen5 16x slot on the motherboard and run the R9700 in the eGPU dock, which for reasons refuses to run at any higher than PCIe Gen3 4x, assuming that the smaller VRAM card would need to do layer swapping and therefore be more sensitive to the PCIe bandwidth. Just last night I decided to swap the cards around and after testing out the exact same workflow running H3 on the 5060 Ti and it went from 10 minutes to 11 minutes. That's a lot less impact than I was expecting. So, honestly, I dunno. Like, maybe... though it might also be fine.
1
u/Savantskie1 17h ago
One is probably behind the chipset and one is probably behind the cpu. That’s probably the issue.
1
u/Cautious_Chicken_604 17h ago
So many issues with eGPU dock. I regret that approach. Should have just upgraded the motherboard and case and been done with it. The only thing I haven't tried yet is a shorter cable, so that'll be my last-ditch effort to get it to run at PCIe Gen4 4x. I currently have the eGPU running via a PCIe -> Oculink adapter that, yes, is chipset connected, but that's not the only issue. I recently got an m.2 -> Oculink adapter and put it in a CPU connected m.2 slot and fucking crickets. It wouldn't detect the card AT ALL. I re-seated it all multiple times and nothing. Gave up and went back to the PCIe -> Oculink adapter. I have it in the PCIe Gen4 4x slot, but unless I force it to Gen3 in the BIOS it falls over the moment I tried to load a model on it. My advice to people now is only get an eGPU dock if you hate yourself.
1
u/Savantskie1 17h ago
Try a plx card setup. And make sure you have the power connection for the little risers you’re going to need for the PCIe slot. That’s my only blocker right now. Once I get a second psu, I will be able to borrow my add2psu adapter for keeping the two psu to turn on at the same time
1
u/Cautious_Chicken_604 16h ago
Next year I'll do a mobo / case upgrade and I'll potentially go this route I think.
1
u/SandySkittle 15h ago
Can you bifurcate your 16 slot to x8 x8? In that case you may want to consider an mcio retimer card with externally facing ports
1
u/Cautious_Chicken_604 15h ago
My board is capable of it. I already spent money going the eGPU dock route and various other things, so it feels a little hard to justify switching approach on that right now. I'll revisit my setup next year though and maybe go the route of one of those open air frames and use risers like this on a higher end motherboard.
1
u/SandySkittle 13h ago
Yeah understand. Risers can be a bit tricky, may not reliable at gen 5, hence going mcio retimer card. Those are used in enterprice setups.
2
u/MKP_Nimilka 18h ago
“Fits in RAM” and “leaves enough RAM to use the PC” really need to be separate benchmark categories. 70 tok/s is impressive, but for a daily driver I’d also want to know the speed with enough memory left for ComfyUI. That’s the result that would actually help me choose a setup.
1
u/Cautious_Chicken_604 18h ago
I've gotten some suggestions from the comments here, that I'll try out later tonight. Hopefully it gets me into that actually usable territory. If not then... I can fall back to 27B when not running Minimax H3, but I'd really like to have my cake and eat it too, so I will explore every option including the crazy ones.
1
u/MKP_Nimilka 18h ago
Fair 😂 If one of the crazy options works, please share the results. I’d be interested in peak RAM usage and tok/s with both workloads running even a slower setup would be a win if you don’t have to keep switching models.
1
u/Cautious_Chicken_604 14h ago
I updated the post body with what worked for me. Needs a decently fast NVMe, but it fucking works!
2
u/Sporebattyl 1d ago
I’m in the 64gb club too. I’m just realizing this as well.
Just curious, why do you run mini max h3 instead of just having qwen do it?
1
u/Cautious_Chicken_604 1d ago
Instead of having Qwen do what exactly? H3 generates videos. I have a workflow in Comfy that I just run to generate videos... for science reasons. The workflow is already built, so there's nothing for Qwen to do, I just click 'run' and queue up 144 odd 15 second clips to be generated throughout the day.
Actually, the thing that makes me really cry in 64GB RAM about this is actually I do sometimes get Qwen to build workflows and try new things out, but that requires the resources to be able to run both workloads at the same time. I can do that with Qwen3.8-27B but it's kinda shitty at it, even with ComfyUI MCP. So, a model capability upgrade is a real win, but... then not being able to run that simultaneously alongside H3 makes that really, really annoying.
1
u/Sporebattyl 1d ago
Ah I see didn’t realize that’s what H3 is for. I’m just starting to look into comfyUI.
3
u/Cautious_Chicken_604 1d ago
H3 is fucking wild. Especially image to video. Prompting it is kind of a bitch though, I've been running llama.cpp one the R9700 and H3 on the 5060 Ti and the ComfyUI workflow I put a node that uses llama.cpp to generate the prompt, so I can get tonnes of variety going etc. I was under the impression that I needed two cards to do that, but I'm starting to realize I can probably just get that to run all on a single card at the expense of some wall-clock time, but if I could run the same workflow in parallel on both GPUs that'd effectively double my output, so it'd be a win.
Lots of fun toys to play with.
1
u/Karmin96 23h ago
How are you all running H3 on a 5060 Ti 16 GB? diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors is 21 GB.
2
u/Cautious_Chicken_604 23h ago
Have you not used Comfy either ever or in a really long time? The layer offloading / memory management is pretty excellent these days. Has been for a while now.
1
u/Karmin96 22h ago
I’ve been using it for just under a year. But when I tried large models a year ago, the experience was terrible - it took five minutes just to edit a single image. That’s why I looked specifically for something that would fit into memory. Qwen Image 2.1 seemed like pure magic. Generation took less than a minute. It fit entirely onto a 5060 Ti (model) and a 3060 12GB (text encoder).
1
u/Cautious_Chicken_604 22h ago
Image models have gotten larger, and image generation slower as a result. I found Z-Image Turbo to be really fast. Krea2 is slower, but much better outputs, so I use that for images now. It's not quite like the old SD1.5 days.
Video is really brutal. I'm using Spectrum as the main speedup method (others use turbo loras etc) and on the 5060 Ti at 0.4MP resolution a 15 second clip takes me 10 ~ 11 minutes on my workflow. Mind you that also includes a few minutes of prompt generation via llama.cpp, so that's not all pure video generation, but still it does take a fucking long-ass time to generate videos. You can absolutely do it though! Especially if you have a good workflow with a decent yield. I can generate around ~140 15 second clips in a given day and half to two thirds of that is usable. I can get anywhere from 15 ~ 20 minutes of novel footage per day. Pretty neat. But it makes multi-gpu basically essentially, else you just can't use your PC for practically anything else at the same time.
1
u/pmttyji 21h ago
Bunch of questions dude. Only recently I got my new rig. Please reply. Thanks.
- How much VRAM(also RAM) needed for H3?
- Why not try H3 on R9700 which's 32GB VRAM? I'm asking this question because I have only one R9700 + 128GB RAM.
- How much time it takes to generate an Image?
- How much time it takes to generate a Video? I don't know what's the possible max length(Maybe 15 seconds)
- What are the suitable models(Image generation, Video generation, Animation) for R9700 + 128GB VRAM?
- Which quant is better? GGUF? I don't see GGUFs for all models so I'm confused.
- Which Quantizers suitable for R9700 + 128GB VRAM? I mean Bartowski, Unsloth for Image generation, Video generation, Animation models.
1
u/tat_tvam_asshole 1d ago
As someone who runs much bigger models than the Qwens, locally, and has them create their own ComfyUI gens, I feel ya. Agents really don't 'get' ComfyUI (i don't think it's in their training data tbh) nor do they really understand creative endeavors that well. I tend to hand them pre-configured workflows that they can riff on easily enough, but it's been a lot of handholding and having them crib my teaching 'lessons' in notes which has helped a bit, but it's still a WIP.
AI has been good for 'copy this but alter it slightly' but not really 'draft a good creative workflow from scratch with the available nodes'. This is partly because vision sucks and partly because they are taught to follow patterns (who woulda thunk) and basically solve the last mile problem from it. Making good workflows is a lot of experiment, guess, check, change. They aren't really OOTB great experimentalists and have poor validation without a human in the loop. That said, I'm sure it'll get better, but yeah, like I can have an AI make a single image pretty reliably without my oversight, but anything image->video is a slog right now lol
1
u/Cautious_Chicken_604 1d ago
A fellow traveler. This matches my experience exactly. Vision being OK-ish, but not good enough to weed out anomalies to autonomously improve workflows, and zero fucking 'sense' plus failing at detecting really obvious failure modes visually. It'll improve man, it will. Until then, we slog. I'd really like to slog 3x faster with a more capable model. Sounds like even you're struggling with it on even better models. That's both good to know and mildly disheartening, but then again... we've come a looooooooong way so far. I'm sure we'll go a long way further.
1
u/tat_tvam_asshole 23h ago
Yeah, I run both GLM5.3-Flash and DS4.1F locally and they just aren't great for creative workflows. I mean, i'm sure I could make it better myself, better deterministic 'indeterministic' types of skills/workflows/rules, but frankly it's just faster and better (right now) if I tell the AI 'make this type of image' then it makes like 6 krea2 images on some theme, I choose a best one, they commit it, then I do the minimax h3 production myself. I'll say that I have had good video gens from AI themselves, but it still is more luck than skill on their part.
Here's my trick to basically how I learned to git gud at Comfyui that I taught my agents. "Go to Lora manager, look up the lora you're going to use, find one of the examples and use that prompt. Make minimal changes to the prompt to reflect the image/video you want to make." I don't allow lora stacking right now, nor do I allow free hand prompt writing as those tend work very poorly. Krea2 gens, like I said, seem to be pretty good at this point, but H3 doesn't really work well because it still has to dissect the video into still frames so it misses alot of poor motion continuity mistakes. It still catches gross fuckups, but because of the render time cost of H3, I just haven't thought of a good way to allow it to catch those on its own.
Don't even get me started on workflows that are even slightly more complicated than that 😂 (reference audio for H3, custom nodes etc) lol
1
u/Cautious_Chicken_604 23h ago
Sounds like we've settled on more or less the same basic workflow / approach. Bootstrap a process to bulk generate starter images from a known decent starting point -> bulk generate starter images with relatively high yield -> manually filter out the BS -> Bulk generate videos from starter images through a handcrafted workflow.
It sucks that it doesn't scale to efficiently doing it for multiple domains.
Lately what I'm thinking of doing is just scraping a tonne of example workflows / prompts of civit and see If I can mine that to build some skills / knowledge base around.
1
u/ambassadortim 23h ago
Yeah I don't feel good spending a day to get a video due to H3 tender times when trying to make longer videos.
1
u/ambassadortim 23h ago
Yeah same boat. I do feed my AI the H3 prompt guides and have it do research in what I'm asking it to create. I run Hermes for these items and that ys fantastic. I have it learn H3 and it gets better. Then I have it review YouTube videos if different fighting styles and it learns how to make better fighting videos. Then even better YouTube videos if specific fight moves or tactics and have it create a motion refmod. That's a huge thing for H3. And I do all that from taking to agents in my phone. I don't use comfyUI web interface at all anymore myself.
1
u/ambassadortim 23h ago
I have similar setup. Sometimes I have a local machine run H3. Then another laptop run Hermes and a cloud API to control H3 via it's API. It makes using H3 very easy. It has sshiti the machine too so it makes and troubleshoots workflows and videos and it's great. Using a model with vision it also analyzes outputs.
I wish I could run local AI but I too can't run both.
1
u/caetydid llama.cpp 1d ago
at least devote a fast nvme ssd to swap between the two in a reasonable amount of time!
1
u/Cautious_Chicken_604 1d ago
I used to do a lot of model training on LTX2.x on the RTX 5060 Ti and that caused a tonne of swap usage. It definitely was not kind to my NVMe. I actually got better performance out of disabling swap because it would sometimes make a decision that it need to swap 20GB of weights into swap and shit would slow to a crawl for like 10 minutes. I realised I had purchased the most dogwater NVMe on the planet to save money. I recently got a proper PCIe Gen5 one to right that wrong. I felt a bit more yolo about this in the past, but after actually seeing the wear on the drive it did cause... I dunno.
1
u/feelspeaceman 1d ago
V100 is the (temporary) savior, just get 1 V100 and everything is solved with SOTA Q38-27B running at 190t/s decode 1000t/s prefill.
Until it's no longer savior because of scalpers, so buy them but try to not be greedy and profit on them like those bad actors.
3
1
u/AdministrativeEmu715 22h ago
I expanded my 16 gigs to 48gb just before this price rally last year. Now running lot of vms in proxmox. Never felt a piece of ram shit can influence my life
1
u/West-Way-All-The-Way 19h ago
Yeah, I am reading the post, I also have 64GB but DDR4 and 16GB in the GPU. I think my problems are not the HW,.I am still struggling to configure the SW. IMHO the models you can currently run are more than good for your home setup, use the rest of the SW suite to make it work. You don't need epic models you need a working setup ( same btw for me ).
1
u/FatheredPuma81 19h ago
I have 48GB and a 4090. Basically the double whammy of barely 4 bit Qwen at usable context and large models render my PC basically unusable even though the models give usable enough tg (25t/s with 4.25bit Flash Next in Unsloth Desktop). Still plenty for every single ComfyUI model I've tried at least.
1
u/Cautious_Chicken_604 18h ago
Try run H3 for 15 second clips at high resolution.
1
u/FatheredPuma81 9h ago
Okay? Worked fine took 40 minutes which is to be expected afaik? What did you want to prove? That the 4090s compute isn't good enough to do it quickly? Already knew that :).
1
u/Cautious_Chicken_604 5h ago
I guess it depends on your launch arguments in ComfyUI but I always find that doing this, especially during the VAE decode stage, the overall VRAM+System RAM usage is basically maxing out my machine. When I first loaded Strata I saw it also basically maxing out my machine, so I figured I wouldn't be able to run two things that simultaneously separately needed all my system RAM. Turns out --fast-disk in ComfyUI and the low ram options in Strata make it work and I can use them together at the same time.
1
u/FatheredPuma81 4h ago
Yea I wasn't at the PC during VAE but the times don't suggest any serious issues. Idk I just launch it with the default bat. The drive the models are saved on is 7GB/s so maybe that's why.
1
1
u/kcksteve 17h ago
I have a worse version of this setup. W6800+rx6800xt+64gb quad channel ddr3. 3.8 27b q6 on the big card assisted by mellum2 on the small card is working well for me. 180 tps for basic operations and 27b at 32tps for heavy thinking is a killer combo.
1
u/mdamour1976 17h ago
I have 128gb RAM but actually choose to run 64gb for XMP and dual channel speeds. 4800 vs 6400 MT/s, and when running with low VRAM it makes a big difference when you've overflowed to system. I may drop it back in to see what I can do with Strata though. I've been quite impressed thus far.
1
u/TopCheddar27 17h ago edited 16h ago
You splitting QFN over the two cards was hurting you more than helping. I have a 4090 and a 5060 and It runs better when I act like the 5060 isn't even there. Other models that stay in VRAM, sure you can split them. But you have to think about where the layers are going and which is your main GPU.
1
u/Cautious_Chicken_604 16h ago
It's all a trade-off at the end of the day. Only way I can run the IQ4_XS is if I split it across the two cards. Otherwise it's too big. So, less speed, but more quality. For me, if I can get 35 t/s as a daily driver I'm happy enough, so I'd kinda prefer that to being able to run IQ3_XXS at 60 t/s tbh. Split across two cards I could get IQ4_XS at 15 t/s, and on a single card only IQ3_XXS at 21 t/s on llama.cpp which was too slow to my meet my 35 t/s target for a daily driver. Actually, I haven't tried running the IQ3_S quant on Strata yet, so I should give that a try, but the recommendation was that was going to be uncomfortably tight.
1
u/Upper_Comparison_908 16h ago
Batching? Do you need it live at the exact same time? Can you treat qfn as an orchestrator perhaps? Or maybe run comfyui on another device if you have one
1
u/JazzlikeLeave5530 16h ago
I have the RAM but I don't have a strong enough card. It's extremely painful to be so close yet so far.
1
u/Cautious_Chicken_604 15h ago
Reminds me of this banger from my youth: https://www.youtube.com/watch?v=36f_hQMwitI
1
u/ColorsOfCosmos 16h ago
Did you try setting up large swap ? Of course it won't help with running them in parallel, but if you run them sequentially you won't have start/stop your apps.
1
u/AvidCyclist250 llama.cpp 14h ago
64gv tam 4080. iq3 s. 80 t/s. kv8. 130k context. i never run comfyui with strata. just have it set up awesome incredible workflows.
1
2
u/sloptimizer 12h ago
Trust me, it's never enough VRAM or compute! That's why you see people dropping large amounts of cash on 8 sparks or 4 apple stuios, etc. Data centers are the extreme case of this condition, but for billionaires. Ultimately, it's the same itch - more compute, more VRAM, more capable AI at your fingertips.
2
u/Cautious_Chicken_604 12h ago
It's a hell of a drug. I'm always impressed when people squeeze more out of what they already have though. Kinda makes you wonder what eventually will be possible to come out entire data centers.
1
u/sloptimizer 10h ago edited 10h ago
Kinda makes you wonder what eventually will be possible to come out entire data centers.
I fear that it will be nothing good for average people like you and I.
1
u/Anixelwhe 12h ago
I brought 128gb of ram and a 5090 last year and while it looks good now, in two years time people will be buying much more powerful hardware for a fraction of what I paid.
1
u/GnosticSon 12h ago
TLDR but I have a 64gb Strix Halo and am very happy with it. Didn't spend that much money and I can easily run Qwen 3.8 27b. For double the money I could have a 128gb machine for very marginal gains. I'd still probably be running Qwen 3.8 27b.
In my opinion 64gb sits at the sweet spot of the curve in terms of money spent to utility provided. Sure if you have minimal budget constraints and a use for larger models go for it. But Qwen 27b more than meets my needs and I didn't have to kill my budget to get it.
1
u/Cautious_Chicken_604 12h ago
Yeah, but Qwen3.8-Flash-Next can be run at excellent speeds now, and it's a better model. I wouldn't call that a modest gain. Especially when Qwen4-Flash comes out, since it'll likely be the exact same architecture, and another improvement in model performance.
Anyway, I figured out my problem, so I can run both QFN and Minimax H3 at the same time on my rig.
1
u/WoodHeartBox 11h ago
I'm in CachyOS and have strata running with a systemd service and had it to write a script that acts as a wrapper. If there isn't enough free ram to run something it shuts strata down and runs the other thing, then when that's done the wrapper script restarts the strata service and keeps going.
The strata service has a setting so if there isn't enough ram or swap strata gets killed by the kernel first before of anything else on the system to prevent system instability, since it uses the most memory when it's running.
The wrapper script needs to eject things from the system completely so it unloads from swap so it doesn't overload the swap. If it doesn't do that then the swap will build up after a while and strata will keep getting OOM killed by the kernel.
It's probably more complicated and detailed than that so if you'd like any extra details just let me know.
1
u/Cautious_Chicken_604 11h ago
That's actually pretty cool. Tbh I kinda wouldn't mind running something like Hermes that way. I never have it setup and running because my compute is always being used for other things, but like if I could only run it when the machine wasn't otherwise busy, that'd work...
1
1
1
u/Global-Elk7271 3h ago
I remember getting a terabyte for 15k. Same memory is more 40k. Hang in there
1
u/chodemunch6969 3h ago
Can you set it up so that it unloads the model after a request or with some TTL? In a similar boat on one of my rigs, and setting low TTLs has given me the ability to have ram/vram free up locally without stuff choking. Not sure how you would do that with strata but i've used this approach w llama-cpp + comfyui locally on my blackwell and it works pretty well.
If what you're asking for is how you can run both concurrently then my guess is as good as yours.
1
u/kartblanch 57m ago
Im running q3 uncensored on 64gb, 5090 and 5070 right now. Getting between 60 tok/s and 95 tok/s
1
u/Forsaken-Ad-519 23h ago
привіт, ти можиш використати форк Strata який буде використовувати одночасно дві твої відеокарти (я потім напишу назву коли доберусь до компа)
1
u/iAccurian 1d ago
Damn, I was hoping I could run the same model using Strata on 48GB system RAM..
1
u/Cautious_Chicken_604 1d ago
I'm pretty sure it'll run on that. I have 16GB used up at the moment idel without running Strata due to other shit that's running. Give it a try.
1
1
u/a4mula 22h ago
640k ought to be enough for anybody.
Different time, different era. But some concepts never die. I doubt it has much to do with constraints and more about how any given system utilizes what they do have.
3
u/Cautious_Chicken_604 22h ago
I think what he meant to say... was $640k worth of GPUs ought to be enough for anybody.
1
u/a4mula 21h ago edited 20h ago
lol, even if it wasn't that was likely what was being pointed out regardless.
It's just funny to me how much things change, while how little things actually change. That was the early 80s and you could see the difference in system thinking even before the gluttony of just throw more compute at it was a thing.
Chinese developers are getting by with a less is more mindset and in the process have pushed the algorithmic side of this equation much harder and faster than their American counterparts, not so much because they wanted to but because they had to.
And that applies just as equally at our scale. You work with the constraints you have. There are people running medium to complex long horizon tasks on edge devices. While others can't get 64GB stacks to do daily driver routines.
edit.
fwiw this wasn't meant to be critical of you, or any that are struggling to find the right scope at their scale. I personally don't even have a working stack. I'm just pointing at generalizations of the space itself. We're all struggling with this to some degree or another regardless of our hardware.
2
u/Cautious_Chicken_604 20h ago
Nah I feel you. I definitely feel like I've been exploring more options for how the squeeze more out of what I have given I'm at my limit of what I can legally acquire right now.
1
u/a4mula 19h ago
Stop drilling up. Start looking for better ways to optimize the stack itself. I have no idea what you're working with directly, but if it's limited to just a few github repos of fun stuff, there's likely a ton of room for your own optimized solutions, better A2A between your systems, better workflows that go beyond just a single agentic harness, better methods of agentic communication and organization, better methods for multi model concurrency or parallelization. There's a lot of component based path optimizations that solve problems that intelligence alone doesn't. The functional side of the operation that keeps the reasoning and intelligence properly aligned.
2
u/Cautious_Chicken_604 18h ago
I do all that already, but when a better model or inference engine comes along it makes sense to get the most out of it and then return back to those other things.
2
u/LankyGuitar6528 14h ago
2
u/a4mula 14h ago
It's funny isn't it, the inversion. We followed the pricing of Moore's Law for a long time. At least through the mid 2010s. Crypto seemed to be the start of the end. AI has taken that and doubled down. We went from proprietary and expensive (IBM) to generalized and affordable (Intel, Cyrix, AMD, Nvidia, LSI, SGI, 3DFX, ATI ect ect ect) and now that the fruits of all of that have been harvested and reconsolidated, we're back to proprietary and expensive. Fuck Huang.
-2
u/zyxciss llama.cpp 19h ago
Hey ChatGPT , Write me a human like post prasing the vibe coded slop "Starata" to get some upvotes.
6
u/Cautious_Chicken_604 18h ago
I don't know what people do with upvotes. Do they stick them up their bum bum? Do they jerk off to them?
On the other hand describing my situation has led to a lot of interesting replies and new information which I was after, so mission accomplished.
Of course, you're probably some sort of anti-bot type of bot yourself.
→ More replies (1)1
3
1
u/kuldan5853 17h ago
It's kinda ironic complaining about "vibe coded slop" on /r/LocalLLaMA when the main reason most people here run their models is as coding assistants..

79
u/Fragrant_Scale6456 1d ago
I was just looking at my amazon history today and saw that I paid $170 for my 64gb of ddr5 6000. I remember thinking at the time that 128gb was not worth the extra $200 lol....