r/LocalLLaMA • • 13h ago

I Built A Thing Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware

I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.

With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).

Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:

Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):

- Prefill: ~6 tok/s (256-token prompt), ~5.5 tok/s (2.3k-token prompt)

- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context

- Output checked against llama.cpp layer by layer

Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):

- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.

- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.

- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.

- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.

Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.

Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting

Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!

Repo here (MIT): https://github.com/Nero7991/llm.vhdl

271 Upvotes

34 comments sorted by

31

u/Redditiskindasilly 12h ago

This is insane work. Really this is the kind of thing that pushes the hobby! Excellent job!

16

u/Open-Adhesiveness-86 11h ago

if you haven't already, worth pinning each compute tile to its local HBM pseudo-channel. the VU35P's built-in AXI switch will route anywhere, but lateral hops across the stack eat into effective bandwidth badly. also at 75MHz you're running way under the HBM controller clock, so you need several wide ports per unit to even come close to saturating those 400GB/s.

6

u/I_am_purrfect 10h ago

Thanks 75MHz was just the initial goal but trying to close timing upto 200 MHz so the memory bandwidth is used more. Although for VU33P we’re compute (fabric limited). For the 35Ps still trying to figure out how to get the weights across. That Jungle Cat Lite board has no high speed interface to load that I’ve found so far. But first trying to get inter-35P GTY pairs going. They have no clock but the PCB has footprints for an oscillator and fanout buffer

3

u/Open-Adhesiveness-86 10h ago

gty refclk has to come off the dedicated mgtrefclk pins and meet the jitter spec, so fabric clocks won't do it, populate that oscillator footprint (156.25 or 161.13 mhz) plus the fanout buffer. for weights, jtag is only a few mb/s, so a qspi or sd-backed load usually wins.

3

u/Open-Adhesiveness-86 10h ago

gty refclk has to be a real differential clock into mgtrefclk, you can't drive it from fabric, so populate that oscillator footprint (si5338 or a 156.25mhz sit9501). for weights jtag tops out around a few mb/s, so pushing them into the 35Ps over gty from the vu33p once links are up is way less painful.

2

u/I_am_purrfect 6h ago

Thanks for the info. I ordered these. Have the 0603 resistors in stock. Was gonna order CDCLVD1204 but lucked out and checked the Vcc to be 3.387V

Weird thing though, Pin 2 (IN_SEL) should be grounded for the correct oscillator (IN0) but it's not and did not find any resistor around. So thought the STM micro might be pulling it low, but it doesn't go to any pins. Maybe goes to one of the FPGA pins? I did do a quick swipe across the connector pins and it find any connection (although I did swipe quickly, maybe the continuity detect didn't latch). Anyway, can solder bodge it to GND since it's adjacent

2

u/Open-Adhesiveness-86 6h ago

IN_SEL on the CDCLVD1204 has an internal pulldown, so floating already selects IN0, that's why there's no resistor anywhere on the net. tying it to GND is still fine and makes it deterministic. just probe it unpowered, the weak pulldown reads like a near-open on a fast continuity swipe.

1

u/I_am_purrfect 6h ago

Floating IN_SEL disables the output

1

u/I_am_purrfect 6h ago

It's a divider, not just a pulldown

7

u/davernow 12h ago

Very cool

7

u/Shoddy-Tutor9563 8h ago

Provided FPGA allows super-effective implementation (or rather hardware acceleration) of any math operation, I was waiting for someone to do something like this. Good luck and looking forward for more updates

15

u/geldonyetich 12h ago

I have full respect for jank hardware solutions but egads carpet is where static electricity lives, you're a braver man than I.

5

u/lolzinventor 12h ago

Nice board. Its almost worth purchasing one to see what else Opus/Fable can do with a board like this. Even a local model (Qwen/Qwen3.8-27B) might be good with VHDL these days?

3

u/TripleSecretSquirrel 12h ago

Hell ya, so cool! I feel like I know enough to know that FPGA inference has potential to be really amazing and offer breakthrough performance, but not nearly enough to do it myself, so hats off to you!

3

u/stormy1one 12h ago

This kind of project is what this sub exists for. Following!

3

u/Primary_Minimum5033 11h ago

Would be awesome for small models that fit completely in memory, but I guess compute power becomes the issue

3

u/I_am_purrfect 10h ago

Yes! I should try some ~4B parameter models

3

u/GalacticDistances 10h ago

Have you considered using it as a separate MTP drafter for the main model?

2

u/I_am_purrfect 10h ago

I might if I get the XCVU35P up with the current architecture, they have more fabric available to possibly add MTP

3

u/bearishmarket 10h ago

This is beautiful.

2

u/GruuMasterofMinions 11h ago

Honestly i am thinking about getting some old phone. 12GB min, more preffered. Wrecked one, just with working motherboard. Then try to see what i can run, but i would assume 7b could run good enough (20tok/s) on that while using minimal power.
IDK still battling this out and trying to find something locally that would make sense.
So far mostly 8gb and ddr4 , but i would aim for ddr5 and sampdragon 8.

2

u/Ecstatic_Airport_837 9h ago

I looked at these a while back but thought you needed Vivado Enterprise to program them, is that correct? Otherwise, how did you generate the bitstream for your design?

2

u/Illustrious_Dish_999 9h ago

LLM ants.... this could be a game changer, like the one years ago when people realized there is no point to dig on GPU cards 😄at 75MHz! imagine pro ants at ghz range... lol. I bow to you.

2

u/txgsync 8h ago

Sweet! Seed search of weights for my SeedLM models would go much faster with a FPGA. Ive avoided buying one due to the complexity but I think its time to invest. Thanks for the inspiration!

2

u/Kahvana 7h ago

Really cool, thanks for sharing!

1

u/No-Craft-7979 7h ago

I have never gotten any rational sense out of 9B… What do you use 9B for.

2

u/I_am_purrfect 6h ago

I've found to be not too good as well. But my end plan was to have a mic and a speaker to create a small assistant box, with TTS, STT models. So this could answer things. Also fine tuning it for specific purposes. And as a test for the Qwen3.5 architecture

1

u/Elouakili_Flexy 2h ago

An LLM wrote the RTL, then another LLM estimated what it would do on N3, and that imaginary chip hits 294 tok/s. Tapeout can't come soon enough.

1

u/Puzzleheaded_Base302 1h ago

llm prefill is compute bond and lesser memory bond. decode is memory bond. for fpga to be effective against gpu, it must first solve the memory bond problem. compute is always abundant on modern gpu during decode.

i hope someone can figure this out.

back in the crypto days, people designed scrypt to be memory bond to suppress fpga acceleration. that did not last past 2013. people figured out how to defeat that memory bond.

-2

u/Technical_Ad_6106 10h ago

no speed, no efficiency... wouldnt take it for free. its scrap metal.

0

u/MotherPotential 6h ago

Can someone extrapolate? Is there a class of ASICS that would cost 50% or less if you somehow manage to harness the power of current of obsolete techj for token gengeration"?

0

u/lookingyou11 5h ago

Man excellent work! I’m working on same proyect with same card! Let’s start work together!