I Built A Thing
Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware
I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.
With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).
Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:
Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):
- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context
- Output checked against llama.cpp layer by layer
Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):
- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.
- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.
- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.
- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.
Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.
Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting
Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!
if you haven't already, worth pinning each compute tile to its local HBM pseudo-channel. the VU35P's built-in AXI switch will route anywhere, but lateral hops across the stack eat into effective bandwidth badly. also at 75MHz you're running way under the HBM controller clock, so you need several wide ports per unit to even come close to saturating those 400GB/s.
Thanks 75MHz was just the initial goal but trying to close timing upto 200 MHz so the memory bandwidth is used more. Although for VU33P we’re compute (fabric limited). For the 35Ps still trying to figure out how to get the weights across. That Jungle Cat Lite board has no high speed interface to load that I’ve found so far. But first trying to get inter-35P GTY pairs going. They have no clock but the PCB has footprints for an oscillator and fanout buffer
gty refclk has to come off the dedicated mgtrefclk pins and meet the jitter spec, so fabric clocks won't do it, populate that oscillator footprint (156.25 or 161.13 mhz) plus the fanout buffer. for weights, jtag is only a few mb/s, so a qspi or sd-backed load usually wins.
gty refclk has to be a real differential clock into mgtrefclk, you can't drive it from fabric, so populate that oscillator footprint (si5338 or a 156.25mhz sit9501). for weights jtag tops out around a few mb/s, so pushing them into the 35Ps over gty from the vu33p once links are up is way less painful.
Thanks for the info. I ordered these. Have the 0603 resistors in stock. Was gonna order CDCLVD1204 but lucked out and checked the Vcc to be 3.387V
Weird thing though, Pin 2 (IN_SEL) should be grounded for the correct oscillator (IN0) but it's not and did not find any resistor around. So thought the STM micro might be pulling it low, but it doesn't go to any pins. Maybe goes to one of the FPGA pins? I did do a quick swipe across the connector pins and it find any connection (although I did swipe quickly, maybe the continuity detect didn't latch). Anyway, can solder bodge it to GND since it's adjacent
IN_SEL on the CDCLVD1204 has an internal pulldown, so floating already selects IN0, that's why there's no resistor anywhere on the net. tying it to GND is still fine and makes it deterministic. just probe it unpowered, the weak pulldown reads like a near-open on a fast continuity swipe.
Provided FPGA allows super-effective implementation (or rather hardware acceleration) of any math operation, I was waiting for someone to do something like this. Good luck and looking forward for more updates
Nice board. Its almost worth purchasing one to see what else Opus/Fable can do with a board like this. Even a local model (Qwen/Qwen3.8-27B) might be good with VHDL these days?
Hell ya, so cool! I feel like I know enough to know that FPGA inference has potential to be really amazing and offer breakthrough performance, but not nearly enough to do it myself, so hats off to you!
Honestly i am thinking about getting some old phone. 12GB min, more preffered. Wrecked one, just with working motherboard. Then try to see what i can run, but i would assume 7b could run good enough (20tok/s) on that while using minimal power.
IDK still battling this out and trying to find something locally that would make sense.
So far mostly 8gb and ddr4 , but i would aim for ddr5 and sampdragon 8.
I looked at these a while back but thought you needed Vivado Enterprise to program them, is that correct?
Otherwise, how did you generate the bitstream for your design?
LLM ants.... this could be a game changer, like the one years ago when people realized there is no point to dig on GPU cards 😄at 75MHz! imagine pro ants at ghz range... lol. I bow to you.
Sweet! Seed search of weights for my SeedLM models would go much faster with a FPGA. Ive avoided buying one due to the complexity but I think its time to invest. Thanks for the inspiration!
I've found to be not too good as well. But my end plan was to have a mic and a speaker to create a small assistant box, with TTS, STT models. So this could answer things. Also fine tuning it for specific purposes. And as a test for the Qwen3.5 architecture
llm prefill is compute bond and lesser memory bond. decode is memory bond. for fpga to be effective against gpu, it must first solve the memory bond problem. compute is always abundant on modern gpu during decode.
i hope someone can figure this out.
back in the crypto days, people designed scrypt to be memory bond to suppress fpga acceleration. that did not last past 2013. people figured out how to defeat that memory bond.
Can someone extrapolate? Is there a class of ASICS that would cost 50% or less if you somehow manage to harness the power of current of obsolete techj for token gengeration"?
31
u/Redditiskindasilly 12h ago
This is insane work. Really this is the kind of thing that pushes the hobby! Excellent job!