r/LocalLLaMA • • 2d ago

New Model Update #3: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Last update for those who may be following: https://www.reddit.com/r/LocalLLaMA/comments/1wv5h8x/update_2_post_training_yandexaliceai80ba3b/

I screwed up guys 😂

Turns out my loss curve during my last run was legitimately unhealthy - as some of you, and myself, were concerned about. After evaluating my QLoRA, I found zero'd gradients in all but two layers. Turns out I had a NaN issue related to my custom v100 kernels that I didn't catch - so that run is cooked, I had to restart. I guess two layers training managed to emulate a loss curve I could at least derive a sensible explanation for until I actually got to evaluate.

Thankfully, checked to make sure gradients were applying again, and restarted the run. Once again, it's live streaming at https://figure-bios-expect-cio.trycloudflare.com/

Loss curve looks much healthier this time and is making me feel more confident that this is going to be okay. Stay tuned! I'm gonna release GGUFs and a llama.cpp patch when I have a working version.

My first epoch loss curve from this run
my first epoch loss curve on the failed first run (note the differences in scale even if the pattern looks similar)
27 Upvotes

3 comments sorted by

4

u/Mysterious_Camp8978 2d ago

lmao that's cool, good luck bro

3

u/jjusko20 2d ago

thanks boss