r/LocalLLaMA • u/jjusko20 • 2d ago
New Model Update #3: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch
Last update for those who may be following: https://www.reddit.com/r/LocalLLaMA/comments/1wv5h8x/update_2_post_training_yandexaliceai80ba3b/
I screwed up guys 😂
Turns out my loss curve during my last run was legitimately unhealthy - as some of you, and myself, were concerned about. After evaluating my QLoRA, I found zero'd gradients in all but two layers. Turns out I had a NaN issue related to my custom v100 kernels that I didn't catch - so that run is cooked, I had to restart. I guess two layers training managed to emulate a loss curve I could at least derive a sensible explanation for until I actually got to evaluate.
Thankfully, checked to make sure gradients were applying again, and restarted the run. Once again, it's live streaming at https://figure-bios-expect-cio.trycloudflare.com/
Loss curve looks much healthier this time and is making me feel more confident that this is going to be okay. Stay tuned! I'm gonna release GGUFs and a llama.cpp patch when I have a working version.


4
u/Mysterious_Camp8978 2d ago
lmao that's cool, good luck bro