Post
16
Beatrix v3's attention has problems. The model at high byte counts simply cannot recall very well, which has been an expensive lesson but a necessary one.
To remedy this fault, I will complete the Beatrix 2s softmax control variant until stable. I will train the curriculum stages with QK normalization similar to Gemma, giving us a proper full model MHA softmax control. The answer is likely faulty bf16, needs fp32 and QK norms on the attention, softmax attention in intervals for the primary model.
After stabilizing the v2 softmax variant, I will train the v3 softmax variant as well.
These are required. There are many open questions about stability at depth and stability at context window capacity, all of which can be answered with the high accuracy recall at depth and context window size.
An expensive lesson, but a necessary lesson for progress.
Even with the faults, Beatrix V3's mechanisms have proven to be invaluable in their own rite. Multiple experiments have yielded since, including the tokenizer processing system - giving Beatrix the ability to directly speak as another model's tokenizer language. It's not perfect yet, but the process is likely reusable for other similar architected models.
The quiet loss has been extensively successful. The structure of each arm trained one after the other shows we can keep each arm nice and quiet while being fully expert in their own field. Activation happens when the byte pattern says so, and the arms stay nice and quiet when the gated pattern doesn't say so.
Modularized arms are successful. This is a process that can likely be utilized on other models as well, but the process hasn't been fully completed yet. The model Beatrix was a good catalyst for this because of the byte format of the model and the anchored behavior allows this to be trained more quickly and stable.
In any case, the 2s-control is upcoming and will begin today. I'll try to get it cleaned up by the end of the week to prepare for the 3s control.
To remedy this fault, I will complete the Beatrix 2s softmax control variant until stable. I will train the curriculum stages with QK normalization similar to Gemma, giving us a proper full model MHA softmax control. The answer is likely faulty bf16, needs fp32 and QK norms on the attention, softmax attention in intervals for the primary model.
After stabilizing the v2 softmax variant, I will train the v3 softmax variant as well.
These are required. There are many open questions about stability at depth and stability at context window capacity, all of which can be answered with the high accuracy recall at depth and context window size.
An expensive lesson, but a necessary lesson for progress.
Even with the faults, Beatrix V3's mechanisms have proven to be invaluable in their own rite. Multiple experiments have yielded since, including the tokenizer processing system - giving Beatrix the ability to directly speak as another model's tokenizer language. It's not perfect yet, but the process is likely reusable for other similar architected models.
The quiet loss has been extensively successful. The structure of each arm trained one after the other shows we can keep each arm nice and quiet while being fully expert in their own field. Activation happens when the byte pattern says so, and the arms stay nice and quiet when the gated pattern doesn't say so.
Modularized arms are successful. This is a process that can likely be utilized on other models as well, but the process hasn't been fully completed yet. The model Beatrix was a good catalyst for this because of the byte format of the model and the anchored behavior allows this to be trained more quickly and stable.
In any case, the 2s-control is upcoming and will begin today. I'll try to get it cleaned up by the end of the week to prepare for the 3s control.