1 / 2
I'm training the model using a self distill, so output loss is KL calculated against the base model token outputs. Model is Qwen3-0.6B, frozen all that gets trained is the encoder. Training on wildchat, running for about 30mins.
1 likes 1 replies
?