Canonical weights:Qwen/Qwen3.6-35B-A3B
Use this lane for the staged flow you asked for:
continue pretrain -> instruction SFT -> preference tuning (DPO/GRPO) -> eval + distill.
The in-app runner handles continued pretrain, SFT, and GRPO. DPO and full distillation recipes are generated below for external stacks.