Qwen 3.6 Training

Idle
Ready to train. Configure and start below.
0%
0
Step
-
Loss
-
Learning Rate
0
Epoch
-
ETA
-
Grad Norm
Training
Charts
Hardware
Models
Export
Inference
Log

Qwen 3.6 Pipeline

Canonical weights: Qwen/Qwen3.6-35B-A3B
Use this lane for the staged flow you asked for: continue pretrain -> instruction SFT -> preference tuning (DPO/GRPO) -> eval + distill. The in-app runner handles continued pretrain, SFT, and GRPO. DPO and full distillation recipes are generated below for external stacks.
Repo: QwenLM/Qwen3.6 HF: Qwen/Qwen3.6-35B-A3B Default lane: AGI-POC + Unsloth

Auto-Create Dataset

Convert source material into a starter SFT dataset without leaving the page.

Model Configuration

Hyperparameters

Dataset

HuggingFace
Local File
Upload
Created
Data Source

LoRA Settings


Reinforcement Learning (GRPO)


Training Loss

Learning Rate

Gradient Norm

Eval Loss

Training Loss

Learning Rate

Gradient Norm

Eval Loss

Loss (MA-20)

GPU Utilization

Loading hardware info...

Trained Models

IDBaseTypeDate
Loading...

Export Model


HuggingFace Hub

Export Reference

FormatUse Case
GGUF Q4_K_Mllama.cpp, Ollama, LM Studio
GGUF Q8_0High quality CPU inference
Merged 16-bitFull precision safetensors
LoRA onlyAdapter weights (smallest)
UD-Q5_K_XLUnsloth Dynamic (best ratio)

Test Model

Output

Run inference to see output...

Training Log

...