LaunchDevelopersJuly 14, 2026

Open-source agent uses RL to train smaller Qwen models autonomously

Agent writes full RL training pipeline (environment, reward, dataset, hyperparameters) and submits to real GPUs. Trained on Qwen3.6-35B-A3B, it produces smaller Qwen models with improved scores. Priced under $1.3k.

2 sources