
Improving Qwen3.8-27B-Medium on DeepSWE with GRPO
Liheng L.
Michael E.
Parth P.
Spencer M.- Base
- Gain from training
Overview
We improved Qwen3.8-27B’s DeepSWE pass@1 (avg 3) from 27.8% to 39.1% (+11.3 points) using GRPO on 500 tasks from AfterQuery’s SWE agent training dataset. AfterQuery builds all tasks from real software engineering workflows within private, enterprise-grade codebases to avoid evaluation contamination and ensure realism.
Training data
AfterQuery’s SWE agent training dataset spans 10,000+ tasks. Experienced software engineers and leading open-source contributors author them, with review by artisan software engineering experts. We selected 500 of these tasks in a learnable distribution for Qwen3.8-27B. These tasks span 422 repositories across Python, TypeScript, Go, Rust, and JavaScript. We used all 500 for training, with none set aside for validation.
By task type, the 500-task training set includes 402 feature requests (80.4%), 58 bug fixes (11.6%), and 40 enhancements that improve existing functionality, such as performance optimizations or refactors (8.0%).
- Python19739.4%
- TypeScript12124.2%
- Go10220.4%
- Rust6913.8%
- JavaScript112.2%
Training methodology
For each task, we generated eight attempts and used GRPO to update all language-model parameters based on each attempt’s reward relative to the others. We chose medium reasoning to speed up rollouts and reduce the overhead of condensing long conversation histories.
Training diagnostics
Reward
Mean shaped reward
Mean training reward, after penalties for incomplete attempts and response token limits, rose from 0.570 to 0.856 over 15 updates. Each update used a different batch of tasks.
Entropy
Entropy (nats/token)
Entropy measures how spread out the model’s next-token probabilities are. It ranged from 0.509 to 0.605 nats/token, with no sustained decline over 15 updates.
Evaluation results
We evaluated the base model and step-15 checkpoint with three attempts per task on DeepSWE and Terminal-Bench 2.1, and six on SWE-Bench-Pro (Hard).
Both models used mini-swe-agent at medium reasoning, with the same prompts and generation settings for each benchmark. Qwen reports 42.2% on DeepSWE 1.1 using Claude Code at xhigh reasoning, so its reported score is not directly comparable to our results.
DeepSWE
Terminal-Bench 2.1
SWE-Bench-Pro (Hard)
Get in touch here to access our off-the-shelf SWE agent training datasets or reach out to us directly at research@afterquery.com to see the full post-training experiment and reward design.
AfterQuery is an applied research lab curating data solutions to accelerate foundation model development.



