Blog

Improving Qwen3.8-27B-Medium on DeepSWE with GRPO

Liheng L.Michael E.Parth P.Spencer M.
pass@1, base Qwen3.8-27B against the step-15 checkpoint, both run with mini-swe-agent at medium reasoning. Averaged over three attempts per task on DeepSWE and Terminal-Bench 2.1, and six on SWE-Bench-Pro (Hard).

Overview

We improved Qwen3.8-27B’s DeepSWE pass@1 (avg 3) from 27.8% to 39.1% (+11.3 points) using GRPO on 500 tasks from AfterQuery’s SWE agent training dataset. AfterQuery builds all tasks from real software engineering workflows within private, enterprise-grade codebases to avoid evaluation contamination and ensure realism.

Training data

AfterQuery’s SWE agent training dataset spans 10,000+ tasks. Experienced software engineers and leading open-source contributors author them, with review by artisan software engineering experts. We selected 500 of these tasks in a learnable distribution for Qwen3.8-27B. These tasks span 422 repositories across Python, TypeScript, Go, Rust, and JavaScript. We used all 500 for training, with none set aside for validation.

By task type, the 500-task training set includes 402 feature requests (80.4%), 58 bug fixes (11.6%), and 40 enhancements that improve existing functionality, such as performance optimizations or refactors (8.0%).

500tasksPython19739.4%TypeScript12124.2%Go10220.4%Rust6913.8%JavaScript112.2%
  • Python19739.4%
  • TypeScript12124.2%
  • Go10220.4%
  • Rust6913.8%
  • JavaScript112.2%
Training tasks by language.

Training methodology

For each task, we generated eight attempts and used GRPO to update all language-model parameters based on each attempt’s reward relative to the others. We chose medium reasoning to speed up rollouts and reduce the overhead of condensing long conversation histories.

Training diagnostics

Reward

Mean shaped reward

0.40.50.60.70.80.9113579111315Training step0.57010.60320.51130.60540.69150.66560.60470.77380.78290.804100.828110.673120.693130.730140.85615
One point per training update, without smoothing. The y-axis starts at 0.4.

Mean training reward, after penalties for incomplete attempts and response token limits, rose from 0.570 to 0.856 over 15 updates. Each update used a different batch of tasks.

Entropy

Entropy (nats/token)

0.40.50.60.713579111315Training step0.53610.55620.50930.55440.57050.52660.56970.54980.58390.537100.560110.605120.599130.553140.57215
One point per training update, without smoothing. The y-axis starts at 0.4.

Entropy measures how spread out the model’s next-token probabilities are. It ranged from 0.509 to 0.605 nats/token, with no sustained decline over 15 updates.

Evaluation results

We evaluated the base model and step-15 checkpoint with three attempts per task on DeepSWE and Terminal-Bench 2.1, and six on SWE-Bench-Pro (Hard).

Both models used mini-swe-agent at medium reasoning, with the same prompts and generation settings for each benchmark. Qwen reports 42.2% on DeepSWE 1.1 using Claude Code at xhigh reasoning, so its reported score is not directly comparable to our results.

DeepSWE

Average success per attempt

pass@1 (avg 3)

Base27.8%
Trained · step 1539.1%

Tasks solved within three attempts

pass@3

Base51.3%
Trained · step 1568.1%

Terminal-Bench 2.1

Average success per attempt

pass@1 (avg 3)

Base57.30%
Trained · step 1561.42%

Tasks solved within three attempts

pass@3

Base73.03%
Trained · step 1575.28%

SWE-Bench-Pro (Hard)

Average success per attempt

pass@1 (avg 6)

Base61.44%
Trained · step 1564.71%

Tasks solved within six attempts

pass@6

Base88.24%
Trained · step 1588.24%

Get in touch here to access our off-the-shelf SWE agent training datasets or reach out to us directly at research@afterquery.com to see the full post-training experiment and reward design.

AfterQuery is an applied research lab curating data solutions to accelerate foundation model development.

Related articles