Scaling the future of RL through reward models
I'm 19. From Malaysia and Nepal, and I grew up in the Bay Area. I swim, and I cook biryani and Philly cheesesteaks. At 15 I built a forex trading algorithm that made mid-five figures.
Luelnow
Data research.
Dropped out and started selling RL and coding data to labs.
Papers I find interesting.
On a 49-task slice of SWE-bench Verified, 28.5% of suites accept a Docker-checked wrong patch. Across 134 submissions those tasks score about 14 points higher inside the same difficulty band. The repair detail is better than the headline: a gold-sanity gate rejects 62% of the LLM-written replacement tests because they fail the gold patch, which a judge reading only the test text lets through.
The bottleneck was executable repos, not more issue text. They break existing tests on purpose across 128 codebases and land 50k tasks without a multi-terabyte image dump. A 32B trained on 5k of them reaches 40.2% pass@1 on SWE-bench Verified. Synthetic bugs only count if a real suite still distinguishes them.
SWEGen builds tasks by generating tests and back-translating commits, so the issue does not have to be human-written. The split I keep coming back to: execution-based verifiers cannot tell two green patches apart, execution-free verifiers get pulled toward the trace instead of the diff. Each saturates near 42%. The hybrid reaches 51% on SWE-bench Verified.
2,438 real Python tasks with a runtime, used twice: once to train the agent, once to train an outcome verifier on its own rollouts. The verifier is the next-token probability of yes against no. Best-of-k with that score, not a fatter scaffold, is what moved the open result to 32% on SWE-bench Verified.
The harness is a file tree, so an edit can be reverted, and every edit has to predict the next round's task outcomes. Ten rounds take Terminal-Bench 2 from 69.7% to 77.0%. The ablation is the part worth stealing: tools, middleware, and memory move the number. The system prompt does not.
GRPO's importance ratio is per token, which fights MoE routing and the numerical gap between the sampler and the trainer. GSPO clips a sequence likelihood instead. Because of that they can skip recomputing old-policy logprobs in the training engine, which is the bill that shows up once rollouts are multi-turn and the two engines are split.
The length bug is the 1/|o| in the loss. A long wrong trace is under-penalized, so incorrect answers grow and get called an aha moment. Drop that term and the std in the group advantage. What is left is a Monte Carlo return with an unbiased baseline, which is what people thought they were running.
Posts I find interesting.
A harness, cut down, is a loop, a tool surface, a context policy, and a stop. mini-swe-agent is the example that matters for training: the only tool is bash, and the history is a flat append. If the scaffold is fatter than that, the policy can memorize the tool API and fall over the day you swap sandboxes.
Read it for the ablation, not the tour. On their AIME setup vanilla GRPO sits near 30 and the four DAPO changes get a 32B base to 50. Token-level loss is the small accuracy bump and the one that stops length and entropy from thrashing. Clip-higher, dynamic sampling, and the soft overlong penalty do the rest.
This one is the outer harness, the guides and sensors a person adds around an agent that already has a scaffold. The useful worry is coverage. A sensor that never fires might mean the agent is clean. It might also mean the sensor cannot see the failure. Same ambiguity as a suite that never rejects a bad patch.