← Back to live feed

Thursday, Oct 1, 2026

1
LFM2.5-2.6B Model Hits 54% Solve Rate With New Multi Harness RL
topics πŸ€– AI tags AIAI ModelsAI Research keywords Adithya S. K.Ben Burtenshaw

A new capture proxy for OpenEnv permits the direct training of AI models via reinforcement learning within tool-using environments like Claude Code, Codex, and OpenCode. Developed by Adithya S.K. and Ben Burtenshaw, the tool sits between the harness and the model to record exact token IDs and logprobs of every call, allowing reinforcement learning on any task set without requiring modifications to the harness or training code.

The LFM2.5-2.6B model demonstrated the system's utility, raising its average solve rate across four harnesses from 42% to 54% while reducing tool calls by 31%. In specific tests within Claude Code, the model's solve rate increased from 33% to 49%. The workflow utilizes Harbor for task sandboxing and the TRL library for asynchronous GRPO training.

Image via @ben_burtenshaw on X
See all 14 tweets β†’