SkillGym improves Qwen3.5-35B-A3B performance by 19 points on Terminal-Bench 2.1
Summary
A new study from the Shanghai AI Laboratory introduces SkillGym, a framework designed to improve the performance of large language models by internalizing human-written agent skills as reusable capabilities. This approach shifts from relying on prompt-dependent skills to creating executable tasks verified by code checkers, which has reportedly boosted performance metrics significantly. For instance, fine-tuning the Qwen3.5-35B-A3B model using SkillGym led to a 19.10-point improvement on the Terminal-Bench 2.1 evaluation. This research taps into a growing trend in LLM development that prioritizes verified training data and empirical skill dependence to enhance problem-solving reliability.
Analysis
SkillGym: SkillGym is a framework introduced in a recent arXiv paper that converts human-written agent skills into executable sandboxed tasks with automated code-based verification. It collects successful trajectories from multiple models to support supervised fine-tuning and reinforcement learning on verified workflows. The approach allows LLMs to internalize procedural skills rather than depending on retrieval and instruction-following at inference time. Claude Code: Claude Code is an agent harness and evaluation environment built around Anthropic's Claude models for running and testing LLM agents on coding and terminal tasks. It served as the primary setup for assessing SkillGym-trained models on benchmarks measuring real-world problem-solving performance. Fine-tuning under this harness demonstrated reduced reliance on external skill files. Qwen3.5-35B-A3B: Qwen3.5-35B-A3B is a large language model variant in Alibaba's Qwen series used for agent capabilities research. The SkillGym paper applied supervised fine-tuning on verified skill trajectories to this model, resulting in improved autonomous performance on terminal and skills benchmarks. It outperformed base versions both with and without prompt-provided skills after training. Training Frameworks: New methods like skill-to-task pipelines with outcome verification enable creation of high-quality trajectory data for fine-tuning agent models across diverse categories. LLM Agent Development: Recent research emphasizes shifting from prompt-dependent skill usage to internalized reusable capabilities in large language models for more reliable problem solving.
Categories
machine_learningaiai_agentstech