GRPO Fine-Tuning Improves 350M Model on IFStruct Structured Output Benchmark
A publicly shared recipe applies Group Relative Policy Optimization via TRL to LiquidAI's LFM2.5-350M, raising its IFStruct score from 22.6% to 29.7% after 100 steps on roughly…

- Custom reward functions targeting JSON format, field count and schema validation drive measurable gains on the hybrid-attention 350M model.
- Targeted prompt augmentation for fenced code blocks and bare-list outputs produces balanced gains, especially in JSON and list compliance.
- The full pipeline, from LoRA training to llama.cpp serving, remains reproducible on consumer hardware without large datasets.
Baseline Evaluation
The unmodified LFM2.5-350M was served locally via llama.cpp in BF16 precision and evaluated on the full 2000-sample IFStruct test set.
It achieved an overall pass rate of 22.6 percent, with JSON at 18.0 percent and YAML at 27.2 percent. Wrapper-key structures scored higher than bare lists.
Frequent errors involved missing required fields, wrong item counts and type mismatches. This result is close to the 21.1 percent reported in the original IFStruct blog.
Training Procedure
Approximately 500 examples were drawn from the Nemotron structured-output dataset. Prompts were augmented so that 40 percent included explicit fenced-code-block instructions and 20 percent targeted top-level array outputs.
A LoRA adapter was attached to LiquidAI/LFM2.5-350M, focusing on the model's hybrid attention and convolution modules and training roughly 6 million parameters.
Three reward functions scored completions for parseability and requested format, exact top-level field count and full schema compliance. These were weighted 1.0, 0.5 and 2.0 respectively. GRPO training ran for 100 steps with eight generations per prompt group.
Post-Training Performance
After merging the LoRA weights, converting to BF16 GGUF and serving with the same llama.cpp configuration, the model reached 29.7 percent overall on IFStruct.
JSON accuracy rose 13.9 points to 31.9 percent while YAML stayed nearly unchanged. Bare-list compliance improved by 13.1 points, aligning it with wrapper-key results.
The tuned model approaches the 33.15 percent reported for Qwen3.5-2B, with gains concentrated in the areas emphasized by the reward signals.
Reproducibility and Scope
The complete notebook, conversion scripts and evaluation harness are available on GitHub, confirming that meaningful structured-output gains are achievable with modest resources.
The work demonstrates the value of task-specific reinforcement learning on smaller open models but remains limited to text-based JSON and YAML schema adherence.
It does not evaluate tool calling, multimodal inputs or downstream application performance.


