Agent skill · davila7

simpo-training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference m

Reference it in any coding agent with:

@skills davila7/simpo-training

Browse the @skills marketplace