Agent skill · zechenzhangagi

simpo-training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference m

Reference it in any coding agent with:

@skills zechenzhangagi/simpo-training

Browse the @skills marketplace