Agent skill · wshobson
preference-optimization
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing bet
Reference it in any coding agent with:
@skills wshobson/preference-optimization