Agent skill · wshobson

preference-optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing bet

Reference it in any coding agent with:

@skills wshobson/preference-optimization

Browse the @skills marketplace