Agent skill · davila7

gptq

Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory

Reference it in any coding agent with:

@skills davila7/gptq

Browse the @skills marketplace