Agent skill · aradotso
turboquant-pytorch
PyTorch implementation of TurboQuant for LLM KV cache compression using two-stage vector quantization (random rotation + Lloyd-Max + QJL residual correctio
Reference it in any coding agent with:
@skills aradotso/turboquant-pytorch