Agent skill · aradotso

turboquant-pytorch

PyTorch implementation of TurboQuant for LLM KV cache compression using two-stage vector quantization (random rotation + Lloyd-Max + QJL residual correctio

Reference it in any coding agent with:

@skills aradotso/turboquant-pytorch

Browse the @skills marketplace