Agent skill · davila7

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/

Reference it in any coding agent with:

@skills davila7/serving-llms-vllm

Browse the @skills marketplace