Agent skill · zechenzhangagi
serving-llms-vllm
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/
Reference it in any coding agent with:
@skills zechenzhangagi/serving-llms-vllm