serving-llms-vllm | Skill Performance & Reviews | TopRankSkills

TopRank Skills

Home / Skills / testing security / serving-llms-vllm

serving-llms-vllm

maintained by benchflow-ai

star 620 account_tree 199 verified_user MIT License
bolt View GitHub

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

Key Features

  • Comprehensive skill evaluation and performance tracking
  • Community-driven ratings and reviews
  • Easy integration with Claude Code
  • Regular updates and maintenance

Quick Start

TopRank Skills install benchflow-ai/serving-llms-vllm

chat Comments (0)

chat_bubble_outline

No comments yet. Be the first to share your thoughts!

Skill Details

GitHub Stars 620
GitHub Forks 199
Created Jan 2026
Last Updated 4 months ago
testing security testing security llm ai

Related Skills

ai-sdk

ai-sdk

vercel
star 22.3k
chevron_right
context-engineering-collection
chevron_right
humanizer
chevron_right
notebooklm
chevron_right
feature-dev
chevron_right

Build your own?

Join 12,000+ developers contributing to the Claude ecosystem.