serving-llms-vllm | Skill Performance & Reviews | TopRankSkills

TopRank Skills

Home / Skills / testing security / serving-llms-vllm

serving-llms-vllm

maintained by benchflow-ai

star 620 account_tree 199 verified_user MIT License
bolt View GitHub

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

Key Features

  • Comprehensive skill evaluation and performance tracking
  • Community-driven ratings and reviews
  • Easy integration with Claude Code
  • Regular updates and maintenance

Quick Start

TopRank Skills install benchflow-ai/serving-llms-vllm

chat Comments (0)

chat_bubble_outline

No comments yet. Be the first to share your thoughts!

Skill Details

GitHub Stars 620
GitHub Forks 199
Created Jan 2026
Last Updated 4个月前
testing security testing security llm ai

Related Skills

ai-sdk

ai-sdk

vercel
star 22.3k
chevron_right
context-engineering-collection
chevron_right
humanizer
chevron_right
notebooklm
chevron_right
feature-dev
chevron_right

Build your own?

Join 12,000+ developers contributing to the Claude ecosystem.