Skip to content
Back to the index

vLLM

vllm.aiInfrastructure

PulseGate's liveness check found it on 12 Sep 2026; it has been in the index since 30 Aug 2026. How this is checked

vLLM is an open-source engine for high-throughput, memory-efficient inference and serving of large language models. It supports multiple hardware backends, continuous batching, PagedAttention, and an OpenAI-compatible API for developers and ML teams.

Inferred · not functionally tested

Open SourceApache-2.0WebCLISelf-hostedAPI
vLLM preview
Visit vllm.ai
90.4kstars
21.4kforks
10features
2023since

Overview

6 features

Purpose: Serving large language models with high throughput and efficient memory use on available hardware.

Inferred · not functionally tested

Audience: AI infrastructure engineers and developers

Inferred · not functionally tested

Functions: Unknown

Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)

Recorded constraints: pricing: open_source · license: Apache-2.0 · platforms: CLI, WEB · deployment: browser, cli, self_hosted, api_only

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: vllm.ai. These links do not verify the individual claims.

vLLM sits in PulseGate's Inference & model serving category. Inferred · not functionally tested: It focuses on serving large language models with high throughput and efficient memory use on available hardware. Inferred · not functionally tested: vLLM is an open-source project aimed at AI infrastructure engineers and developers. Basis unknown · not verified: The project is open source (Apache-2.0). Basis unknown · not verified: It runs on the web, the command line, and API, and it can be self-hosted.

It is developed by vLLM Community, and it first shipped in 2023. Development happens publicly on GitHub with 90.4k stars and 3.4k commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are OpenAI-compatible API, pagedAttention, and continuous batching. Inferred · not functionally tested: Catalogued interfaces include a public API.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • OpenAI-compatible API
  • PagedAttention
  • Continuous batching
  • Multi-hardware support
  • Open-source model support
  • High-throughput serving

Topics: Inferred · not functionally tested

Tags
llm-servinginference-enginepagedattentioncontinuous-batchingopenai-compatible-api
AI capabilities
Text
Inference: Local

JSON profile · Text profile · Access guide

Built with & integrations

Framework
Next.js
Hosting
Vercel
AI providers
local_ossmeta_llama
Written with
Claude CodeCodexGemini CLI
Connectors
API
Runs on
BrowserCLISelf-hostedAPI-only
Written with — evidence
Claude Code
commit b87339888d29 · since Sep 2026
Codex
commit 988d9b6777d0 · since Sep 2026
Gemini CLI
.gemini/
Detected from
Next.js
x-nextjs-prerender header · /_next/static/ in the HTML · __next_f in the HTML
Vercel
x-vercel-id header · x-vercel-cache header
local_oss
bvllm in the HTML
meta_llama
llama- in the HTML

Trust & compliance

License
Apache-2.0
Public signals
HTTPSOpen SourceFree tierActive maintenance

Indexing history

3

What PulseGate has recorded for this listing

  1. Indexed12 Sep · 03:25 UTC
    Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin seen via Hacker News firehose (Algolia)
    Source: Hacker News firehose (Algolia) · Open
  2. Indexed7 Sep · 22:46 UTC
    Speculative Decoding in vLLM on AMD GPUs seen via Hacker News firehose (Algolia)
    Source: Hacker News firehose (Algolia) · Open
  3. Indexed30 Aug · 00:09 UTC
    Efficient Decode Context Parallelism with vLLM for Long Context Workloads seen via Hacker News firehose (Algolia)
    Source: Hacker News firehose (Algolia) · Open

Frequently asked questions about vLLM

What does vLLM do?
Inferred · not functionally tested: VLLM focuses on serving large language models with high throughput and efficient memory use on available hardware. It is catalogued under Inference & model serving on PulseGate.
Who should use vLLM?
Inferred · not functionally tested: vLLM is an open-source project built for AI infrastructure engineers and developers.
Does vLLM have a free plan?
Basis unknown · not verified: Yes — vLLM is open source under the Apache-2.0 license and free to use.
What platforms does vLLM run on?
Basis unknown · not verified: vLLM runs on the web, the command line, and API. It can also be self-hosted.
Is vLLM still maintained?
PulseGate's liveness check found it on 12 Sep 2026. Its GitHub repository shows 3.4k commits in the last 90 days.
What projects are similar to vLLM?
Similar projects tracked by PulseGate include llm-d, vLLM Optimizer, and Mesh LLM.llm-dvLLM OptimizerMesh LLM
Who makes vLLM?
vLLM is developed by vLLM Community.
When did vLLM launch?
vLLM first shipped in 2023.

Similar projects

Closest matches by what these projects do