resource · seed

vllm-project / vllm

A high-throughput and memory-efficient inference and serving engine for LLMs.

github.com/vllm-project/vllm · Python

Everything I ship today is API-first. This is the bookmark for the day self-hosting an open model becomes the cheaper answer.

#inference #self-hosting

See this note on the whiteboard →