resource · seed
vllm-project / vllm
A high-throughput and memory-efficient inference and serving engine for LLMs.
github.com/vllm-project/vllm · Python
Everything I ship today is API-first. This is the bookmark for the day self-hosting an open model becomes the cheaper answer.