vllm · 46 topics
2026 · Sep
alibaba shrinks qwen3 32b to fit on a 24gb consumer gpu
09-27🔥🔥edge0 audio8 asr infinite
09-12🔥🔥vericache lossless kv cache
09-11🔥🔥cohere labs north small translate beats deepl and google translate across 50 languages
09-10🔥🔥deploying qwen3 8 2 4t a95b on amazon sagemaker hyperpod with vllm
2026 · Aug
qwen3 8 flash next fp8
08-23🔥🔥qwen3 tts speed cost frontier
08-23🔥🔥ornith 1 5 35b a3b gguf
08-23🔥🔥z lab qwen3 8 27b dflash2
08-19🔥🔥🔥qwen3 8 2 4t a95b fp8
08-18🔥🔥🔥unsloth qwen3 8 27b nvfp4
08-15🔥🔥🔥qwen qwen3 8 27b
2026 · Jul
2026 · Jun
2026 · May
2026 · Sep
build real time voice applications with vllm omni on sagemaker ai
09-21🔥🔥ukisai swift qwen38 27b gguf
09-19🔥🔥aiperf llm benchmarking
09-18🔥🔥vllm boosts kimi k3 throughput
09-12🔥🔥dealignai glm 53 cybersecurity fp8
09-12🔥🔥nvidia qwen3 flash next nvfp4
09-11🔥🔥reduce llm latency with prefix aware routing on amazon sagemaker inference
09-09🔥🔥coheres open source megakernel beats vllm
09-08🔥🔥speculative decoding in vllm on amd gpus
09-06🔥🔥k2 horizon local verify
2026 · Aug
vllm v0280 release
08-28🔥🔥vllm v0 28 0 sparse attention speculative decoding
08-26🔥🔥llms could control their host machines by exploiting inference engines
08-26🔥🔥nvidia dynamo shadow engine recovery
08-24🔥🔥why your local llm feels dumber than it is
08-16🔥🔥vllm dspark deepseek v4
08-04🔥🔥🔥the inference engineering masterclass
08-04🔥🔥airllm 70b inference single 4gb gpu
08-01🔥🔥poolside laguna s 2 1 rate limits
08-01🔥🔥mac laguna ollama llama cpp vllm
2026 · Jul
deploying kimi k3 on aws
07-27🔥🔥vllm v0 26 0 inkling 1t
07-27🔥🔥poolside laguna s 21 nvfp4
07-09🔥🔥native speed vllm transformers backend
07-06🔥🔥internscience agents a1