Excited to be co-hosting a happy hour with DigitalOcean and NVIDIA on Aug 26 as the vLLM Conference comes to a close 🎉 https://lnkd.in/p/g8t3Hr6F
Inferact
Software Development
San Francisco, CA 3,991 followers
Building the future of inference
About us
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
- Website
-
https://inferact.ai
External link for Inferact
- Industry
- Software Development
- Company size
- 11-50 employees
- Headquarters
- San Francisco, CA
- Type
- Privately Held
- Founded
- 2025
Locations
-
Primary
Get directions
San Francisco, CA, US
Employees at Inferact
Updates
-
Proud to be co-hosting this one alongside the vLLM community and the NVIDIA Dynamo team 🙌 If you're in SF during vLLM conference week, come talk inference optimization, distributed serving, and everything that breaks at scale. We'd love to meet you there 🙏 Details and RSVP below 👇 https://lnkd.in/p/gC3WBUU8
We're excited to co-host the vLLM x Dynamo Meetup with the NVIDIA team during vLLM conference week in San Francisco! 🚀 Join us for an evening of tech talks and conversation on what it takes to serve LLMs efficiently at scale — from inference optimization and distributed serving to the practical challenges of running these systems in production. 📅 Monday, August 24 · 6:00–9:00 PM PT On the agenda: • 6:30 PM — vLLM tech talk • 6:45 PM — NVIDIA Dynamo tech talk • 7:00 PM — Food, drinks, and time to connect with engineers, researchers, and builders across the inference community Space is limited and registration is subject to approval — reserve your spot early: https://luma.com/r8o604o0 #LLM #Inference #vLLM #NVIDIA #MachineLearning #AIInfrastructure
-
We're proud to be hosting the first vLLM Conference at Ray Summit later this month 🌉 If you're working on inference, come spend three days with the team and the broader vLLM community. See you in SF, Aug 24–26.
The vLLM Conference is coming up in 3 weeks! 🎉 Come learn about the current state and future of AI inference, Aug 24–26 in San Francisco 🌉, hosted by Inferact at Anyscale Ray Summit. We'll have speakers from Inferact, NVIDIA, AMD, Google TPU, Anyscale, PyTorch, Meta, Red Hat, and key builders around vLLM. The talks on the roadmap deep dive into the latest on accelerators, training and serving pipelines, and production-scale inference 🚀 Link to the full schedule below 👇
-
-
Inferact reposted this
We're hiring a Founding Product Marketing Manager at Inferact! 🎉 Come work with the team and project at the heart of AI inference. What you'll do: • Define the brand and identity of both vLLM and Inferact • Own marketing partnerships with leading hardware, model, and cloud companies • Run global events for vLLM and Inferact, including the vLLM conference You may be a good fit if you are: • AI native and good at context-switching • Scrappy + biased towards action • Opinionated about branding, design, and how to host a good event You'll report directly to Simon Mo, and we'll also work closely together 🥳 . If this sounds interesting, please DM me or apply below :) https://lnkd.in/gqawaevd
-
Kimi K3 is one of the most powerful open-weight models ever released: 2.8T params, 1M context, native vision. 🚀 Getting it to serve well on day 0 took real engineering. The Inferact team led the vLLM integration: KDA-aware caching, MXFP4 MoE kernels, and a DSpark speculative-decoding model we trained and open-sourced. Huge thanks to Kimi (Moonshot AI), NVIDIA, AMD, and the broader vLLM community for the partnership. Link to our DFlash speculative decoding draft model in the comments below. Serve K3 on vLLM today.
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀 At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public. What K3 brings: 🧠 2.8T-parameter Mixture-of-Experts (16 of 896 experts active per token) 📚 1M-token context window 🛠️ Native multimodal understanding, including vision ⚡ Kimi Delta Attention: a hybrid of linear and full attention that makes million-token context affordable Huge thank you to Kimi (Moonshot AI) AI for the model release and partnership, Inferact for leading the vLLM optimizations, and to our partners at NVIDIA, AMD, and the broader vLLM community. Recipes, blog, and DSpark draft model in comments below 👇
-
Welcome to the team Joe Cotant!! Very excited to have you 🚀
After a great run at Anyscale, I started my next opportunity last week by joining Inferact (https://inferact.ai/). Inferact was founded by creators and core maintainers of vLLM (https://vllm.ai/), the most widely-used open-source LLM inference engine. Open models are tightening the gap on their closed counterparts, and optimizing the inference layer is becoming increasingly critical for delivering value in production. We're hiring for many roles across the stack. Please reach out if you or someone you know is interested! https://lnkd.in/gTCykeba
-
Day 0 and vLLM already runs Inkling at full feature parity. ⚡ We worked with Thinking Machines Lab to bring their 1T-parameter multimodal model to vLLM on launch day — text, image, and audio in, with up to 1M context. Performance: up to 380 tok/s/user with MTP on 4× GB200s. Getting there was a real lift — sconv-aware TP sharding, low-latency fused collectives, and direct integration of TML's new FA4 sheared-bias kernel. Huge shoutout to the team at Inferact that pulled this off on a Day-0 timeline. 🙌 Write-up + repro instructions 👇 https://lnkd.in/gp2h3D9d
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://lnkd.in/gY4NvS5h Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. Cost and latency are important in real-world use cases. Inkling's continuous thinking effort lets you pick your point on the cost/performance curve — reaching the same score with a fraction of the tokens. Inkling natively understands and reasons across text, audio, and images. It’s strong on audio in particular, ranking among the strongest open-weights models on VoiceBench, MMAU, and AudioMC. We’re grateful to our partners for their day-0 support across the open-source ecosystem: NVIDIA, Together AI, Fireworks AI, Databricks, Unsloth AI, Modal, Baseten, LightSeek Foundation, Inferact on VLLM, RadixArk on SGLang. Inkling is the first in a family. We’ve included some details of Inkling-Small, a lighter-weight model trained on a similar recipe, with full weights to follow. We hope you enjoy Inkling, and as always we’re keen to see what you build.
-
Bringing the vLLM community together in one room is exactly why we started Inferact. Proud to host the first vLLM Conference at Ray Summit 2026! August 24 - 26 in San Francisco, one ticket gets you access to both events. Read more 👇 https://lnkd.in/gBNFDCZM
The first vLLM Conference runs as its own track inside Ray Summit 2026, hosted by Inferact. One ticket gets you into both. Speakers include the people building vLLM and the ecosystem around it: → Woosuk Kwon, Creator of vLLM and Cofounder/CTO, Inferact → Simon Mo, Core Maintainer of vLLM and Cofounder/CEO, Inferact → Siyuan Fu, Senior Engineer, NVIDIA → Douglas Lehr, Principal Engineer, AMD → Qi Zhou, Senior Staff Software Engineer, Google TPU → Richard Zou, Senior Staff Software Engineer, Meta → Debarshi Raha, VP/Fellow Engineer, DigitalOcean → Greg Pereira, Senior Machine Learning Engineer, Red Hat Plus Hugging Face, Prime Intellect, Anyscale, Google DeepMind, and Together AI. San Francisco, Marriott Marquis, August 24–26. Early bird pricing ends this Friday, July 17. https://lnkd.in/gcdu7UVy
-
-
Proud to host the first vLLM Conference at Ray Summit, Aug 24–26 in SF 🚀🎉 We're bringing the community together to talk about what's next: 🏗️ Building and scaling AI in production 🌐 Deploying vLLM to run inference across your own cloud and hardware 🤖 What open source means for the future of AI 🌉 Come learn where the future of inference, open source, and AI is heading — and meet the leading builders driving it 👇 https://lnkd.in/gBNFDCZM
Announcing the first-ever vLLM Conference — hosted by Inferact at Ray Summit, Aug 24–26 in San Francisco 🎉🌉 This is where we'll get into the work pushing open, high-performance inference forward, such as: 🗺️ Where the vLLM roadmap is headed ⚡ Getting the most out of accelerators including NVIDIA, AMD, TPU 🔗 Wiring vLLM into training and serving pipelines 🚀 Running inference on production scale The summit features Speakers from Inferact, NVIDIA, AMD, Google TPU, Anyscale, PyTorch, Meta, Red Hat, and more 🎤 Come learn where the future of inference, open source, and AI is heading — and meet the leading builders driving it 👇 https://lnkd.in/gBNFDCZM
-
Inferact cohosted the AI Engineer World's Fair closing happy hour with vLLM & Novita AI last week 🍵 Matcha, egg tarts, and the Portugal v. Croatia match ⚽ in the background — plus conversations spanning vLLM, the latest frontier models, and where AI is headed. 140+ of you showed up; new connections made, ideas swapped 🙏 thank you all!
-