1. X
  2. LMCache Lab
Log inSign up
LMCache Lab
257 posts
user avatar
LMCache Lab
@lmcache
๐Ÿงช Open-Source Team that maintains LMCache and Production Stack ๐Ÿค– Democratizing AI by providing efficient LLM serving for ALL
Github, Online
lmcache.ai
Joined September 2024
51
Following
1,920
Followers
RepliesRepliesMediaMedia
  • user avatar
    LMCache Lab
    @lmcache
    Aug 7
    LMCacheย ๐Ÿคย storage at #FMS2026 Nearly every storage vendor at FMS used LMCache to show off their KV cache offload performance this week. This is a great example of what a common, open source KV cache layer can enable as KV cache becomes a bigger part of the AI inference stack.
    user avatar
    Tensormesh
    @tensormesh
    Aug 7
    That's a wrap on #FMS2026! ๐ŸŽ‰ Thanks to everyone who stopped by to connect with Tensormesh. A highlight: Yihua's panel on managing KV cache across GPUs, host memory & storage as context windows and inference workloads scale. Great conversations all around. Plus, it was great to
    Image
    Image
    Image
    Image
  • user avatar
    LMCache Lab
    @lmcache
    Aug 6
    ๐Ÿš€ ๐—Ÿ๐— ๐—–๐—ฎ๐—ฐ๐—ต๐—ฒ ๐˜ƒ๐Ÿฌ.๐Ÿฑ.๐Ÿฏ ๐—ถ๐˜€ ๐—ผ๐˜‚๐˜! โœจ New model support โ€ข Kimi K3 (day-0, incl. MLA-layer KV shapes) โ€ข MiniMax M3 and other Mamba models fixed for vLLM 0.26 ๐Ÿ—‚๏ธ Fleet-wide cache directory (MP mode) โ€ข MP servers stream cache events into a shared key + token directory,
    Image
    Made with AI
  • user avatar
    LMCache Lab
    @lmcache
    Jul 29
    ๐—Ÿ๐— ๐—–๐—ฎ๐—ฐ๐—ต๐—ฒ ๐—ต๐—ฎ๐˜€ ๐˜€๐˜‚๐—ฝ๐—ฝ๐—ผ๐—ฟ๐˜๐—ฒ๐—ฑ ๐—ž๐—ถ๐—บ๐—ถ ๐—ž๐Ÿฏ ๐—ณ๐—ฟ๐—ผ๐—บ ๐——๐—ฎ๐˜† ๐Ÿฌ! Kimi K3 combines Kimi Delta Attention (KDA) recurrent-state layers with Multi-head Latent Attention (MLA) layers. LMCache can store and reuse both cache types, enabling prefix reuse across repeated
    Image
    Made with AI
  • user avatar
    LMCache Lab
    @lmcache
    Jul 28
    ๐—ช๐—ต๐—ฎ๐˜ ๐—ต๐—ฎ๐—ฝ๐—ฝ๐—ฒ๐—ป๐˜€ ๐˜๐—ผ ๐˜†๐—ผ๐˜‚๐—ฟ ๐—ž๐—ฉ ๐—ฐ๐—ฎ๐—ฐ๐—ต๐—ฒ ๐˜„๐—ต๐—ฒ๐—ป ๐—ฎ ๐—ฅ๐—”๐—š ๐—พ๐˜‚๐—ฒ๐—ฟ๐˜† ๐—ป๐—ฒ๐—ฒ๐—ฑ๐˜€ ๐˜๐˜„๐—ผ ๐—ฑ๐—ผ๐—ฐ๐˜‚๐—บ๐—ฒ๐—ป๐˜๐˜€ ๐˜๐—ต๐—ฎ๐˜ ๐˜„๐—ฒ๐—ฟ๐—ฒ ๐—ป๐—ฒ๐˜ƒ๐—ฒ๐—ฟ ๐—ฐ๐—ฎ๐—ฐ๐—ต๐—ฒ๐—ฑ ๐˜๐—ผ๐—ด๐—ฒ๐˜๐—ต๐—ฒ๐—ฟ? Prefix caching runs into two problems: โ†’ Document B's cached state was computed without attending
    Image
  • user avatar
    LMCache Lab
    @lmcache
    Jul 22
    ๐Ÿš€ ๐—Ÿ๐— ๐—–๐—ฎ๐—ฐ๐—ต๐—ฒ ๐˜ƒ๐Ÿฌ.๐Ÿฑ.๐Ÿฎ ๐—ถ๐˜€ ๐—ผ๐˜‚๐˜! โœจ New model support โ€ข MiniMax M3 ๐Ÿ”Œ New L2 backends (MP mode) โ€ข Azure Blob Storage, Valkey (standalone + cluster), Cloud Bigtable, SageMaker HyperPod ๐Ÿ”ง Highlights โ€ข Async engine-driven store, Pin/Unpin + token-based delete APIs
    Image
    Made with AI

Log in or sign up for X

See whatโ€™s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsยทPrivacyยทCookiesยทAccessibilityยทAds Infoยทยฉ 2026 X Corp.
Advertisement
Advertisement