Kymata Labs/The Living IndexesBuilt by tekvisions ↗
The Gateway Index / Gateways & Proxies / #52
yalun753

yalun753/moe-l2

by yalun753 · Gateways & Proxies · updated 2d ago

MoE expert offload for low-VRAM GPUs — run 100B+ MoE models (DeepSeek, Qwen, Mixtral) on 8 GB cards. Expert cache with LRU hot experts, OpenAI-compatible proxy, GGUF multi-shard.

64
momentum
142
stars
2
forks
#52
rank
consumer-gpudeepseekexpert-cachingllama-cppllm-inferencemoevram-optimization
View on GitHub →