lyogavin/airllm
減少推論記憶體用量,讓 70B LLM 在單張 4GB GPU 上運行。 AirLLM 70B inference with single 4GB GPU
chinese-llmchinese-nlpfinetunegenerative-aiinstruct-gptinstruction-setllamallmloraopen-models
AI 分析 AI analysis
AI 閱讀專案 README 後撰寫的導讀。 Written by AI after reading the project README.
AirLLM 讓你用單張 4GB 顯卡就能跑 70B 大模型,甚至 405B Llama 3.1 只需 8GB、671B DeepSeek-V3 只要約 12GB,關鍵是不做量化或剪枝。如果你的硬體有限卻想玩超大模型,這專案值得一試;對初學者來說,pip install 後用 AutoModel 就能推理,上手門檻很低。
AI 推薦理由 Why AI picked it
記憶體優化機制是逐層或逐 expert 串流載入,並有可驗證的 VRAM 數據(405B 在 8GB、Kimi K3 低於 4GB),支援多種開源模型與 pip 安裝、notebook 範例。適合 GPU 資源受限但仍想執行大型模型的開發者,頻繁更新顯示成熟度佳。 AirLLM streams layers or MoE experts to cut memory usage, with measured figures like 405B on 8GB and Kimi K3 under 4GB. It supports many open models via pip and notebooks, targeting developers with limited GPU resources. Frequent updates indicate mature maintenance.
上榜紀錄 Trending history
- 2026-08-31 每月 Monthly #21
- 2026-08-30 每月 Monthly #18
- 2026-08-29 每月 Monthly #18
- 2026-08-28 每月 Monthly #20
- 2026-08-24 每月 Monthly #19
- 2026-08-22 每月 Monthly #18
- 2026-08-21 每月 Monthly #21
- 2026-08-20 每月 Monthly #20
- 2026-08-19 每月 Monthly #21
- 2026-08-18 每月 Monthly #18
- 2026-08-17 每月 Monthly #17
- 2026-08-16 每月 Monthly #16
- 2026-08-12 每週 Weekly #14
- 2026-08-11 每週 Weekly #6
- 2026-08-10 每週 Weekly #4
- 2026-08-09 每週 Weekly #3
- 2026-08-08 每週 Weekly #5
- 2026-08-07 每週 Weekly #4
- 2026-08-06 每日 Daily #13
- 2026-08-06 每週 Weekly #3
- 2026-08-05 每日 Daily #8
- 2026-08-05 每週 Weekly #7
- 2026-08-04 每日 Daily #1
- 2026-08-04 每週 Weekly #17
- 2026-08-03 每日 Daily #3
- 2026-07-20 每日 Daily #17
- 2026-07-19 每日 Daily #8