GitHub Unveils Breakthrough: Runs 2.8T-Parameter Model on Single Apple Silicon Mac with LLM Technology, Featuring MXFP4 On-Demand Experts and NEON Kernels via HTTP and MPS Compute, with OpenAI-Compatible API Server
agents apple meta openai
| Source: Mastodon | Original article
A massive 2.8T-parameter language model can now run on a single Apple Silicon Mac.
GitHub has introduced a new project called Deltafin, which enables users to run Kimi K3, a massive 2.8T-parameter Mixture-of-Experts LLM, on a single Apple Silicon Mac. This project is significant as it allows for the deployment of a large language model on a local machine, leveraging technologies such as NEON kernels, Metal/MPS compute, and an OpenAI-compatible API server.
The ability to run such a large model locally matters because it provides users with more control over their data and computing resources. Additionally, it demonstrates the potential for running complex AI models on consumer-grade hardware, which could have implications for the development of more accessible and private AI applications.
As this project is still in its early stages, it will be interesting to watch how it evolves and whether it can achieve faster performance on newer Apple Silicon chips. The Deltafin project's focus on exact, reproducible, and open-source solutions may also attract attention from researchers and developers looking to push the boundaries of AI capabilities on local machines.
Sources
Back to AIPULSEN