ggml.ai: Local Model Inference, llama.cpp, Hugging Face Deal

ggml.ai built GGML, llama.cpp, and whisper.cpp for local inference; its team joined Hugging Face, which pledged to keep the open-source projects community-run.

kindCompany
founded2023
Event history

A tensor library for hardware you already own

Georgi Gerganov founded ggml.ai in 2023 to fund continued development of GGML, a tensor library for machine learning built to run large models efficiently on commodity hardware rather than requiring a data-center GPU. Nat Friedman and Daniel Gross provided the company's pre-seed funding. The project describes its guiding principles as minimalism, a preference for plain C and C++ over heavier frameworks, and an open-source-first approach released under the MIT license. [1]

GGML is the engine underneath two much more widely used projects: llama.cpp, which runs large language models locally, and whisper.cpp, which does the same for OpenAI's Whisper speech recognition model. Between them, these are among the most common ways a developer runs an open-weight model on a laptop, phone or single server rather than through a hosted API. [1]

Why local inference matters

Atlas interpretation: Most attention in AI goes to training frontier models. GGML and llama.cpp sit on the other side of that: once a lab publishes model weights, someone still has to write the code that makes those weights run efficiently on ordinary hardware, with quantization and memory tricks that trade a small amount of accuracy for a large drop in compute and memory requirements. That work is what made open-weight releases from Meta, Mistral, DeepSeek and others usable by individual developers rather than only by organizations with server farms, and it is largely why llama.cpp became a common dependency across the local-AI tooling ecosystem rather than one hobbyist project among many. [1]

Joining Hugging Face

On February 20, 2026, ggml.ai announced that Gerganov and the GGML team were joining Hugging Face, with Hugging Face acquiring the company. Both companies said GGML and llama.cpp would remain fully open source and community-run, with Hugging Face providing long-term, sustainable resources rather than absorbing the projects into a closed product. Hugging Face committed to specific technical work as part of the deal: closer compatibility between its Transformers model definitions and llama.cpp, and easier packaging for local deployment. [3]

Atlas interpretation: The stated commitments are promises about how the acquisition will be run, not yet a settled track record, so this page treats them as intentions the projects can be checked against rather than results already delivered. [3]

ggml.ai on the timeline

ggml.ai's one event on the timeline is the acquisition itself, ggml.ai joins Hugging Face. The timeline also lists Hugging Face as its parent organization following the deal, reflecting the acquisition rather than a change in what the ggml and llama.cpp projects are or how they are licensed. [3]

Sources

  1. ggml.ai

    ggml.ai · Aug 20, 2026

  2. ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

    ggml.ai · Feb 20, 2026

  3. GGML and llama.cpp join HF to ensure the long-term progress of Local AI

    Hugging Face · Feb 20, 2026