creators
Georgi Gerganov - creator of llama.cpp and ggml
Georgi Gerganov wrote ggml and llama.cpp, the engine behind most local LLM inference, and now maintains them at Hugging Face from Sofia, Bulgaria.
Georgi Gerganov is a Bulgarian software engineer based in Sofia and the author of ggml, a dependency-free C tensor library, and of llama.cpp, the C/C++ inference engine that made it practical to run large language models on ordinary CPUs and consumer hardware. He started ggml towards the end of September 2022 and released llama.cpp on 10 March 2023, days after Meta's LLaMA weights appeared. The project became the base layer for local inference, alongside whisper.cpp, his earlier C/C++ port of OpenAI's Whisper, and the GGUF file format that grew out of it. He founded ggml.ai in 2023, with pre-seed funding from Nat Friedman and Daniel Gross, to keep the work going. On 20 February 2026 ggml.ai joined Hugging Face, with Gerganov and his team — among them Xuan-Son Nguyen and Aleksander Grygier — keeping full autonomy over llama.cpp's technical direction. Nvidia agreed in September 2026 to acquire Hugging Face; Gerganov said the project will stay hardware-agnostic and community-driven.
- Specialization
- on-device LLM inference, tensor libraries, C/C++ systems programming, open-source maintenance
- Country
- Bulgaria