llama.cpp

Efficient C++ inference engine with quantization, powers Ollama and other runtimes

Open FounderOS desktop