Magnitude launches local inference engine for AI agents
The open-source project says its software tunes kernels on a user’s device and can run open models up to twice as fast as llama.cpp, a claim not independently verified in the supplied material.
Magnitude has introduced an open-source inference engine designed to run open models for AI agents. Its GitHub repository says the software compiles and tunes kernels on a user’s device before running a model, adapting them to the hardware.
The project claims speeds of up to twice those of llama.cpp. The repository includes no independent benchmark results or testing methodology in the supplied material, so the claim’s performance across hardware, models and workloads is unclear.
Magnitude says the engine supports Apple Silicon, NVIDIA and AMD GPUs, as well as CPUs, and is available for macOS, Windows and Linux. It packages a desktop app with the Magnitude command-line interface.
The software can connect to tools including Pi, OpenCode, Hermes, Codex and Cline, according to the repository. Magnitude also offers an OpenAI-compatible API for other tools.
Magnitude says prompts, files and models remain on the user’s machine, and that an internet connection is not required after a model has been downloaded.


