Hardware

Magnitude Launches Self-Optimizing Inference Engine

Y Combinator startup Magnitude has launched an open-source local inference engine that compiles custom kernels on a user's device to run AI models up to twice as fast as llama.cpp.

Hacker News1 day agoHardware
Image: Hacker News

The Y Combinator-backed startup Magnitude has released its self-optimizing, open-source inference engine designed specifically for running artificial intelligence agents. Unlike traditional engines like Ollama, LM Studio, or llama.cpp, which distribute precompiled kernels for broad hardware categories, Magnitude compiles and tunes its kernels directly on the user's local machine before running a model. This hardware-tailored approach allows the engine to run on Apple Silicon, NVIDIA GPUs, AMD GPUs, or standard CPUs without requiring a fixed minimum specification.

According to the developers, this on-device optimization delivers performance up to 2x faster than llama.cpp. Specifically, benchmarks show a 92% faster decode speed on Apple's Metal framework and a 19% speedup on NVIDIA's CUDA platform. Additionally, the engine is designed to be highly resource-efficient, utilizing 27% less memory per agent and immediately freeing up those system resources once the agent stops running. It also supports fast concurrent sessions by sharing prefix caches to prevent performance slowdowns.

For practitioners, the Apache 2.0-licensed engine functions as a local desktop application for macOS, Windows, and Linux, which includes a command-line interface. It offers one-click integration with popular agent frameworks and tools such as Pi, OpenCode, Hermes, Codex, OpenClaw, Claude Code, Oh My Pi, and Cline. Any unsupported tools can connect via an OpenAI-compatible API. Because the software runs entirely locally, all prompts, files, and downloaded models remain private on the user's machine, eliminating ongoing token costs and the need for an active internet connection after the initial download.

This is our own summary of reporting by Hacker News

More in Hardware