llama.cpp launches official platform for local frontier AI deployment
The new platform supports hardware ranging from laptops to clusters, utilising hand-tuned kernels to execute dense and mixture-of-experts models locally.
The llama.cpp project has officially launched its dedicated platform, llama.app, providing a centralised hub for users to deploy frontier artificial intelligence models directly on their own hardware. The release marks a significant step in the open-source ecosystem, offering a solution that allows users to run complex AI workloads without relying on external application programming interfaces, telemetry services, or usage limits.
According to the platform’s documentation, the software is engineered to operate across a wide spectrum of computing resources, from standard laptops to large-scale clusters. It utilises a single binary and hand-tuned kernels compatible with all graphics processing units and central processing units, ensuring consistent performance regardless of the underlying infrastructure.
The platform supports a diverse array of models from major technology firms. This includes Alibaba’s next-generation multimodal reasoning models, which feature both dense and mixture-of-experts variants designed for coding and vision tasks. The site also lists support for Google’s open models, which are built using Gemini 3 technology and are capable of handling agentic workflows and over 140 languages.
Additionally, the launch highlights OpenAI’s first open-weight models since GPT-2. These models are described as being built for reasoning and agentic tasks, with specific capabilities for developer use, including function calling and tool use. The platform also notes support for Google’s multimodal models, which offer up to 128K context for both edge and cloud deployments.
To streamline the user experience, the platform introduces a 'pi-llama' plugin that automatically discovers local models without requiring manual configuration. The service emphasises data sovereignty, stating that files remain on the user’s machine and requests never leave the local environment, addressing growing concerns regarding privacy and data ownership in the AI sector.
This development occurs against a backdrop of rapid expansion in the broader AI landscape. Earlier this month, Google’s Gemini app reached one billion monthly active users, while OpenAI continued to roll out updates to its voice mode and launched ChatGPT Health in the US. The launch of llama.app provides an alternative infrastructure layer, allowing institutions and individuals to maintain control over their AI deployments amidst this consolidation of market power.
