Tech

AI Labs Strip Knowledge from Models to Boost Reasoning and Cut Costs

Leading laboratories are deliberately reducing world knowledge in large language models to lower per-token compute costs and improve reasoning efficiency, a move that redefines how AI handles information retrieval and hallucination.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Strategic shift prioritises procedural logic over factual recall, enabling high-performance models to run on consumer hardware

Leading artificial intelligence laboratories are intentionally stripping world knowledge from large language model weights to prioritise reasoning capabilities and reduce per-token compute costs. This strategic shift allows smaller models, such as Qwen3.5 and DeepSeek V4-Flash, to achieve high performance on mathematical and coding benchmarks while significantly lowering hardware requirements. By offloading factual recall to external retrieval systems and tools, developers aim to create models that are less prone to hallucination, easier to update, and capable of running on consumer-grade hardware.

Performance metrics illustrate the efficacy of this approach. GLM-5.2 achieved a 99.2% score on the AIME 2026 benchmark with approximately 40 billion active parameters per token. Qwen3.5 scored 91.3% with 17 billion active parameters, fitting into 6GB of VRAM when quantised, which roughly doubles the score of the next best model under 10 billion parameters on Artificial Analysis's intelligence index. DeepSeek V4-Flash operates with 13 billion active parameters. In contrast, GPT-4, released in 2023, was rumoured to use around 280 billion active parameters and struggled with similar AIME problems.

The trade-off becomes apparent when examining factual recall. On the SimpleQA benchmark, which measures factual recall without allowing tools, the current leader is Gemini 2.5 Pro at 53%, indicating that even the best current models miss half of factual questions when no tools are allowed. Small models perform poorly in this regard; Artificial Analysis measures hallucination rates for Qwen3.5 4B and 9B models at 80% to 82% on its knowledge benchmark. When these models do not know a fact, they frequently generate plausible but incorrect answers.

Research suggests factual knowledge capacity is approximately two bits per parameter, whereas reasoning procedures compress more efficiently. Facts require significant storage space, contributing to the growth of frontier models to trillions of parameters. Reasoning, however, involves a relatively small set of procedures such as breaking problems into parts and checking intermediate states, which transfer well into smaller models through distillation and reinforcement learning.

This design decouples the expensive, slow artifact of the trained model from the rapidly changing state of the world. Facts baked into weights have a short shelf life and become outdated quickly, whereas reasoning procedures remain stable. By keeping only generalist knowledge in the weights and retrieving specific details at runtime via external tools, models can be updated without costly retraining.

The trend points toward frontier-quality reasoning models running on single consumer GPUs, such as 24GB cards, within a couple of years. This is achieved by stripping out expert layers that primarily store facts, allowing total model size to shrink toward active size. While these models may not know much without tools, they offer a local, cost-effective solution where incorrect facts can be traced and corrected in external knowledge bases rather than within the model weights.

Continue reading

More from Tech

Read next: Tiny386 PC emulator ported to Raspberry Pi Pico 2-class hardware
Read next: Readers turn to AI chatbots for personalised fiction and role-play
Read next: Septuagint’s contested history comes into focus in review of Timothy Michael Law’s book