Dust algorithm offers a search-based alternative to backpropagation for transformer training
A new zeroth-order method estimates learning updates by perturbing activations rather than weights, showing competitive results in limited pretraining experiments.
A research paper introduces Dust, a zeroth-order optimisation method designed to pretrain transformer models without relying on backpropagation. The algorithm estimates learning updates by perturbing transformer activations and scoring the resulting effects on token losses, effectively using a search-based approach to credit assignment.
The authors frame backpropagation as an effective low-compute bias that has shaped the co-evolution of deep learning architectures, optimisers, and hardware. They argue that as compute becomes more abundant, more generic and brute-force learning algorithms based on search could offer advantages over methods constrained by differentiability. Dust aims to replace the analytic structure of backprop with a method that relies heavily on computation.
In the reported experiments, the team tested models ranging from 2 million to 243 million parameters with training budgets up to 20 million tokens. The authors report that Dust approached the performance of backpropagation and, in some experiments, surpassed it as the perturbation population increased. At smaller token budgets of 100k and 1M, Dust ended below backpropagation, while at 10M and 20M tokens, the gap narrowed significantly with larger populations.
Despite these results, the authors caution that Dust is currently far less compute-efficient than backpropagation. They state that the method would need orders of magnitude greater efficiency to become a practical replacement for standard training procedures. A fitted estimate of Dust’s limiting loss at 20 million tokens is described as loosely constrained, suggesting the gap is still closing rather than having reached a confirmed lower limit.
The method uses a concept called a virtual population, where perturbations are applied independently at every token rather than to model weights. This allows a single forward pass to evaluate a population at least three orders of magnitude larger than traditional weight-space evolution strategies. The authors note that this approach could eventually make it possible to train architectures that are difficult to optimise with backpropagation, such as those with external programs in the loop.
The paper leaves open questions regarding whether Dust can scale to larger models and training budgets. While the cosine similarity between Dust’s estimates and backprop gradients remains stable across two orders of magnitude in tokens, the mechanism by which it might find better directions than first-order gradients is not yet fully clear.


