Q Labs published Dust (Hacker News front page), a zeroth-order, forward-only method that perturbs activations independently per token so one forward pass evaluates thousands of virtual population members, then estimates gradients from loss changes. Pretraining GPT-style transformers on FineWeb, Dust beats tuned backprop at 100k and 1M tokens and closes the gap with population at 10M and 20M; the largest model tested is 243M parameters. The authors say it is not yet compute-efficient enough to replace backprop.

Key Takeaways

  • ✓At 1M tokens Dust's test loss drops below backprop from ~1k draws; at 100k from a few hundred (paper)
  • ✓20M tokens: power-law fitted limit 4.431 (95% CI 3.89–4.58) vs backprop 4.633; authors call it loosely constrained trend evidence
  • ✓Efficiency: ~10³–10⁴× better than weight-space ES (EGGROLL); EGGROLL at 256× the population still trails Dust at 64 draws
  • ✓Counterintuitive: at 10M tokens a 243M model beats a 120× smaller one at most population sizes
  • ✓qlabs-eng/dust (MIT): default population 16,384, 8-GPU torchrun for 1M/10M runs, plus a smaller single-GPU setup
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Q Labs' Dust is a zeroth-order training method that replaces backprop's chain rule with a population estimate. It adds Gaussian noise to every linear layer's output independently at each token, rewards each perturbation by the change in that token's loss (with decayed future-token credit for keys/values), averages reward-weighted noise into an output-error estimate, and takes its outer product with the layer input to form the weight gradient. Because each token is a virtual population member, one forward pass evaluates thousands of members. On GPT-style transformers trained on FineWeb, Dust beats tuned backprop at 100k and 1M tokens, closes the gap with population at 10M and 20M (fitted 20M limit 4.431 vs backprop 4.633, loosely constrained), is about 10^3 to 10^4 times more efficient than the weight-space ES method EGGROLL, and gets more population-efficient with size up to 243M parameters. The authors say it is not yet compute-efficient enough to replace backprop; the MIT-licensed qlabs-eng/dust repo provides a minimal implementation and a backprop baseline.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.