Gökdeniz Gülmez shipped MLX-LM-LoRA v5.8.6 on Oct 5 (GitHub Release and PyPI), the first release after 3.1.3. It adds DSLA with DPO/ORPO/CPO objectives and critic-free KLPO with token/sequence routes and several KL estimators, improves GRPO stability and microbatching for Online DPO, XPO, RLHF REINFORCE and PPO, and adds fast VJP for gated-delta layers. Minimums rise to mlx>=0.32.3, mlx_lm>=0.32.0 and Python>=3.11; license standardized on Apache 2.0.

Key Takeaways

  • ✓Jump from v3.1.3 (Sep 16) straight to v5.8.6, merging nine PRs (#64, #94, #96–#102).
  • ✓Two new modes: --train-mode dsla (--dsla-loss dpo|orpo|cpo, default --latent-weight 0.1) and --train-mode klpo (--klpo-route token|sequence, estimators mc/binary/topk/full).
  • ✓KLPO skips GRPO group normalization, PPO ratio clipping and the reference model, reusing GRPO reward callbacks and generation/scoring.
  • ✓Raised minimums: mlx>=0.32.3, mlx_lm>=0.32.0, Python >=3.11.
  • ✓QAT now covers DSLA; new GitHub Pages site with a Python API reference.
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

MLX-LM-LoRA is an Apple MLX library for training LLMs on Apple Silicon. v5.8.6 (Oct 5, 2026) is the first release after v3.1.3 and focuses on post-training. New: DSLA, which adds prompt-response similarity and batch-direction latent supervision on top of a DPO, ORPO or CPO objective (DSLA-DPO uses a frozen reference; ORPO/CPO are reference-free), and KLPO, a critic-free RL method with token- or sequence-level regression and MC, binary, top-k or full KL estimators, with no GRPO group normalization, PPO clipping or reference model. GRPO scoring is more stable; Online DPO, XPO, RLHF REINFORCE and PPO gain microbatching; gated-delta layers get a fast VJP with a checkpointed fallback. QAT now covers DSLA, a new docs site includes a Python API reference, and the license is Apache 2.0. No speed or quality benchmarks were published. Upgrade with pip install -U mlx-lm-lora; it needs mlx>=0.32.3, mlx_lm>=0.32.0 and Python>=3.11.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.