Tag: reinforcement-learning

All the articles with the tag "reinforcement-learning".

Inside GLM-5.2: IndexShare, KVShare, and the End-to-End TV Loss

21 Jun, 2026

A deep dive into GLM-5.2 — a 753B open-weight MoE that serves a 1M-token context. We walk the three innovations that make it cheap to run: IndexShare (cross-layer sparse-attention index reuse), KVShare + rejection sampling for speculative decoding, and a novel end-to-end TV loss that breaks the entropy bound on MTP acceptance. Plus the slime RL stack behind its long-horizon agentic skills.
GRPO and DAPO: A Deep Dive into RL for Reasoning LLMs

28 May, 2026

An end-to-end walkthrough of Group Relative Policy Optimization (GRPO) and Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) — the two RL algorithms that drive open reasoning models in 2025–2026. Full math, every design choice motivated, and a head-to-head comparison.
From GRPO to GSPO: Group-Based Policy Optimization for LLMs

28 May, 2026

A complete walkthrough of Group Relative Policy Optimization (GRPO) and Group Sequence Policy Optimization (GSPO) — the policy-gradient algorithms behind DeepSeek-R1 and Qwen3. Full math, the failure mode that motivated GSPO, the MoE story, and a side-by-side comparison.
GRPO and Dr.GRPO: The Math, the Biases, and the Fix

28 May, 2026

An end-to-end derivation of Group Relative Policy Optimization (GRPO) from DeepSeekMath and the Dr.GRPO correction from Liu et al. Covers the full objective, the gradient, the two biases (length and question difficulty), the unbiased fix, and the practical recipe behind R1-Zero–style training.

Tag: reinforcement-learning

Inside GLM-5.2: IndexShare, KVShare, and the End-to-End TV Loss

GRPO and DAPO: A Deep Dive into RL for Reasoning LLMs

From GRPO to GSPO: Group-Based Policy Optimization for LLMs

GRPO and Dr.GRPO: The Math, the Biases, and the Fix