High-Throughput Asynchronous RLHF with In-VRAM Tensor-Native Rewards & Second-Moment Off-Policy Control (M2PO / GRPO).
Description excerpted from the original listing, which is linked below.