Yeah. This work is over claiming the novelty quite a bit.
Any of the Best of N papers that exploded in popularity after GRPO. Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work. That sequential LLMs have a step-wise probability is…
You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling. GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free.
That's fair. On the other hand, Minibatch OT (optimal transport) was one of the more fundamental advancements early on in flow matching & rectified flow models. This best of K approach effectively discards matches that…
This is just GRPO (proposed by DeepSeek), which similarly samples many plausible generations, selects the best of K, and trains that sample. Minibatch OT in flow matching also has a very similar mechanism, where samples…
Yeah. This work is over claiming the novelty quite a bit.
Any of the Best of N papers that exploded in popularity after GRPO. Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work. That sequential LLMs have a step-wise probability is…
You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling. GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free.
That's fair. On the other hand, Minibatch OT (optimal transport) was one of the more fundamental advancements early on in flow matching & rectified flow models. This best of K approach effectively discards matches that…
This is just GRPO (proposed by DeepSeek), which similarly samples many plausible generations, selects the best of K, and trains that sample. Minibatch OT in flow matching also has a very similar mechanism, where samples…