AnalysisAI ModelsOctober 8, 2026

Paper proposes Q-Learning with Scalar Adjoint Matching for flow policies

Read original source →arxiv.org

The method targets off-policy RL fine-tuning of flow policies, which is difficult because the policy generates actions over many steps against a learned value function. Author listed as Yonghoon Dong.

1 source

More stories today

Open the live feed