Concept detail
prime-RL
1 tagged videos · 1 connection videos · degree 0
Publication timeline
- 2026-071 video
Source videos
Observed neighbors
Custom post-training algorithms
The framework is designed to make it easier to implement on-policy distillation, self-distillation, and other emerging algorithms.
1 source videos
Modern Post-Training: A Deep Dive — Will Brown, Prime IntellectVerifiers library
Complementary open-source tools covering the environment side and the training infrastructure side, respectively.
1 source videos
Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect