Each point is one post-trained version of the same 120B model on the safety (x) versus benign compliance (y) plane — up and to the right is better on both. Drag the slider to raise the reward-shaping strength α: the orange curve is a dose, trading falling compliance for rising safety. The two grey markers (plain DPO, data-filtered DPO) are reference arms that used no reward shaping.