Each point is one post-trained version of the same 120B model on the safety (x) versus benign compliance (y) plane — up and to the right is better on both. Drag the slider to raise the reward-shaping strength α: the orange curve is a dose, trading falling compliance for rising safety. The two grey markers (plain DPO, data-filtered DPO) are reference arms that used no reward shaping.

reward shaping (α sweep)   plain DPO   data-filtered DPO
0−1−2−5