Reading, steering, and training model behavior toward safer, more reliable outputs.
A live monitor that flags an unsafe tool call before it runs, powered by a linear probe that reads the model's own activations rather than its words
A hallucination probe, then RL that uses it as a reward to cut the hallucination rate
Reward shaping is a tunable safety dose for DPO — a small α-sweep maps how much over-refusal each unit of added safety costs
A reader-driven tour of the features alignment amplifies and suppresses, and how much of it is visible in the preference data before training starts