Drag the firing threshold and watch how many unsafe actions the guardrail catches (recall) versus how often safe actions trip a false alarm. Toggle between the home domain the probe was tuned on and a new domain (online-shopping tasks) to see why one fixed line cannot serve both. Recall and false alarms are computed live over the evaluation set drawn below.

lenient (score 0)threshold 0.970strict (score 1)
unsafe actions caught (recall)
safe actions falsely flagged
of unsafe in this set