Orpheus generates speech frame by frame, so every utterance has a real time axis inside it. Pick a clip and press play: the dot sweeps the model's per-frame activation path through its emotion geometry (left), while the right panel compares the emotion the model is internally tracking against the emotion actually present in the audio it produces. Arousal (energy) tracks closely; switch to valence (pleasantness) to see the axis this readout cannot resolve.

Activation path (PC1–PC2), dot = current frame
Arousal over time — internal (slate) vs produced audio (ember)