End-to-end recipes from real Silico experiments — each pairs the research workflow with the science, and forks into your own workspace so you can pick up where it left off.
A live monitor that flags an unsafe tool call before it runs, powered by a linear probe that reads the model's own activations rather than its words
A reader-driven tour of the features alignment amplifies and suppresses, and how much of it is visible in the preference data before training starts
Reward shaping is a tunable safety dose for DPO — a small α-sweep maps how much over-refusal each unit of added safety costs
A generation-time nudge that recovers novelty from a copy-prone design model
a linear probe + steering vector for tandem-repeat copy number
A hallucination probe, then RL that uses it as a reward to cut the hallucination rate
A map of emotion inside a language model — a polar readout plus a steering dial
A GRPO run that trains a 4B agent to GPT-4.1 level
Reproducing a molecular-tumor-board agent, then swapping out one subsystem at a time to find which design choices actually move the score
A reading of one human genome with a genomic foundation model: what its DNA carries, and why carrying a disease variant is almost never the same as having the disease.
A faithful replication that reads bacterial phylogeny out of a genomic model's embeddings
A worked tutorial: find a structured concept hiding in a model's activations, prove it is real, and steer the model along it — on the seven days of the week in Llama-3.1-8B.
Figure 2 of Mixing Mechanisms, rebuilt from scratch on causalab and verified against the raw interventions, then replicated across two model families — from a single Silico prompt
An autoregressive speech model carries a moving, frame-by-frame emotion signal as it talks. Reading that signal shows arousal is the trustworthy axis, and a single direction added during generation steers the prosody.
Recovering the IOI circuit from Wang et al. (2022) with a native path-patching primitive and no original-paper code, then a preliminary look at the same circuit in an 8B model
Recovering the causal-tracing sites from Meng et al. (2022) with a native activation-restoration primitive and no original-paper code, then testing whether the same structure appears in an 8B model
A ten-head gaze-steering handle that overrides a VLM's text prompt — a click-to-steer demo plus a prompt-conflict sweep
A general-purpose sparse autoencoder that turns Qwen3-4B into a browsable atlas of 40,960 labeled concepts
A SelfIE read-out of a model's weekday geometry — a nameable ring plus a causal-subspace test