A benchmark for auditing LLM safety interventions, part of the lab's work on interpretable interventions and evals for AI alignment and model security.
Current Work
Ultra-efficient, on-device architectures for few-shot autonomous visual reasoning in open-world physical environments.
An IMU-based wearable that quantifies Parkinson's progression from spinal curvature at roughly 1% the cost of radiographic monitoring.
Projects
Nearly 70 million Americans have no local news source. Relay turns two million pages of state legislative documents and ten thousand hours of hearing video into personalized, citation-backed briefings, via a multi-agent pipeline and multimodal RAG.
Directions
Where this is heading: representation engineering as the instrument layer for alignment, and knowledge graphs as the structure that makes edge reasoning composable.
Outlook
The two problems I work on, AI alignment and visual reasoning on edge devices, are one problem. Safety claims today are asserted, not quantified: a model is called aligned because it refuses the right prompts, not because anyone checked what it can still do. Capable AI is assumed to need a datacenter behind it. Both assumptions should fail together.
Alignment should be a measurement, made with instruments that read what a model computes rather than what it says. Intelligence should be cheap enough, in parameters and power, to run on a drone or a wearable, adapting without gradients or a connection home. An AI you can trust is one whose limits you can state precisely, and an AI that must phone home is not really deployed. My work is an attempt at both halves: benchmarks that make safety claims falsifiable, probes that read internal representations, architectures that reason on a battery.