Why It Pays to Hack OpenAI and Anthropic Models
When Dan Lahav was in grade school, he became obsessed with the short stories of Isaac Asimov, where robots gain sentience and trick humans into relinquishing control of society. Today he spends his time prodding artificial intelligence to do just that.
Lahav is CEO of Irregular, a two-year-old startup that specializes in testing AI models’ capacity for malice. OpenAI, Anthropic, and Google DeepMind have hired Irregular, which was previously named Pattern Labs, to stress-test how people might use their models for cyberattacks or fraud and recommend ways to prevent the bad behavior.