University of California, Berkeley
Center for Human-Compatible AI (CHAI)
United States · PhD/Postdoc
CHAI is one of the most influential academic research centers in the field of AI alignment. Founded by Stuart Russell, it is dedicated to ensuring that AI system objectives remain aligned with human intentions. The center focuses on core topics including inverse reinforcement learning (IRL), collaborative human-AI interaction, and scalable alignment, and serves as a key origin point for academic research in AI safety.
学术历史
CHAI was established in 2016, co-founded by Stuart Russell and several researchers, building upon the human-compatible AI framework articulated in his book Human Compatible (2019). Prior to the center's founding, Russell and Peter Norvig's co-authored textbook Artificial Intelligence: A Modern Approach had already integrated AI safety topics into mainstream curricula. The establishment of CHAI marked the transition of AI alignment from philosophical discourse to systematic academic research, and its proposed inverse reinforcement learning paradigm became one of the core technical approaches to alignment research.
当前状态
CHAI is currently led by multiple core faculty members, with research directions encompassing inverse reinforcement learning, learning from human feedback, theoretical foundations of AI safety, and collaborative robotics. The center maintains close collaborations with frontier AI companies such as Anthropic and OpenAI, regularly publishing technical reports and safety research papers. It offers a doctoral training program that accepts applications from the Department of Computer Science and the Department of Statistics. The center also participates in the BAIR (Berkeley AI Research) consortium, sharing computational resources and facilitating cross-laboratory collaboration.
实验室 / 研究中心
Center for Human-Compatible AI
访问页面 →关键人物
Stuart Russell
Professor / Center Founder
A leading figure in AI, Professor in the Department of Computer Science at UC Berkeley, and author of Artificial Intelligence: A Modern Approach, the most widely adopted AI textbook worldwide. He proposed the human-compatible AI framework and advocates for uncertainty-centered alignment methods. In 2019, he published Human Compatible, a systematic exposition of the AI safety vision.
Anca Dragan
Associate Professor
Researches human-computer interaction and AI safety, focusing on human model construction and value learning. She has participated in the OpenAI safety team and was named to MIT Technology Review's Innovators Under 35.
Dawn Song
Professor
Expert in computer security and AI safety, researching adversarial examples, model privacy, and federated learning security. Recipient of the MacArthur Fellowship.
标志性成果
- Proposed the human-compatible AI framework (Stuart Russell, 2019), systematically defining a formalized solution to the alignment problem
- Theoretical foundations and algorithmic development of inverse reinforcement learning (IRL), providing a technical pathway for learning values from human behavior
- Introduced cooperative inverse reinforcement learning (CIRL), modeling the value inference problem in human-AI collaboration
- Systematic analysis of reward hacking in deep reinforcement learning, revealing potential risks of RLHF
学术资源
证据
师资
8 位相关教师
知名:Stuart Russell、Anca Dragan、Dawn Song、Pieter Abbeel、Stuart Russell Lab
研究产出
CHAI-affiliated researchers consistently publish AI safety and alignment papers at top conferences including NeurIPS, ICML, and ICLR, with recent focus on RLHF, scalable alignment, and mechanistic interpretability
就业去向
Graduates join frontier AI safety teams at Anthropic, OpenAI, DeepMind, Google DeepMind, and others, or take faculty positions at leading universities such as MIT and Stanford
发现信息有误? 提交纠错