Hadrien Pouget will make a presentation on Red-teaming as an avenue for international cooperation on AI, followed by a discussion with the audience.
♦ Attendance is free but registration online is required.
Summary
Red-teaming has emerged as a best-practice among developers of frontier AI systems, especially large language models (LLMs). The voluntary commitments made by seven leading AI labs in partnership with the White House featured red-teaming as an essential tool. By having participants test a system adversarially, undesirable behaviours can be better identified and mitigated. Opening up this exercise to the international community could offer a compelling first step in establishing more formal form of international AI governance. A successful project would contribute to a shared appreciation of the risks and benefits these systems provide, and a platform to develop and disseminate best practices. However, there are a number of public and private interests at play that must be acknowledged, and many possible arrangements.
About Hadrien Pouget
Hadrien Pouget is an Associate Fellow in the Technology and International Affairs Program at the Carnegie Endowment for International Peace. He leverages his combined technical and policy backgrounds to analyze the course of AI regulation. He takes a particular interest in the technical and political challenges faced by those setting AI technical standards, which are set to underpin regulation. Previously, he worked as a research assistant at the computer science department at the University of Oxford. In this role, he published several papers on the testing and evaluation of machine learning systems.