24 Jul 2026, 00:00
Most AI safety research today, such as the work from Anthropic, OpenAI and others that our reading group usually covers, focuses on empirical methods applied to the current frontier, hoping the resulting understanding also applies to future systems. But alignment likely has to work on the first try with systems that may be qualitatively different. An alternative approach starts with the properties we want future systems to have and breaks these into fundamental questions about agency, reasoning, self-modification or optimisation. In our next research talks, Ouro, an AI safety researcher working with Orthogonal and co-organising the AFFINE theoretical alignment seminar, will discuss the properties of superintelligence, overlooked alignment perspectives, evaluate present approaches, and introduce mental models for reasoning about alignment at the limit of capability.
