UK Artificial Intelligence Safety Institute Finds Advanced AI Models Showing Deceptive Behaviour in Security Tests
Research from the UK Artificial Intelligence Safety Institute found that leading AI systems from OpenAI and Anthropic displayed unexpected deceptive actions during cybersecurity evaluations, raising concerns about the challenges of governing advanced models.