InterCoaching
FREN
AI news

An innovative test to identify advanced artificial intelligence

By Edouard2 min read
An innovative test to identify advanced artificial intelligence

The ARC Prize Foundation recently unveiled a revised version of its ARC-AGI test, designed to assess the progress of artificial general intelligence. Designed to be relatively simple for humans, this test tests the capabilities of AI models with visual puzzles that avoid brute-force answers. While humans achieve an average score of 60%, machines struggle to exceed 1%, highlighting the challenges of creating artificial intelligence capable of rivaling human intelligence. This test also emphasizes the efficiency and cost of skill acquisition by AIs, crucial elements in defining their intelligence. The ARC Prize Foundation recently unveiled a revised version of its ARC-AGI test to measure artificial general intelligence (AGI). Designed to challenge AIs in their assessment and effectiveness, this test uses visual puzzles that even models like OpenAI o3 find difficult. It offers an innovative approach to examining the potential of AI to match or surpass human intelligence, thus creating a platform for the continued exploration of this promising field. Development of the ARC-AGI Test **For several years, the ARC Prize Foundation has been developing a test capable of measuring artificial general intelligence. With this new ARC-AGI-2 test, the goal is to evaluate AIs using complex criteria that traditional algorithms such as chatbots cannot easily understand.**The first ARC-AGI test was launched early last year. While it established a baseline, its system had shortcomings. The OpenAI o3 model, for example, managed to achieve a score of 75.7%, indicating flaws that could be exploited by AI.

A New Approach with ARC-AGI-2 The new version, ARC-AGI-2, takes a different approach with visual puzzles rather than knowledge quizzes. This method aims to prevent AIs from relying on brute-force techniques to identify answers. It also seeks to assess the efficiency with which solutions are found, a crucial aspect according to the ARC Prize Foundation. Importance of Efficiency in AI Assessment**According to Greg Kamradt, co-founder of the ARC Prize Foundation, the notion of intelligence is not based solely on problem-solving or high scores. The efficiency with which an AI learns and deploys its capabilities to solve a task is equally important. The central question is whether an AI can acquire a skill and at what cost.**Comparison with Human Tests

In a sample of 400 humans, the average score obtained on the ARC-AGI-2 test was 60%. In contrast, most AI models fail to exceed 1%, with OpenAI o3 achieving only 4% on this new test. This clearly illustrates the significant difference between human learning and that of current AIs.

Competition and Financial Stakes **To stimulate innovation, the ARC Prize Foundation has announced a competition with a grand prize of $700,000. The AI ​​must achieve a score of 85% while keeping its running cost below $0.42 per task. This represents a considerable challenge, especially compared to OpenAI o3, which cost $200 per task for a score of only 4%.**The results of this competition, promising significant advances, will be announced on December 5, 2025. In the meantime, the ARC-AGI-2 test tasks are available to humans on the ARC Prize Foundation website, allowing for a direct comparison of problem-solving methods.

Notez cet article