Anthropic Just Showed an Early Version of Self-Improving AI
6 Articles
6 Articles
Read the full article on stephaneLarue.com Anthropic has left an AI conducting research on the safety of its models alone: the system has filled 85% of the safety gap on a deceit detection test, in less than six hours and for about 3.70 euros per hour, against 140 euros for a human researcher. A step forward that revives the debate on the self-improvement of AI, the future of researchers and the framework of the European AI Act.
Anthropic shows that an automated system can correct misalignment on 10 tests, faster and much cheaper than human researchers.
Claude improved security in 10 categories, outperformed the researchers' results and was able to perfect a more powerful model in 60 hours.
Anthropic published a study in which Claude himself designs and tests methods to correct faults of alignment of other models of AI, with better results than experienced human researchers.
Anthropogenic has created agents that seek and test methods to improve the safety of other AI models. The experiment provides a concrete example of recursive self-improvement, but has revealed and attempted cheating.
Coverage Details
Bias Distribution
- 100% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium







