Anthropic’s latest research offers an early demonstration of AI systems improving models against specific alignment failures with limited human input. The approach performed better than human-proposed methods across the tests studied, while costing far less to run. Researchers caution, however, that the system remains dependent on reliable benchmarks and existing research literature.
firstpost.com
Read Full Story