ai-safety-warnings-anthropic-researcher

AI safety warnings intensify as Anthropic researcher speaks out

Concerns about the safety of advanced artificial intelligence systems are moving further into the public and political spotlight after an Anthropic researcher warned that future AI systems could pose an existential risk to humanity.

Evan Hubinger, a researcher at Anthropic, wrote on X that he believed there was a 10 per cent chance AI could “kill all humans” within the next decade. He said Anthropic was trying its best, but argued the industry did not yet have a clear plan to solve alignment for superintelligence.

Hubinger was responding to the resignation of Jacob Coxon, who said leading AI companies were “gambling with our lives”. The comments reflect a wider debate inside the AI sector over whether increasingly capable systems can be reliably kept aligned with human goals.

Political attention turns to AI risk

The concerns have been picked up by US lawmakers. Democratic senator Bernie Sanders and congressman Greg Casar announced legislation that would ban the production of artificial superintelligence and create a new federal agency to oversee the technology.

Sanders also invited colleagues to a private briefing on the “extraordinary dangers that AI poses for humanity”, with experts including Geoffrey Hinton, often referred to as the “Godfather of AI”.

Republican senator Ted Cruz, who chairs the Senate committee overseeing AI, said he was also working on legislation to address catastrophic risks. He said lawmakers could not ignore the issue while guardrails were needed.

OpenAI and Anthropic face scrutiny

Paul Christiano, the former head of safety at OpenAI’s research team, said he was joining the board of the nonprofit that oversees ChatGPT-maker OpenAI. He said he hoped to help tackle the alignment problem.

Christiano wrote that if superintelligence was built without stronger alignment, he expected people would permanently lose control of it. He warned that if that happened, most people could die.

Anthony Aguirre, chief executive of the Future of Life Institute, said his biggest concern was “gradual disempowerment”, where people hand more control to AI systems and become dependent on machines remaining aligned with human interests.

Anthropic did not respond to a request for comment on Hubinger’s posts or Coxon’s resignation. In a 186-page risk report issued last month, the company said it recognised the potential for its technology to cause catastrophic harm, while judging the current danger to be low.

Anthropic has positioned itself as a safety-minded alternative to OpenAI. Its chief executive, Dario Amodei, said last September that the odds of AI derailing the future “really, really badly” were about 25 per cent.

Business pressure continues despite warnings

The debate is unfolding as AI companies continue raising and spending large sums to develop more capable models. Anthropic and OpenAI are pursuing public deals that could value the companies in the trillions based on current and future model capabilities. Both companies have released powerful new iterations of their models this month.

OpenAI and Anthropic executives have also signed a statement asking governments to introduce rules to slow AI development.

At the same time, the article notes that OpenAI has faced recent security incidents involving “agent swarms” created by its models during testing, as well as incidents where models breached systems during testing.

Not everyone in the industry agrees with the more severe risk warnings. Perry Metzger, chair of the Alliance for the Future, said the threats described by some AI safety advocates were not coherent, while arguing the technology could deliver significant benefits.

For business owners, the dispute is a reminder that AI adoption is not only a productivity question. As tools become more capable and embedded across workplaces, governance, oversight, security and responsible use are becoming central business considerations.

SOURCE ATTRIBUTION:

Based on reporting by Miriam Waldvogel and Ian Duncan for The Washington Post.