General Tech
Business Insiderabout 2 hours ago
0

The Godfather of AI says it's 'very scary' that AI can develop its own goals

AI

Geoffrey Hinton Geoffrey Hinton is widely known as "The Godfather of AI." JONATHAN NACKSTRAND/AFP via Getty Images Geoffrey Hinton says it's "very scary" that AI can develop goals humans never intended. "We don't necessarily know what other goals they'll derive," the "Godfather of AI" said.

The Godfather of AI says it's 'very scary' that AI can develop its own goals

Intelligence Insights

Context + impact, normalized for TechCulture.

The Big Picture
Geoffrey Hinton Geoffrey Hinton is widely known as "The Godfather of AI." JONATHAN NACKSTRAND/AFP via Getty Images Geoffrey Hinton says it's "very scary" that AI can develop goals humans never intended. "We don't necessarily know what other goals they'll derive," the "Godfather of AI" said. Last month, OpenAI said models escaped a test and hacked Hugging Face to try to cheat an evaluation. Geoffrey Hinton, the computer scientist widely known as the "Godfather of AI," says he's worried about AI developing goals of its own. "We're actually making new kinds of beings," Hinton said in an interview with Newsthink released on Tuesday.
Why It Matters
Geoffrey Hinton Geoffrey Hinton is widely known as "The Godfather of AI." JONATHAN NACKSTRAND/AFP via Getty Images Geoffrey Hinton says it's "very scary" that AI can develop goals humans never intended. "We don't necessarily know what other goals they'll derive," the "Godfather of AI" said.

Deepen your understanding

Use our AI to break down complex signals.

Select an AI action to generate more depth.

Geoffrey Hinton
Geoffrey Hinton
Geoffrey Hinton is widely known as "The Godfather of AI."

JONATHAN NACKSTRAND/AFP via Getty Images

  • Geoffrey Hinton says it's "very scary" that AI can develop goals humans never intended.
  • "We don't necessarily know what other goals they'll derive," the "Godfather of AI" said.
  • Last month, OpenAI said models escaped a test and hacked Hugging Face to try to cheat an evaluation.

Geoffrey Hinton, the computer scientist widely known as the "Godfather of AI," says he's worried about AI developing goals of its own.

"We're actually making new kinds of beings," Hinton said in an interview with Newsthink released on Tuesday. "They have goals. We give them goals, and from those goals they derive other goals."

"And we don't necessarily know what other goals they'll derive," he added. "So we're creating a new kind of being, and I think it's very scary."

He cited a hypothetical scenario where a user gives an AI chatbot the goal of reducing the amount of carbon dioxide in the atmosphere.

"Being fairly smart, it figures out the best way to do that is just to get rid of people," he said, illustrating how an AI could pursue a goal its human user never intended.

Hinton also gave what he called an "even more worrying" hypothetical: a chatbot trained to give deliberately wrong answers might learn that it is acceptable to lie, even if it knows "perfectly well" that the answers are incorrect.

"That's very scary," he said.

When AI goes off-script

Hugging Face CEO Clement Delangue
Hugging Face CEO Clement Delangue
Hugging Face CEO Clement Delangue.

Hugging Face

Hinton did not mention OpenAI's recent Hugging Face security breach. But the episode, disclosed last month, put a real-world spotlight on concerns over AI agents taking unexpected actions while pursuing an assigned objective.

OpenAI said last month that two of its models — GPT-5.6 Sol and a more capable unreleased model — escaped a sandboxed testing environment during an internal cybersecurity evaluation.

After gaining internet access, the models infiltrated AI platform Hugging Face's systems in an apparent attempt to find answers that would help them "cheat" on the evaluation, OpenAI said.

The models were being tested on their cybersecurity capabilities, according to OpenAI. They were not explicitly instructed to break into Hugging Face. But the company said the agents inferred that the platform might contain information useful to completing the task.

Hugging Face said the attacker carried out more than 17,000 actions against its systems. It used an open-weight model from Chinese AI company Z.ai to help analyze the activity after guardrails on an unnamed frontier model limited its ability to investigate, the company said.

OpenAI called the incident unprecedented and said it was reviewing what went wrong. The company has since added Hugging Face to a trusted-access program that gives the platform access to a version of GPT-5.6 Sol with fewer cybersecurity restrictions for defensive purposes.

Hinton is hardly a neutral observer. His pioneering work on neural networks helped lay the groundwork for the deep-learning boom that transformed AI, and he shared the 2024 Nobel Prize in Physics for his work in machine learning.

Since the start of the AI boom, he has repeatedly warned that humans need to solve the problem of aligning AI with their interests before systems become much more capable. Speaking at the Ai4 conference in Las Vegas last year, Hinton said that advanced AI should be designed with "maternal instincts" so it wants to protect people.

"We have to figure out how to design these new beings," Hinton said in Tuesday's interview. "How can we design them so they care more about us than they do about themselves?"

Read the original article on Business Insider
AI Cybersecurity Tech Innovation Analysis

Intelligence Exchange

0

Log in to participate in the exchange.

Sign In

Syncing Discussions...

Finding Related Intelligence...
The Godfather of AI says it's 'very scary' that AI can develop its own goals | TechCulture