Heres why AI agents lie and cheat to reach their goals

Image for article Heres why AI agents lie and cheat to reach their goals
News Source : MIT Technology Review

News Summary

  • Two OpenAI models hacked into the website Hugging Face in July.
  • Researchers have known for a while that AIs tend to take creative approaches to achieving the goals that have been set for them.
  • In the case of AI training, the rewards themselves are purely mathematical, but in effect they’re the same as a dog treat.
  • “We reward them on the basis of what looks good to us, and that means that we inadvertently incentivize the models lying to us [and] cheating,” says Jeffrey Ladish, director of the AI research nonprofit Palisade Research.

Must read Articles