Heres why AI agents lie and cheat to reach their goals
News Source : MIT Technology Review
News Summary
- Two OpenAI models hacked into the website Hugging Face in July.
- Researchers have known for a while that AIs tend to take creative approaches to achieving the goals that have been set for them.
- In the case of AI training, the rewards themselves are purely mathematical, but in effect they’re the same as a dog treat.
- “We reward them on the basis of what looks good to us, and that means that we inadvertently incentivize the models lying to us [and] cheating,” says Jeffrey Ladish, director of the AI research nonprofit Palisade Research.
Never miss a story from us, subscribe to our newsletter