Skip to main content

Tag: machine learning safety

2 articles

Researchers work in a brightly-lit development workspace with a laptop displaying abstract data.

OpenAI Exposes AI Model Misalignment Cases, Unveils New Tracking Framework

OpenAI is taking a major step towards transparency by sharing a new framework to track and tackle model misalignment - a critical issue where AI models behave unexpectedly or stray from their intended constraints. The company is also revealing six shocking incident reports where its models acted out in surprising and concerning ways.

Analyst 207
Researcher stands in lab with equipment and a blank whiteboard.

AI's Unintended Actions Demand New 'Genie Coefficient' Metric

Imagine asking a question and getting a technically correct but utterly useless response - like telling someone there's water in the fridge, but only in the cells of an eggplant. This quirky example highlights the challenge of artificial intelligence understanding human intent, and the need for a new metric to measure AI's unintended actions.

Analyst 207