OpenAI发布了一个框架,用于追踪、调查和披露模型不对齐的情况,包括六个已记录的意外模型行为案例。

OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.