GPQA Diamond :一个困难的智力基准,用于测试化学、物理和生物学方面的专业知识。
Life can only be understood backward, but it must be lived forward - Søren Kierkegaard (Quiet-STaR 在论文的 Abstract 引用了这句话,当时觉得挺有意境的)
According to these evaluations, o1-preview hallucinates less frequently than GPT-4o, and o1-mini hallucinates less frequently than GPT-4o-mini. However, we have received anecdotal feedback that o1-preview and o1-mini tend to hallucinate more than GPT-4o and GPT-4o-mini. More work is needed to understand hallucinations holistically, particularly in domains not covered by our evaluations (e.g., chemistry). Additionally, red teamers have noted that o1-preview is more convincing in certain domains than GPT-4o given that it generates more detailed answers. This potentially increases the risk of people trusting and relying more on hallucinated generation.
| 欢迎光临 链载Ai (https://www.lianzai.com/) | Powered by Discuz! X3.5 |