How accurate are your AI outputs?
OpenAI warns us that the outputs generated by their models may not be accurate in their “Terms and Conditions.” Their “Terms and Conditions” also indicate that we should not rely on their AI outputs as a sole source of truth or factual information. That is excellent advice!
https://openai.com/policies/terms-of-use
Hallucination Evaluation by OpenAI
On April 16, 2025, OpenAI published a report on its own internal evaluation of hallucinations by OpenAI’s o3 and o4-mini models in comparison to its earlier o1 model. In addition to some hard data the report also provides a brief overview of the very complex nature of AI hallucinations.
https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf
SimpleQA
The first of the two datasets used in the evaluation, named “SimpleQA,” included 4,326 concise questions across many areas. Each question had only one correct answer. See Table 1 below.
In the SimpleQA test the older o1 model was accurate 47% of the time and hallucinated 44% of the time. The o1 model declined to respond for 9% of the questions. In comparison the o3 model was accurate 49% of the time and hallucinated 51% of the time. This “new and improved” o3 model responded more often but hallucinated more often in its output. The o4-mini model, on the other hand, was accurate only 20% of the time and hallucinated an incredible 79% of the time.
So, OpenAI was spot on – We absolutely should not rely on their AI outputs as a sole source of truth or factual information!
PersonQA
The other dataset, named “PersonQA,” asked well-publicized factual questions about public figures. See Table 2 below.
The o1 model was accurate 47% of the time and hallucinated only 16% of the time. But that left 37% of questions that it declined to answer. In comparison the o3 model was accurate 59% of the time and hallucinated 33% of the time. The o4-mini model was accurate 36% of the time and hallucinated 48% of the time.
The parting thought: “Don’t trust and always verify the AI output.”


