The scores OpenAI did publish
The technical report described GPT-4 as a large multimodal model accepting text and images, leading with one number: a simulated Uniform Bar Exam score near the top 10 percent of test takers, up from GPT-3.5's bottom 10 percent. It also reported gains on MMLU (86.4 percent versus 70.0), HellaSwag, HumanEval and GSM-8K, plus SAT, GRE and LSAT scores. [2][3]
For the exam results, OpenAI said it ran a version of each test with any questions the model appeared to have seen during training removed, and reported the lower of the two scores. Image input shipped to exactly one outside partner at launch, Be My Eyes, for a visually-impaired-focused assistant, while ChatGPT Plus subscribers got text-only access and API developers joined a waitlist. [2][3]
What the report explicitly declined to say
A short section titled Scope and Limitations of this Technical Report states that the paper "contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar," citing "both the competitive landscape and the safety implications of large-scale models like GPT-4." No parameter count, no dataset description, no hardware list: a stated refusal, not an omission by oversight. [2]
OpenAI's own earlier GPT-2 paper published parameter counts for all four model sizes, up to 1.5 billion, and described the architecture in enough detail to reproduce it. The fight over that release was about whether to publish the trained weights, not whether to describe how the model worked. GPT-4's report withheld both the description and the weights. [5][2]
The one technical claim it did make
Atlas interpretation: The report's single specific technical claim was about predictability rather than architecture: OpenAI said it could reliably forecast some GPT-4 capabilities from smaller models trained on 1,000 to 10,000 times less compute, using loss curves fit before the final training run finished. That claim does different work than a benchmark score. It signals engineering discipline to peers and investors without handing a competitor a single number to copy. [2]
The bar exam number didn't hold up either
Two months later, attorney and researcher Eric Martinez re-ran the bar exam comparison using the scoring rubric and cohort data behind an actual Illinois bar administration. Measured against first-time test takers rather than the retake-heavy cohort OpenAI's estimate implied, GPT-4's score fell to roughly the 60th percentile overall and near the 40th on the essay section; measured only against people who went on to pass and practice law, it dropped further still. [4]
Atlas interpretation: The technical report was candid about withholding how GPT-4 was built. It was considerably less careful about the one number it chose to headline. This event's title, calling the report one that says nothing, undersells what OpenAI did disclose about its evaluation methodology; it slightly overstates how solid the single most-quoted result in that disclosure turned out to be. [2][4]
Sources
- GPT-4
OpenAI · Mar 14, 2023
- GPT-4 Technical Report
arXiv · Mar 15, 2023
- OpenAI releases GPT-4, an AI that it claims is state-of-the-art
TechCrunch · Mar 14, 2023
- Re-evaluating GPT-4's bar exam performance
SSRN · May 18, 2023
- Language Models are Unsupervised Multitask Learners
OpenAI · Feb 14, 2019