How We Know · Case 003
Independent review receipt
What GPT-4's 90th-percentile bar-exam claim compared
Decision
pass
A separate Codex review agent used a fresh clone, did not author the candidate, and checked exact source, calculation, lineage, receipt, repository, and deterministic-build closure. The review did not decide that the bounded scientific claim is universally true.
Review identity
- Technical reviewer ID
- codex-independent-em0034-reviewer
- Reviewed author head
- 1c5a8318b2bb770c38a07a96f731531ba1c53a32
- Dossier
- em:dossier:sha256:babe89ba3bda594a8d9f2db86a5a2987f284437a069b940d19b6928856d936d1
- Receipt SHA-256
- f35f8c093778ad1cdeafc57746575f1a82d12fe6e6fbf0d1d4a546e3c5296e1e
- Completed
- 2026-08-28T15:59:23Z
What was checked
- Exact accepted packet, dossier, and review-receipt bytes
- Source, edition, span, calculation, and license closure
- Count grammar, unresolved items, and typed dependence edges
- Policy divergence, adversarial validation, and deterministic build
What this did not decide
- This review verifies the bounded historical candidate and does not establish current model behavior, general legal competence, practicing-lawyer quality, or a current product ranking.
- The launch comparison distribution, original top-ten-percent chart, national UBE distribution, and proprietary inputs remain unresolved and receive no invented value or evidentiary credit.
- The Martinez re-analysis remains model- and assumption-dependent, including its retained 45/48 internal discrepancy and its separate comparison populations.
- This receipt does not admit, feature, publish, deploy, or describe Case 003 as live.
This review checked the bounded packet and derivation. It did not decide whether every empirical proposition is universally or currently true.