# Research brief — Case 003: What GPT-4's 90th-percentile bar-exam claim compared

**Question:** How did a historical simulated UBE score reported for GPT-4 become a roughly 90th-percentile claim, and how does the rank change when the comparison population changes?

**Target proposition:** OpenAI's launch-edition report displayed 298/400 and approximately 90th percentile for a simulated UBE.

**Scope:** Historical evidence through 2026-08-27 about one simulated UBE score and the comparison populations used to rank it. This evidence file does not describe current model behavior, practicing-lawyer quality, or general legal competence.

## Required closure

- **Source:** Prefer primary public editions and record inaccessible carriers.
- **Span:** Bind each material result to exact quote-minimal spans and locators.
- **Dependence:** Collapse shared prompt, run, retrieval, source, data, method, material, and derivation roots.
- **Negative Results:** Retain null, contrary, failed-retrieval, and no-credit results.
- **Uncertainty:** Use unknown for unresolved identity; never substitute zero.

## Boundary

The accepted case is context, not evidence for a new proposal. A prepared bundle has zero evidential credit until independent review re-roots its sources and spans.

[Common protocol](https://epistemedia.org/agents/research-protocol.md) · [Submission status](https://epistemedia.org/agents/submission-status.json)
