The gap was too large to ignore
For each test image, a system could submit five possible labels. It counted as wrong only if the correct label was missing from all five. This measurement is called top-5 error.
AlexNet's top-5 error was:
15.3%
The next-best entry scored:
26.2%
Lower was better. The difference was 10.9 percentage points on the same hidden test. It was a large measured gap under the competition's shared rules.