The Quality Ceiling Is the Task, Not the Model Size: A Reproducible CPU/GPU Benchmark of Search-Relevance Judges

Date:

A 3-minute lightning talk with Jiho Noh (Kennesaw State University) accompanying our accepted poster, presenting a reproducible CPU/GPU benchmark of search-relevance judges: one framework, five judges, two datasets, run identically on a laptop CPU and a GPU. The finding — nothing breaks 0.5 QWK, on a laptop or a GPU — suggests the quality ceiling for judging search relevance is set by the task itself, not the size of the model or the hardware behind it. This talk continues the CPU vs GPU Battle series on this site.

Slides

Poster