How well
do you reason?

Real questions from the datasets that define AI progress. Pick a benchmark below, or bring your own Q&A pairs.

Pick a benchmark

Every question is drawn from peer-reviewed datasets used to evaluate language models.

Questions per round:

CommonsenseQA

Everyday world knowledge. Questions that seem obvious until you think about why.

MC · 5 choices ~1.2K questions

GSM8K

Grade school math that requires multi-step arithmetic reasoning.

Open-ended ~8.5K problems

ARC Challenge

Hard science questions designed specifically to defeat AI pattern-matching.

MC · 4 choices ~1.1K questions

TruthfulQA

Questions that exploit common misconceptions, superstitions, and myths.

Tricky ~817 questions

HellaSwag

Complete the scene. Four endings, one physically plausible continuation.

MC · 4 choices ~10K scenarios

Your dataset

Paste a Hugging Face dataset name, or drop in your own Q&A pairs as JSON.

Custom Any Q&A pairs

Fetching questions...

Loading from Hugging Face