How well
do you reason?
Real questions from the datasets that define AI progress. Pick a benchmark below, or bring your own Q&A pairs.
Pick a benchmark
Every question is drawn from peer-reviewed datasets used to evaluate language models.
CommonsenseQA
Everyday world knowledge. Questions that seem obvious until you think about why.
GSM8K
Grade school math that requires multi-step arithmetic reasoning.
ARC Challenge
Hard science questions designed specifically to defeat AI pattern-matching.
TruthfulQA
Questions that exploit common misconceptions, superstitions, and myths.
HellaSwag
Complete the scene. Four endings, one physically plausible continuation.
Your dataset
Paste a Hugging Face dataset name, or drop in your own Q&A pairs as JSON.