Hugging Face Datasets: The Complete Directory & Live Ranking
A live directory of Hugging Face Hub datasets, ranked by real adoption, growth, and maintenance - not opinion.
What Is a Hugging Face Dataset, and How Do You Use One?
What it is
A Hugging Face dataset is a published collection of data (text, images, audio, or tabular) hosted on the Hugging Face Hub, along with a dataset card documenting its schema, splits, and license. Datasets are used to train, fine-tune, and evaluate AI models.
How to use a Hugging Face dataset
Every dataset on this list can be loaded directly with the datasets library (load_dataset("org/dataset-id")), which handles download, caching, and format conversion automatically. Click a dataset's name to open its official Hugging Face page, where the dataset card documents the exact schema, splits, and any license terms to accept.
As of August 4, 2026, this is a live, free directory ranking the best and most-used datasets on the Hugging Face Hub by real data. zplatform.ai tracks datasets that clear a minimum adoption floor (150 ranked today) using the free, public Hugging Face Hub API. Ranked by downloads, likes, and how recently each dataset was updated - not opinion - and the list refreshes automatically every week. How we choose and rank →
The best Hugging Face datasets as of August 4, 2026, ranked by our composite score:
- assets - 764.4k downloads, quality score 90/100.
- HPLT2.0_cleaned - 1.6M downloads, quality score 90/100.
- samples - 577.0k downloads, quality score 90/100.
- seedance-2-prompts-datasets - 337.0k downloads, quality score 89/100.
- PhysicalAI-Robotics-GR00T-X-Embodiment-Sim - 12.6M downloads, quality score 89/100.
Every Ranked Dataset
Every Hugging Face dataset that clears our adoption floor, scored on a transparent composite of adoption, growth, maintenance, and trust. Select up to 4 to compare side by side.
| Compare | # | Dataset | Category | Downloads | Likes | License | Updated | Quality |
|---|---|---|---|---|---|---|---|---|
| 1 | assets | Other | 764.4k | 3 | mit | 28d ago | 90/100 | |
| ||||||||
| 2 | HPLT2.0_cleaned | Text & Language Data | 1.6M | 44 | cc0-1.0 | 2mo ago | 90/100 | |
| ||||||||
| 3 | samples | Other | 577.0k | 9 | cc0-1.0 | 9d ago | 90/100 | |
| ||||||||
| 4 | seedance-2-prompts-datasets | Vision & Video | 337.0k | 17 | cc-by-4.0 | 2d ago | 89/100 | |
| ||||||||
| 5 | PhysicalAI-Robotics-GR00T-X-Embodiment-Sim | Robotics & Simulation | 12.6M | 253 | cc-by-4.0 | 5mo ago | 89/100 | |
| ||||||||
| 6 | OpenFake | Vision & Video | 190.7k | 30 | cc-by-nc-4.0 | yesterday | 88/100 | |
| ||||||||
| 7 | prompts.chat | Question Answering & Reasoning | 654.6k | 9.8k | cc0-1.0 | today | 88/100 | |
| ||||||||
| 8 | WaxalNLP | Speech & Audio | 210.4k | 264 | cc-by-sa-4.0cc-by-4.0 | 2d ago | 88/100 | |
| ||||||||
| 9 | LLaVA-OneVision-1.5-Mid-Training-85M | Other | 3.5M | 86 | apache-2.0 | 19d ago | 88/100 | |
| ||||||||
| 10 | Ultra-FineWeb | Text & Language Data | 770.4k | 406 | apache-2.0 | 2mo ago | 88/100 | |
| ||||||||
| 11 | yahoo-finance-data | Other | 737.9k | 106 | odc-by | today | 87/100 | |
| ||||||||
| 12 | stack-v3-train | Text & Language Data | 136.2k | 296 | odc-by | yesterday | 87/100 | |
| ||||||||
| 13 | American-Sign-Language-Dataset | Other | 575.9k | 2 | mit | 2mo ago | 87/100 | |
| ||||||||
| 14 | video-vec2wav2-tokenizer | Other | 2.4M | 14 | - | 13d ago | 87/100 | |
| ||||||||
| 15 | RoboDojo | Other | 160.5k | 6 | apache-2.0 | 9d ago | 87/100 | |
| ||||||||
| 16 | L2D | Robotics & Simulation | 437.6k | 49 | apache-2.0 | 2mo ago | 87/100 | |
| ||||||||
| 17 | 2026-challenge-demos | Other | 111.3k | 3 | mit | 7d ago | 86/100 | |
| ||||||||
| 18 | community-pipelines-mirror | Other | 1.5M | 9 | - | 27d ago | 86/100 | |
| ||||||||
| 19 | PhysicalAI-Robotics-Open-H-Embodiment | Robotics & Simulation | 351.5k | 41 | cc-by-4.0 | 1mo ago | 86/100 | |
| ||||||||
| 20 | humanoid-everyday | Robotics & Simulation | 289.2k | 43 | apache-2.0 | 2mo ago | 86/100 | |
| ||||||||
| 21 | zoengjyutgaai | Speech & Audio | 218.3k | 0 | cc0-1.0 | 3mo ago | 85/100 | |
| ||||||||
| 22 | opencs2_dataset | Vision & Video | 157.4k | 33 | cc-by-4.0 | 1mo ago | 85/100 | |
| ||||||||
| 23 | MultiCities | Vision & Video | 169.3k | 0 | other | 3mo ago | 85/100 | |
| ||||||||
| 24 | fleurs | Speech & Audio | 1.7M | 438 | cc-by-4.0 | 3mo ago | 85/100 | |
| ||||||||
| 25 | dom-pi-pdfs-2025 | Other | 180.6k | 0 | cc-by-4.0 | 2mo ago | 85/100 | |
| ||||||||
| 26 | Maritime_Visual_Tracking_Dataset_MVTD | Other | 203.7k | 0 | cc0-1.0 | 3mo ago | 85/100 | |
| ||||||||
| 27 | monet | Vision & Video | 716.4k | 144 | apache-2.0 | 1mo ago | 85/100 | |
| ||||||||
| 28 | license-plates-700k | Other | 189.8k | 0 | mit | 2mo ago | 85/100 | |
| ||||||||
| 29 | VideoChat3-LV116k | Vision & Video | 68.3k | 15 | apache-2.0 | 16d ago | 85/100 | |
| ||||||||
| 30 | OpenAlex | Tabular & Time-Series | 203.4k | 5 | cc0-1.0 | 2mo ago | 85/100 | |
| ||||||||
| 31 | DoseRAD2026 | Other | 197.9k | 0 | cc-by-nc-4.0 | 3mo ago | 85/100 | |
| ||||||||
| 32 | PhysicalAI-Robotics-Locomanipulation-GRAIL | Other | 74k | 23 | apache-2.0 | 19d ago | 85/100 | |
| ||||||||
| 33 | ProgramBench-Tests | Text & Language Data | 231.3k | 10 | mit | 3mo ago | 85/100 | |
| ||||||||
| 34 | knowledge-base | Other | 73.8k | 13 | cc-by-4.0 | today | 85/100 | |
| ||||||||
| 35 | infini-news-corpus | Text & Language Data | 183k | 31 | cc-by-4.0 | 1mo ago | 85/100 | |
| ||||||||
| 36 | ubuntu_osworld_file_cache | Other | 7.2M | 43 | apache-2.0 | 3mo ago | 85/100 | |
| ||||||||
| 37 | JL1-CUP-2024-Second-Format | Vision & Video | 104.6k | 1 | other | 2mo ago | 84/100 | |
| ||||||||
| 38 | dave_sonar | Other | 137.9k | 0 | cc-by-4.0 | 3mo ago | 84/100 | |
| ||||||||
| 39 | papas-nativas-peru-83-variedades | Vision & Video | 142.2k | 0 | cc-by-nc-4.0 | 3mo ago | 84/100 | |
| ||||||||
| 40 | X2C | Vision & Video | 127.9k | 0 | cc-by-4.0 | 2mo ago | 84/100 | |
| ||||||||
| 41 | figofigofigofigo | Other | 1.3M | 22 | - | 3mo ago | 84/100 | |
| ||||||||
| 42 | DogSpeak_Dataset | Other | 142k | 0 | cc-by-nc-sa-4.0 | 3mo ago | 84/100 | |
| ||||||||
| 43 | OpenAoE-2000h | Other | 52.7k | 17 | other | today | 84/100 | |
| ||||||||
| 44 | ETCI-2021-Flood-Detection | Vision & Video | 152.5k | 0 | unknown | 3mo ago | 84/100 | |
| ||||||||
| 45 | USDT-M_Perpetual_Futures | Tabular & Time-Series | 45.5k | 3 | mit | 12d ago | 84/100 | |
| ||||||||
| 46 | 4D-Lung | Vision & Video | 148.5k | 1 | cc-by-3.0 | 1mo ago | 84/100 | |
| ||||||||
| 47 | LA-33K | Other | 47.3k | 1 | mit | 27d ago | 84/100 | |
| ||||||||
| 48 | LLaVA-OneVision-1.5-Instruct-Data | Other | 1.3M | 77 | apache-2.0 | 19d ago | 84/100 | |
| ||||||||
| 49 | PhysicalAI-Robotics-GR00T-X-Embodiment-Sim | Robotics & Simulation | 179.7k | 0 | cc-by-4.0 | 2mo ago | 84/100 | |
| ||||||||
| 50 | openwebtext | Text & Language Data | 5.5M | 528 | cc0-1.0 | 13d ago | 84/100 | |
| ||||||||
| 51 | SWE-bench_Multilingual | Other | 1.2M | 21 | mit | 13d ago | 84/100 | |
| ||||||||
| 52 | openpath-corpus | Vision & Video | 43.4k | 1 | other | 25d ago | 84/100 | |
| ||||||||
| 53 | osworld_v2_assets | Other | 493.3k | 2 | - | 12d ago | 84/100 | |
| ||||||||
| 54 | flatpak | Other | 889k | 1 | - | today | 83/100 | |
| ||||||||
| 55 | dartlab-data | Question Answering & Reasoning | 215.3k | 11 | cc-by-4.0 | today | 83/100 | |
| ||||||||
| 56 | challenge | Other | 44.5k | 1 | mit | today | 83/100 | |
| ||||||||
| 57 | OmniWorld | Vision & Video | 624.5k | 95 | cc-by-nc-sa-4.0 | 4mo ago | 83/100 | |
| ||||||||
| 58 | kuka_lerobot | Robotics & Simulation | 131.9k | 0 | apache-2.0 | 2mo ago | 83/100 | |
| ||||||||
| 59 | prophet-mosque-library | Vision & Video | 808.9k | 2 | mit | 6mo ago | 83/100 | |
| ||||||||
| 60 | MMMU | Question Answering & Reasoning | 2.4M | 331 | apache-2.0 | 24d ago | 83/100 | |
| ||||||||
| 61 | SWE-rebench | Other | 12.4M | 69 | cc-by-4.0 | 7mo ago | 83/100 | |
| ||||||||
| 62 | PhysicalAI-Autonomous-Vehicles | Other | 2.8M | 965 | other | 3mo ago | 83/100 | |
| ||||||||
| 63 | SAGE-10k | Other | 791k | 80 | apache-2.0 | 6mo ago | 83/100 | |
| ||||||||
| 64 | MacroLens | Tabular & Time-Series | 137.1k | 0 | cc-by-4.0 | 2mo ago | 82/100 | |
| ||||||||
| 65 | Berkeley-Function-Calling-Leaderboard | Other | 385k | 117 | apache-2.0 | 3mo ago | 82/100 | |
| ||||||||
| 66 | OpenVid-1M | Vision & Video | 902.6k | 276 | cc-by-4.0 | 4mo ago | 82/100 | |
| ||||||||
| 67 | AgiBotWorld2026 | Robotics & Simulation | 151.2k | 53 | cc-by-nc-sa-4.0 | 13d ago | 81/100 | |
| ||||||||
| 68 | hd_tmp | Other | 6.3M | 29 | - | 1mo ago | 81/100 | |
| ||||||||
| 69 | ProObjaverse-300K | Vision & Video | 278.9k | 0 | apache-2.0 | 3mo ago | 81/100 | |
| ||||||||
| 70 | Submission-Archive | Other | 874.5k | 0 | - | today | 81/100 | |
| ||||||||
| 71 | CS2-10k | Other | 140.6k | 35 | cc-by-nc-4.0 | 1mo ago | 81/100 | |
| ||||||||
| 72 | xperience-10m | Vision & Video | 2.7M | 222 | other | 4mo ago | 81/100 | |
| ||||||||
| 73 | gaia | Other | 1.7M | 14 | gpl-3.0 | 5mo ago | 81/100 | |
| ||||||||
| 74 | MMLU-Pro | Question Answering & Reasoning | 2.0M | 508 | mit | 3mo ago | 81/100 | |
| ||||||||
| 75 | 2.1tbofdata | Robotics & Simulation | 295.5k | 0 | cc-by-nc-4.0 | 4mo ago | 81/100 | |
| ||||||||
| 76 | DDOS | Vision & Video | 245.6k | 0 | cc-by-nc-4.0 | 4mo ago | 81/100 | |
| ||||||||
| 77 | financial-lakehouse-test | Other | 235.8k | 0 | other | 3mo ago | 81/100 | |
| ||||||||
| 78 | dolma3_dolmino_mix-100B-1125 | Other | 954k | 23 | odc-by | 5mo ago | 80/100 | |
| ||||||||
| 79 | molmospaces | Other | 118.4k | 50 | odc-bycc-by-4.0 | 2mo ago | 80/100 | |
| ||||||||
| 80 | CC-Bench | Other | 210.1k | 0 | other | 3mo ago | 80/100 | |
| ||||||||
| 81 | SwissCube | Other | 127.9k | 0 | mit | 3mo ago | 80/100 | |
| ||||||||
| 82 | images | Other | 493.2k | 0 | other | today | 80/100 | |
| ||||||||
| 83 | DOM | Robotics & Simulation | 201.9k | 12 | other | 3mo ago | 80/100 | |
| ||||||||
| 84 | aime_2025 | Other | 237.9k | 17 | cc-by-nc-sa-4.0 | 3mo ago | 80/100 | |
| ||||||||
| 85 | PhysicalAI-SmartSpaces | Other | 709.5k | 84 | cc-by-4.0 | 7d ago | 80/100 | |
| ||||||||
| 86 | gsm8k | Text & Language Data | 14.1M | 1.5k | mit | 4mo ago | 80/100 | |
| ||||||||
| 87 | sat-image-boundingbox-sft-full | Other | 176.4k | 1 | apache-2.0 | 3mo ago | 80/100 | |
| ||||||||
| 88 | pile-of-law | Text & Language Data | 525.0k | 278 | cc-by-nc-sa-4.0 | 17d ago | 80/100 | |
| ||||||||
| 89 | SurgBench_NIPS25 | Other | 122.4k | 0 | apache-2.0 | 2mo ago | 80/100 | |
| ||||||||
| 90 | American-Sign-Language-Dataset | Other | 207.9k | 0 | mit | 4mo ago | 80/100 | |
| ||||||||
| 91 | MDC | Other | 552.2k | 0 | - | today | 80/100 | |
| ||||||||
| 92 | Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting | Other | 190.7k | 0 | mit | 3mo ago | 80/100 | |
| ||||||||
| 93 | wikipedia-monthly | Text & Language Data | 163.8k | 0 | cc-by-sa-4.0 | 3mo ago | 80/100 | |
| ||||||||
| 94 | certificates | Other | 384.8k | 88 | apache-2.0 | today | 79/100 | |
| ||||||||
| 95 | IndoLepAtlas | Vision & Video | 76.8k | 3 | cc-by-nc-sa-4.0 | 17d ago | 79/100 | |
| ||||||||
| 96 | MR-RATE | Vision & Video | 392.0k | 94 | cc-by-nc-sa-4.0 | 21d ago | 79/100 | |
| ||||||||
| 97 | KakologArchives | Text & Language Data | 29.5M | 72 | mit | today | 79/100 | |
| ||||||||
| 98 | results | Other | 2.6M | 15 | - | yesterday | 79/100 | |
| ||||||||
| 99 | DogSpeak_Dataset | Other | 133.2k | 0 | cc-by-nc-sa-4.0 | 3mo ago | 79/100 | |
| ||||||||
| 100 | dolma3.5_pool | Text & Language Data | 41.5k | 5 | odc-by | 25d ago | 78/100 | |
| ||||||||
| 101 | WeatherSynthetic | Vision & Video | 158.0k | 0 | cc-by-nc-4.0 | 4mo ago | 78/100 | |
| ||||||||
| 102 | text-to-image-2M | Vision & Video | 196k | 172 | mit | 7d ago | 78/100 | |
| ||||||||
| 103 | sts12-sts | Text & Language Data | 1.3M | 8 | unknown | 5mo ago | 78/100 | |
| ||||||||
| 104 | picbreeder-vlm-archive | Vision & Video | 205.6k | 0 | cc-by-nc-4.0 | 24d ago | 78/100 | |
| ||||||||
| 105 | SynData | Other | 563.6k | 185 | cc-by-4.0 | 2mo ago | 78/100 | |
| ||||||||
| 106 | molmobot-data | Other | 134.7k | 7 | odc-by | 11d ago | 77/100 | |
| ||||||||
| 107 | fineweb-tokenized | Text & Language Data | 8.8M | 27 | odc-by | 2mo ago | 77/100 | |
| ||||||||
| 108 | flores | Text & Language Data | 4.3M | 109 | cc-by-sa-4.0 | 2mo ago | 77/100 | |
| ||||||||
| 109 | ReActor | Other | 3.3M | 303 | mit | 3mo ago | 77/100 | |
| ||||||||
| 110 | finephrase | Text & Language Data | 2.1M | 138 | odc-by | 4mo ago | 77/100 | |
| ||||||||
| 111 | DSERT-RoLL | Other | 63.3k | 2 | cc-by-4.0 | 29d ago | 77/100 | |
| ||||||||
| 112 | RMISC | Tabular & Time-Series | 50.8k | 6 | mit | 21d ago | 77/100 | |
| ||||||||
| 113 | hacker-news | Text & Language Data | 135.5k | 338 | odc-by | today | 77/100 | |
| ||||||||
| 114 | Syn4D | Other | 60k | 17 | cc-by-4.0 | 4d ago | 77/100 | |
| ||||||||
| 115 | tartanair2 | Other | 181.1k | 9 | bsd-3-clause | 2mo ago | 77/100 | |
| ||||||||
| 116 | Multi-SWE-bench | Text & Language Data | 91.2k | 42 | other | 27d ago | 76/100 | |
| ||||||||
| 117 | wmt24pp | Text & Language Data | 89.2k | 90 | apache-2.0 | 4d ago | 76/100 | |
| ||||||||
| 118 | OpenMind | Vision & Video | 97.1k | 0 | cc-by-4.0 | 11d ago | 76/100 | |
| ||||||||
| 119 | leaderboard-dataset | Other | 101.3k | 22 | cc-by-4.0 | today | 76/100 | |
| ||||||||
| 120 | MathVision | Question Answering & Reasoning | 308.9k | 169 | mit | 2mo ago | 76/100 | |
| ||||||||
| 121 | dcvlm-baseline-200b | Other | 79.7k | 3 | other | 9d ago | 76/100 | |
| ||||||||
| 122 | Nemotron-CC-v2 | Text & Language Data | 621.4k | 133 | other | 28d ago | 76/100 | |
| ||||||||
| 123 | product-database | Other | 118.6k | 137 | agpl-3.0odbl | yesterday | 76/100 | |
| ||||||||
| 124 | dataset | Vision & Video | 137.8k | 9 | other | yesterday | 76/100 | |
| ||||||||
| 125 | SYNTH | Text & Language Data | 319.8k | 272 | cc-by-4.0 | 3mo ago | 76/100 | |
| ||||||||
| 126 | TexVerse-1K | Other | 811.5k | 7 | odc-by | 1mo ago | 76/100 | |
| ||||||||
| 127 | TexVerse-Skeleton-Animation | Other | 192.2k | 5 | odc-by | 1mo ago | 76/100 | |
| ||||||||
| 128 | EconomicIndex | Other | 195.0k | 559 | mit | 1mo ago | 75/100 | |
| ||||||||
| 129 | the-stack | Text & Language Data | 446.5k | 1k | other | yesterday | 75/100 | |
| ||||||||
| 130 | deform360 | Robotics & Simulation | 58.3k | 9 | mit | 8d ago | 75/100 | |
| ||||||||
| 131 | IndEgo | Other | 199.7k | 7 | cc-by-4.0 | 1mo ago | 75/100 | |
| ||||||||
| 132 | egoexo4d-hf | Other | 63.2k | 2 | mit | 27d ago | 75/100 | |
| ||||||||
| 133 | legalbench | Text & Language Data | 1.7M | 187 | cc-by-4.0 | 4mo ago | 75/100 | |
| ||||||||
| 134 | gpic | Other | 401.9k | 148 | mit | 4d ago | 75/100 | |
| ||||||||
| 135 | TexVerse | Other | 2.1M | 47 | odc-by | 1mo ago | 75/100 | |
| ||||||||
| 136 | neuripsED_2701 | Vision & Video | 48.3k | 0 | cc-by-4.0 | 7d ago | 74/100 | |
| ||||||||
| 137 | HIW-500 | Other | 158.6k | 47 | cc-by-4.0 | 1mo ago | 74/100 | |
| ||||||||
| 138 | cs2_dataset_render | Vision & Video | 77.5k | 4 | cc-by-4.0 | 3mo ago | 74/100 | |
| ||||||||
| 139 | WAM_Psi_Egoverse_dataset | Other | 103.5k | 0 | apache-2.0 | 3mo ago | 74/100 | |
| ||||||||
| 140 | dartlab-data | Question Answering & Reasoning | 72.0k | 0 | apache-2.0 | 2mo ago | 74/100 | |
| ||||||||
| 141 | gpt-image-2-prompts-datasets | Vision & Video | 42.3k | 2 | cc-by-4.0 | yesterday | 74/100 | |
| ||||||||
| 142 | fineweb-edu-translated | Text & Language Data | 2.1M | 17 | odc-by | 3mo ago | 74/100 | |
| ||||||||
| 143 | Scientific-Summaries | Other | 117.5k | 5 | cc-by-4.0 | 3mo ago | 74/100 | |
| ||||||||
| 144 | MNBVC | Text & Language Data | 1.4M | 647 | mit | 4mo ago | 74/100 | |
| ||||||||
| 145 | SWE-rebench-V2 | Text & Language Data | 107.0k | 56 | cc-by-4.0 | 3mo ago | 74/100 | |
| ||||||||
| 146 | vacuum_peptides | Other | 43.8k | 4 | cc-by-4.0 | 12d ago | 74/100 | |
| ||||||||
| 147 | SenseNova-Vision-Corpus-50M | Other | 45.4k | 49 | cc-by-nc-4.0 | 20d ago | 74/100 | |
| ||||||||
| 148 | American-Sign-Language-Dataset | Other | 40.3k | 0 | mit | 10d ago | 74/100 | |
| ||||||||
| 149 | qb-audio | Other | 41.8k | 2 | other | 4d ago | 74/100 | |
| ||||||||
| 150 | mediach | Vision & Video | 40.7k | 0 | apache-2.0 | 11d ago | 74/100 | |
| ||||||||
No datasets match your search.
Comparison
Best Hugging Face Datasets by Use Case
Best Speech & Audio Datasets
WaxalNLP and zoengjyutgaai lead speech and audio datasets, used to train transcription and audio-classification models.
- WaxalNLP 88/100 quality
- zoengjyutgaai 85/100 quality
- fleurs 85/100 quality
Best Robotics & Simulation Datasets
PhysicalAI-Robotics-GR00T-X-Embodiment-Sim and L2D lead robotics and simulation datasets, a fast-growing category behind today's agentic and embodied-AI benchmarks.
- PhysicalAI-Robotics-GR00T-X-Embodiment-Sim 89/100 quality
- L2D 87/100 quality
- PhysicalAI-Robotics-Open-H-Embodiment 86/100 quality
Best Text & Language Datasets
HPLT2.0_cleaned and Ultra-FineWeb lead text and language datasets, covering generation, classification, translation, and embedding training data.
- HPLT2.0_cleaned 90/100 quality
- Ultra-FineWeb 88/100 quality
- stack-v3-train 87/100 quality
Best Vision & Video Datasets
seedance-2-prompts-datasets and OpenFake top vision and video datasets, spanning image generation, classification, and multi-frame understanding.
- seedance-2-prompts-datasets 89/100 quality
- OpenFake 88/100 quality
- opencs2_dataset 85/100 quality
Best Tabular & Time-Series Datasets
OpenAlex and USDT-M_Perpetual_Futures top tabular and time-series datasets.
- OpenAlex 85/100 quality
- USDT-M_Perpetual_Futures 84/100 quality
- MacroLens 82/100 quality
Best Question Answering & Reasoning Datasets
prompts.chat and dartlab-data top question-answering and reasoning benchmarks.
- prompts.chat 88/100 quality
- dartlab-data 83/100 quality
- MMMU 83/100 quality
Browse by Category
Question Answering & Reasoning
6 datasets · 5.7M downloads
Filter the directory to Question Answering & Reasoning ↓State of the Datasets Ecosystem
Most Popular
- KakologArchives 29.5M
- gsm8k 14.1M
- PhysicalAI-Robotics-GR00T-X-Embodiment-Sim 12.6M
- SWE-rebench 12.4M
- fineweb-tokenized 8.8M
- ubuntu_osworld_file_cache 7.2M
- hd_tmp 6.3M
- openwebtext 5.5M
- flores 4.3M
- LLaVA-OneVision-1.5-Mid-Training-85M 3.5M
Fastest Growing (30-day downloads trend)
- video-vec2wav2-tokenizer +321.3%
- RoboDojo +190.1%
- 2026-challenge-demos +176.2%
- SWE-rebench +156.3%
- 4D-Lung +147.3%
- opencs2_dataset +139.7%
- stack-v3-train +132.3%
- OpenAlex +121.3%
- Berkeley-Function-Calling-Leaderboard +117.5%
- SAGE-10k +108.9%
By Category
- Other 73 · 65.9M
- Vision & Video 29 · 9.4M
- Text & Language Data 25 · 76.3M
- Robotics & Simulation 10 · 14.7M
- Question Answering & Reasoning 6 · 5.7M
- Tabular & Time-Series 4 · 436.9k
- Speech & Audio 3 · 2.1M
How we rank, and why you can trust it
- Data source: the free, public Hugging Face Hub API (downloads, likes, license, and last-modified date for every dataset). Refreshed weekly.
- Inclusion is automatic: a dataset must clear an adoption floor - at least 1,000 downloads in the last 30 days, or at least 50 likes - so the list stays credible rather than padded with unused uploads. We do not hand-pick winners.
- Adoption is measured by all-time downloads (when available) and likes, log-scaled so a handful of enormous outliers do not flatten the rest of the ranking.
- Growth is measured by the 30-day change in downloads, compared against our own weekly snapshots - shown once enough history has accrued.
- Maintenance reflects how recently a dataset was updated - a dataset updated in the last 30 days scores highest, decaying the longer it has gone untouched.
- Ranking is data-driven only. No dataset pays to be listed or ranked.
The Quality Score formula
The Quality Score (0-100) is a transparent weighted blend: 40% Adoption (downloads and likes, log-scaled), 20% Growth (30-day downloads trend), 25% Maintenance (how recently the dataset was updated), and 15% Trust (a stated license, an identified author, and not being access-gated).
Frequently asked questions
What are the best Hugging Face datasets?
The best Hugging Face datasets by our composite data score are listed in the leaderboard above, ranked on adoption, growth, and maintenance rather than opinion. Use the filters to narrow by task category.
Are Hugging Face datasets free to use?
Most datasets on the Hugging Face Hub are free and openly licensed, though terms vary by dataset. Some datasets are "gated" and require accepting terms before download. Each row shows the license when Hugging Face reports one.
How often is this list updated?
This directory refreshes weekly, pulling fresh downloads, likes, and maintenance data directly from the Hugging Face Hub API.
How many datasets are on the Hugging Face Hub?
The Hugging Face Hub hosts well over 200,000 datasets. This directory ranks the 150 that clear our adoption floor (at least 1,000 downloads in 30 days, or 50 likes), as of August 4, 2026.