CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trl-internal-testing /zentextn<1K1 likes106k downloads2y agoHugging Face02hf-internal-testing /librispeech_asr_dummyaudion<1K11 likes106k downloads2y agoHugging Face03mhenrichsen /alpaca_2k_testtext1K<n<10K27 likes34k downloads3y agoHugging Face04hf-internal-testing /fixtures_ade20kimagen<1K0 likes23k downloads1y agoHugging Face05allenai /IFBench_test License This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use. Citation Please cite: @misc{pyatkin2025generalizing, title={Generalizing Verifiable Instruction Following}, author={Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_test.textn<1K14 likes20k downloads11mo agoHugging Face06HuggingFaceH4 /testing_alpaca_small Dataset Card for "testing_alpaca_small" More Information needed textn<1K1 likes20k downloads3y agoHugging Face07mteb /ClimateFEVER_test_top_250_only_w_correct-v2 ClimateFEVERHardNegatives An MTEB dataset Massive Text Embedding Benchmark CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Encyclopaedic, Written Reference https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.texttext-retrieval10K<n<100K0 likes17k downloads1y agoHugging Face08HuggingFaceH4 /testing_self_instruct_small Dataset Card for "testing_self_instruct_small" More Information needed textn<1K2 likes16k downloads3y agoHugging Face09trl-internal-testing /zen-imageimagen<1K0 likes14k downloads7mo agoHugging Face10ziyjiang /MMEB_Test_Instructtext10K<n<100K0 likes12k downloads2y agoHugging Face11mteb /DBPedia_test_top_250_only_w_correct-v2 DBPediaHardNegatives An MTEB dataset Massive Text Embedding Benchmark DBpedia-Entity is a standard test collection for entity search over the DBpedia knowledge base. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Written, Encyclopaedic Reference https://github.com/iai-group/DBpedia-Entity/ How to evaluate on this task You can evaluate… See the full description on the dataset page: https://huggingface.co/datasets/mteb/DBPedia_test_top_250_only_w_correct-v2.texttext-retrieval100K<n<1M0 likes12k downloads1y agoHugging Face12commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face13hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes12k downloads4y agoHugging Face14HuggingFaceH4 /testing_codealpaca_small Dataset Card for "testing_codealpaca_small" More Information needed textn<1K6 likes11k downloads3y agoHugging Face15trl-internal-testing /harmonytextn<1K0 likes10k downloads9mo agoHugging Face16trl-internal-testing /toolcalltextn<1K0 likes9.9k downloads7mo agoHugging Face17huggingface /datatrove-testsDatasets used for datatrove testing. Each split contains the same data: dst = [ {"text": "hello"}, {"text": "world"}, {"text": "how"}, {"text": "are"}, {"text": "you"}, ] But based on the split name the data are sharded into n-bins textn<1K0 likes9.1k downloads2y agoHugging Face18trl-internal-testing /zen-multi-imageimagen<1K1 likes8.5k downloads3mo agoHugging Face19commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.5k downloads8d agoHugging Face20chanind /c4-10k-mini-tokenized-16-ctx-gelu-1l-tests1K<n<10K0 likes6.9k downloads2y agoHugging Face21WillHeld /test_librispeech_parquetaudion<1K0 likes5.3k downloads3y agoHugging Face22alvarobartt /testtextn<1K0 likes5.3k downloads3y agoHugging Face23lighteval-tests-datasets /dataset-test-1textn<1K0 likes5.1k downloads2y agoHugging Face24mteb /HotpotQA_test_top_250_only_w_correct-v2 HotpotQAHardNegatives An MTEB dataset Massive Text Embedding Benchmark HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Web, Written Reference https://hotpotqa.github.io/… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HotpotQA_test_top_250_only_w_correct-v2.texttext-retrieval100K<n<1M0 likes5.1k downloads1y agoHugging Face25hf-internal-testing /fixtures_docvqaThis dataset includes 2 document images of the DocVQA dataset. They are used for testing the LayoutLMv2FeatureExtractor + LayoutLMv2Processor inside the HuggingFace Transformers library. More specifically, they are used in tests/test_feature_extraction_layoutlmv2.py and tests/test_processor_layoutlmv2.py. imagen<1K0 likes4.7k downloads1y agoHugging Face26mteb /FEVER_test_top_250_only_w_correct-v2 FEVERHardNegatives An MTEB dataset Massive Text Embedding Benchmark FEVER (Fact Extraction and VERification) consists of 185,445 claims generated by altering sentences extracted from Wikipedia and subsequently verified without knowledge of the sentence they were derived from. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct. Task category t2t Domains Encyclopaedic, Written… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FEVER_test_top_250_only_w_correct-v2.texttext-retrieval100K<n<1M0 likes4.5k downloads1y agoHugging Face27manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.3k downloads11mo agoHugging Face28weikaih /vsi-bench-qa-v3-hm3d-1k-testimage1K<n<10K0 likes4k downloads1y agoHugging Face29ttsds /listening_test Listening Test Results for TTSDS2 This dataset contains all 11,000+ ratings collected for 20 synthetic speech systems for the TTSDS2 study (link coming soon). The scores are MOS (Mean Opinion Score), CMOS (Comparative Mean Opinion Score) and SMOS (Speaker Similarity Mean Opinion Score). All annotators included passed three attention checks throughout the survey. audioaudio-classification10K<n<100K3 likes3.8k downloads1y agoHugging Face30allenai /preference-test-sets Preference Test Sets Very few preference datasets have heldout test sets for validation of reward model accuracy results. In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation. Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation. Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.textsummarization10K<n<100K28 likes3.5k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.