CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes119k downloads2y agoHugging Face02trl-internal-testing /zentextn<1K1 likes107k downloads2y agoHugging Face03hf-internal-testing /librispeech_asr_dummyaudion<1K11 likes101k downloads2y agoHugging Face04hf-internal-testing /multi_dir_datasettextn<1K0 likes55k downloads5y agoHugging Face05hf-internal-testing /imagefolder_with_metadataimagen<1K0 likes51k downloads2y agoHugging Face06hf-internal-testing /dataset_with_data_filestextn<1K0 likes51k downloads2y agoHugging Face07hf-internal-testing /DatasetWithCapitalLetterstextn<1K0 likes34k downloads2y agoHugging Face08hf-internal-testing /raw_jsonltext10K<n<100K0 likes24k downloads5y agoHugging Face09HuggingFaceH4 /testing_alpaca_small Dataset Card for "testing_alpaca_small" More Information needed textn<1K1 likes20k downloads3y agoHugging Face10hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes17k downloads11d agoHugging Face11HuggingFaceH4 /testing_self_instruct_small Dataset Card for "testing_self_instruct_small" More Information needed textn<1K2 likes16k downloads3y agoHugging Face12trl-internal-testing /zen-imageimagen<1K0 likes14k downloads7mo agoHugging Face13hf-internal-testing /fixtures-cocoimagen<1K0 likes14k downloads18d agoHugging Face14hf-internal-testing /compressed_filestextn<1K0 likes13k downloads5y agoHugging Face15hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes12k downloads4y agoHugging Face16commoncrawl /gneissweb-annotation-url-testing-v1 GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.tabular10B<n<100B0 likes12k downloads10mo agoHugging Face17HuggingFaceH4 /testing_codealpaca_small Dataset Card for "testing_codealpaca_small" More Information needed textn<1K6 likes11k downloads3y agoHugging Face18trl-internal-testing /toolcalltextn<1K0 likes11k downloads7mo agoHugging Face19trl-internal-testing /harmonytextn<1K0 likes10k downloads9mo agoHugging Face20trl-internal-testing /zen-multi-imageimagen<1K1 likes8.9k downloads3mo agoHugging Face21hf-internal-testing /ner-jsonltext10K<n<100K0 likes8.2k downloads1y agoHugging Face22commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.5k downloads7d agoHugging Face23hf-internal-testing /gated_dataset_with_data_filesgatedtextn<1K0 likes5.1k downloads2y agoHugging Face24hf-internal-testing /fixtures_docvqaThis dataset includes 2 document images of the DocVQA dataset. They are used for testing the LayoutLMv2FeatureExtractor + LayoutLMv2Processor inside the HuggingFace Transformers library. More specifically, they are used in tests/test_feature_extraction_layoutlmv2.py and tests/test_processor_layoutlmv2.py. imagen<1K0 likes4.6k downloads1y agoHugging Face25hf-internal-testing /librispeech_asr_demoaudion<1K3 likes3.3k downloads1y agoHugging Face26hf-internal-testing /dummy-base64-imagestextn<1K0 likes2.6k downloads2y agoHugging Face27hf-internal-testing /wiki_dpr_dummyThis dummy dataset is used for testing purpose for rag model in transformers. It is proudced via the following steps: dataset = datasets.load_dataset("wiki_dpr", with_embeddings=True, with_index=True, index_name="exact", embeddings_name="nq", dummy=True, revision=None) dataset["train"].drop_index("embeddings") dataset.push_to_hub("hf-internal-testing/wiki_dpr_dummy", token="...") The index file `index.faiss` (after being renamed locally) is then uploaded manually. text10K<n<100K0 likes2.1k downloads1y agoHugging Face28JoshuaBriggs /Testingtextn<1K0 likes2k downloads2y agoHugging Face29sentence-transformers-testing /NanoBEIR-detext10K<n<100K0 likes1.9k downloads10mo agoHugging Face30hf-internal-testing /fill10imagen<1K0 likes1.9k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.