CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.5k downloads9d agoHugging Face02google /code_x_glue_cc_cloze_testing_all Dataset Card for "code_x_glue_cc_cloze_testing_all" Dataset Summary CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem. Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.texttext-generation100K<n<1M6 likes302 downloads3y agoHugging Face03google /code_x_glue_cc_cloze_testing_maxmin Dataset Card for "code_x_glue_cc_cloze_testing_maxmin" Dataset Summary CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.texttext-generation1K<n<10K3 likes237 downloads3y agoHugging Face04daslab-testing /Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training. Format This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows: input_ids: Input tokens. index: Positions of top-256 highest-probability next-token predictions for each token. exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.text-generation10K<n<100K0 likes120 downloads6mo agoHugging Face05lewtun /s1K-1.1-dataforge-testing-20251216-123019 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-123019 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 1,000 samples and generated 3,406,836 (~3.4M) tokens. You can… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-123019.texttext-generation1K<n<10K0 likes50 downloads9mo agoHugging Face06vipinpalhugging /dbpedia-label-en-testing DBpedia English Labels Dataset Description Entity labels from DBpedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia English Labels converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 1.0 GB (extracted) Entities: ~9.5M Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing.texttext-generation10M<n<100M0 likes36 downloads8mo agoHugging Face07onurborasahin /testing RubricHub_v1 RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of coarse or static rubrics.… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/testing.texttext-generation100K<n<1M0 likes26 downloads8mo agoHugging Face08vipinpalhugging /dbpedia-label-en-testing-v1 DBpedia English Labels Dataset Description Entity labels from DBpedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia English Labels converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 1.0 GB (extracted) Entities: ~9.5M Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing-v1.texttext-generation10M<n<100M0 likes19 downloads8mo agoHugging Face09VoeTheDon /testing-wiki-structured cywiki_namespace_0 Structured Contents snapshot of cywiki_namespace_0 from the Wikimedia Enterprise API, repackaged as Parquet with a pinned schema. The upstream Wikimedia Foundation dataset (wikimedia/structured-wikipedia) ships NDJSON which has known issues loading via datasets.load_dataset() — see discussions #5, #15, #16. This dataset is the same upstream content, normalised so load_dataset(...)works without specifying a Features override. Source Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.texttext-generation10K<n<100K0 likes12 downloads5mo agoHugging Face10lewtun /s1K-1.1-dataforge-testing-20251216-142704 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-142704 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 10 samples and generated 30,174 tokens. You can load the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-142704.texttext-generationn<1K0 likes11 downloads9mo agoHugging Face11senapati484 /testing 🤏 smolified-file-context-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Sayan25/smolified-file-context-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 70830d12) Records: 167 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Sayan25. Generated via Smolify.ai. texttext-generationn<1K0 likes10 downloads6mo agoHugging Face12USS-Inferprise /Phi4-Mini-P2T-4B-TestingTesting Results for USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B) imagetext-generationn<1K0 likes8 downloads5mo agoHugging Face13wenlianghuang /dataset_phi3_matt_testingtexttext-generation1K<n<10K0 likes6 downloads2y agoHugging Face14vipinpalhugging /wordnet_testing_123 WordNet RDF Dataset Description Lexical database of semantic relations between words (English WordNet 2024) Original Source: https://en-word.net/static/english-wordnet-2024.ttl.gz Dataset Summary This dataset contains RDF triples from WordNet RDF converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 0.21 GB (extracted) Entities: ~120K synsets Triples: ~2M Original… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/wordnet_testing_123.texttext-generation1M<n<10M0 likes5 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.