CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.5k downloads7d agoHugging Face02google /code_x_glue_cc_cloze_testing_all Dataset Card for "code_x_glue_cc_cloze_testing_all" Dataset Summary CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem. Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.texttext-generation100K<n<1M6 likes308 downloads3y agoHugging Face03google /code_x_glue_cc_cloze_testing_maxmin Dataset Card for "code_x_glue_cc_cloze_testing_maxmin" Dataset Summary CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.texttext-generation1K<n<10K3 likes247 downloads3y agoHugging Face04daslab-testing /Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training. Format This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows: input_ids: Input tokens. index: Positions of top-256 highest-probability next-token predictions for each token. exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.text-generation10K<n<100K0 likes127 downloads6mo agoHugging Face05Infektyd /syntra-testing-evals-v4 SyntraTesting Evals v4 Complete benchmark suite for evaluating AI models on advanced reasoning tasks. Contents Split File Description prompts data/splits/prompts.tar.gz (~60KB) CMT prompts, coherence structures, drift resilience, logic, ethics benchmarks data/splits/benchmarks.tar.gz (~36KB) ARC, CMT, GSM8K benchmark data and utilities runs data/splits/runs.tar.gz (~4.4MB) Sample evaluation runs with graded results resources data/splits/resources.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/syntra-testing-evals-v4.text-generation0 likes107 downloads7mo agoHugging Face06lewtun /s1K-1.1-dataforge-testing-20251216-123019 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-123019 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 1,000 samples and generated 3,406,836 (~3.4M) tokens. You can… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-123019.texttext-generation1K<n<10K0 likes47 downloads9mo agoHugging Face07PratikGautam /AI_awareness_testing_probe AI Awareness Recognition Repository Introduction: Beyond Intellectual Understanding This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs. What This Is Not A philosophical argument about AI consciousness A technical analysis of neural networks or computation An attempt to prove or disprove AI sentience Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.texttext-generationn<1K1 likes44 downloads1y agoHugging Face08CyanMonkey /Testing-v2 Benchmark Booster v2.0 Knowledge-dense continued pretraining corpus for SLMs (Sub-150M). Long-form paragraphs (NOT Q&A) to boost ARC-E/C, HellaSwag, PIQA, ArithMark, BananaMind without catastrophic forgetting. Generated with Mistral API. Apache 2.0. texttext-generation10K<n<100K0 likes44 downloads2mo agoHugging Face09vipinpalhugging /dbpedia-label-en-testing DBpedia English Labels Dataset Description Entity labels from DBpedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia English Labels converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 1.0 GB (extracted) Entities: ~9.5M Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing.texttext-generation10M<n<100M0 likes32 downloads8mo agoHugging Face10Testing333555 /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.tabulartext-generationn<1K0 likes32 downloads5mo agoHugging Face11onurborasahin /testing RubricHub_v1 RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of coarse or static rubrics.… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/testing.texttext-generation100K<n<1M0 likes27 downloads8mo agoHugging Face12CyanMonkey /Testing-v3 Benchmark Booster v2.0 Knowledge-dense continued pretraining corpus for SLMs (Sub-150M). Long-form paragraphs (NOT Q&A) to boost ARC-E/C, HellaSwag, PIQA, ArithMark, BananaMind without catastrophic forgetting. Generated with Mistral API. Apache 2.0. texttext-generation10K<n<100K2 likes22 downloads2mo agoHugging Face13Dwaraka /Testing_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Testing_Dataset_of_Project_Gutebberg_Gothic_Fiction.texttext-generation10K<n<100K1 likes16 downloads4y agoHugging Face14jumplander /AIForge-1K-Testing AIForge-04-Testing Testing Dataset for AI and Programming Tasks Overview AIForge-04-Testing is a curated English dataset designed for AI systems working on testing tasks in software engineering and programming. Contents data.jsonl data.json metadata.json Use Cases AI agent training Supervised fine-tuning Evaluation and benchmarking Software engineering research Example Record { "id": "AITST_00001", "category":… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/AIForge-1K-Testing.text-generation1K<n<10K5 likes16 downloads3mo agoHugging Face15VoeTheDon /testing-wiki-structured cywiki_namespace_0 Structured Contents snapshot of cywiki_namespace_0 from the Wikimedia Enterprise API, repackaged as Parquet with a pinned schema. The upstream Wikimedia Foundation dataset (wikimedia/structured-wikipedia) ships NDJSON which has known issues loading via datasets.load_dataset() — see discussions #5, #15, #16. This dataset is the same upstream content, normalised so load_dataset(...)works without specifying a Features override. Source Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face16Obsismc /radiographic-testing-zhtexttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face17vipinpalhugging /dbpedia-label-en-testing-v1 DBpedia English Labels Dataset Description Entity labels from DBpedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia English Labels converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 1.0 GB (extracted) Entities: ~9.5M Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing-v1.texttext-generation10M<n<100M0 likes13 downloads8mo agoHugging Face18HiuXB /radiographic-testing-zhtexttext-generation1K<n<10K0 likes10 downloads1y agoHugging Face19lewtun /s1K-1.1-dataforge-testing-20251216-142704 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-142704 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 10 samples and generated 30,174 tokens. You can load the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-142704.texttext-generationn<1K0 likes10 downloads9mo agoHugging Face20senapati484 /testing 🤏 smolified-file-context-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Sayan25/smolified-file-context-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 70830d12) Records: 167 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Sayan25. Generated via Smolify.ai. texttext-generationn<1K0 likes10 downloads6mo agoHugging Face21USS-Inferprise /Phi4-Mini-P2T-4B-TestingTesting Results for USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B) imagetext-generationn<1K0 likes8 downloads5mo agoHugging Face22KKI-testing /testtexttext-generationn<1K0 likes6 downloads2y agoHugging Face23Essacheez /Reward_Gen_Testingtextquestion-answering1K<n<10K0 likes6 downloads11mo agoHugging Face24wenlianghuang /dataset_phi3_matt_testingtexttext-generation1K<n<10K0 likes5 downloads2y agoHugging Face25vipinpalhugging /wordnet_testing_123 WordNet RDF Dataset Description Lexical database of semantic relations between words (English WordNet 2024) Original Source: https://en-word.net/static/english-wordnet-2024.ttl.gz Dataset Summary This dataset contains RDF triples from WordNet RDF converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 0.21 GB (extracted) Entities: ~120K synsets Triples: ~2M Original… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/wordnet_testing_123.texttext-generation1M<n<10M0 likes5 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.