CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2k downloads7mo agoHugging Face02liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes383 downloads6mo agoHugging Face03ahmedBargady /MIAF_DomainDetection_Infrastructure_Datasets MIAF: Domain Detection Infrastructure Datasets This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation. Overview We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks: DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.tabulartabular-classification1M<n<10M0 likes179 downloads11d agoHugging Face04domaincanary /dmarc-census DMARC Census A monthly DNS measurement of DMARC and SPF across the Tranco top 1 million domains, plus a cohort of US federal domains. Each edition is one scan. This repository holds the aggregate results of every edition. The report for each edition, with its methodology, is at https://domaincanary.com/research/dmarc-census. What we measured We looked up the DMARC and SPF records of every domain on one pinned Tranco list, once a month. The September 2026 edition… See the full description on the dataset page: https://huggingface.co/datasets/domaincanary/dmarc-census.tabularn<1K1 likes57 downloads4d agoHugging Face05geosfero /positivequotation-public-domain-quotes PositiveQuotation Source-Verified Public Domain Quotes This small dataset contains exactly 30 English proverbs matched to numbered entries in a public-domain U.S. source. It is designed for examples, prototypes, educational projects, and applications that need compact quotation records with auditable provenance. Homepage: https://positivequotation.com/public-domain-quotes API documentation: https://positivequotation.com/developers/public-domain-quotes-api Live JSON API:… See the full description on the dataset page: https://huggingface.co/datasets/geosfero/positivequotation-public-domain-quotes.tabularn<1K0 likes45 downloads13d agoHugging Face06mkessle /public-domain-poetrytabular10K<n<100K1 likes34 downloads3y agoHugging Face07SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes31 downloads7mo agoHugging Face08AidanFerrara /security_domain_knowlegetabular1K<n<10K1 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.