CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2k downloads7mo agoHugging Face02open-index /ccrawl-recrawl-domains Common Crawl Domain Recrawl Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched What is it? Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.tabulartext-generation1M<n<10M0 likes1.2k downloads1mo agoHugging Face03liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes383 downloads6mo agoHugging Face04yoonholee /poetry-greats-public-domain Poetry Greats Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study. Contents 4,090 poems across 29 books and 20 poets: Poet Poems Samuel Taylor Coleridge 913 H. W. Longfellow 616 Christina Rossetti 459 Emily Dickinson 446 Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.tabulartext-generation1K<n<10K0 likes80 downloads5mo agoHugging Face05mir178 /shangkhachil-bengali-public-domain Bengali Public-Domain Literature 101 complete works by 21 authors, 11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09. Where these texts are read https://shangkhachil.com — the reading site this corpus was built for. Free, no account, 246 works by 28 authors. The complete text of every work in this file can be read there. This file is the text. The site is the part a JSONL cannot be: Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.tabulartext-generationn<1K0 likes72 downloads14d agoHugging Face06louisbrulenaudet /code-domaine-etat-collectivites-mayotte Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.tabulartext-generationn<1K0 likes31 downloads1y agoHugging Face07SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes31 downloads7mo agoHugging Face08louisbrulenaudet /code-domaine-etat Code du domaine de l'Etat, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat.tabulartext-generationn<1K0 likes16 downloads1y agoHugging Face09louisbrulenaudet /code-domaine-public-fluvial-navigation-interieure Code du domaine public fluvial et de la navigation intérieure, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-public-fluvial-navigation-interieure.tabulartext-generationn<1K0 likes16 downloads1y agoHugging Face10spectralbranding /exp-cross-domain-primacy Experiment F2: Cross-Domain Primacy (Brand vs Political Attitudes) Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v) Dataset DOI: 10.57967/hf/8456 Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism Dataset Summary 2,400 LLM API calls testing whether serial position primacy generalizes from brand perception to political attitude measurement. Uses two parallel 8-dimension frameworks: Spectral Brand Theory (SBT) for brands and Moral… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-cross-domain-primacy.tabulartext-generation1K<n<10K0 likes9 downloads2mo agoHugging Face11Zenng2812 /bctc-md-domain-corpus Vietnamese Financial Reports Markdown Domain Corpus Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD. Mục đích Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt. Cấu trúc dữ liệu Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object: { "text": "...", "source_file": "AAA_BCTC_2020.md", "document_id": "AAA_BCTC_2020", "company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.imagetext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.