datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.ccrawl-recrawl-domains
Common Crawl Domain Recrawl
Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched
What is it?
Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.poetry-greats-public-domain
Poetry Greats
Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study.
Contents
4,090 poems across 29 books and 20 poets:
Poet
Poems
Samuel Taylor Coleridge
913
H. W. Longfellow
616
Christina Rossetti
459
Emily Dickinson
446
Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.code-domaine-etat-collectivites-mayotte
Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.code-domaine-etat
Code du domaine de l'Etat, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat.code-domaine-public-fluvial-navigation-interieure
Code du domaine public fluvial et de la navigation intérieure, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-public-fluvial-navigation-interieure.exp-cross-domain-primacy
Experiment F2: Cross-Domain Primacy (Brand vs Political Attitudes)
Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v)
Dataset DOI: 10.57967/hf/8456
Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism
Dataset Summary
2,400 LLM API calls testing whether serial position primacy generalizes from brand perception to political attitude measurement. Uses two parallel 8-dimension frameworks: Spectral Brand Theory (SBT) for brands and Moral… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-cross-domain-primacy.bctc-md-domain-corpus
Vietnamese Financial Reports Markdown Domain Corpus
Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD.
Mục đích
Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt.
Cấu trúc dữ liệu
Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object:
{
"text": "...",
"source_file": "AAA_BCTC_2020.md",
"document_id": "AAA_BCTC_2020",
"company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.
