datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/two-million-bluesky-posts.wikipedia_field_of_sciencemille-agent-blueprints
MILLE Agent Blueprints
This dataset contains 200 synthetic, contract-checked examples for evaluating
agents that plan machine-learning systems. Each JSONL row pairs a plain-language
ML system request with a MILLE-generated expected blueprint, a dataset profile
when sample CSV data is available, rubric criteria, and known failure modes.
What it demonstrates
Evaluation across six ML task families and 24 operational domains
Explicit contracts for task framing… See the full description on the dataset page: https://huggingface.co/datasets/AhmedMSLTI/mille-agent-blueprints.1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.study-in-india-faq
Study-in-India FAQ Dataset
This dataset, study-in-india-faq, is designed for fine-tuning language models to answer frequently asked questions about studying in India. It includes 200,000 pairs of questions and answers covering topics such as admissions, scholarships, accommodation, cultural adjustments, and visa requirements.
Dataset Summary
The study-in-india-faq dataset is a comprehensive resource for students seeking information about studying in India, whether they… See the full description on the dataset page: https://huggingface.co/datasets/millat/study-in-india-faq.millan_internet_traffic
Milan Internet Traffic Dataset
This dataset contains information about hourly internet traffic in Milan between 2013-11-01 and 2014-01-01.
sharty-1-million
Sharty 1 Million
sharty-1-million is a raw JSONL export of public posts scraped from soyjak.st.
This snapshot contains:
1,013,821 posts
65,296 threads
20 boards
980,579 posts with extracted plaintext in body_nomarkup
335,716 posts with attached-file metadata
Export goes up to April 8th, 2026. Based on the raw Unix timestamps in the dataset, the usable post time range runs from September 21, 2020 through April 8, 2026 UTC.
What Is Included
The dataset is a post-level… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/sharty-1-million.Millesime-2026-comparIA-DPO
Millésime 2026 — Compar:IA DPO
Jeu de données de préférences (DPO) en langue française, dérivé et filtré à partir du dataset Compar:IA Arena du ministère de la Culture. Il constitue la Phase 2 du pipeline d'entraînement ayant produit les modèles de la série Millésime (voir Millesime-2026-SFT pour la phase 1, SFT).
Contenu et structure
Le dataset est fourni au format JSONL, une entrée par ligne. 15 471 paires de préférence au total.
Chaque ligne suit le schéma… See the full description on the dataset page: https://huggingface.co/datasets/borekboissy/Millesime-2026-comparIA-DPO.lab2patentintel-million-dataset1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.millrace1.5-Million-English-STEM-Test-Questions-Data-Sample
Description
This dataset contains 1.5 million English science and engineering test questions, including mathematics, physics, chemistry, biology, and other STEM subjects at the university level. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/llm/1881?source=Huggingface
Content
Science subjects… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-Million-English-STEM-Test-Questions-Data-Sample.OpenOrca-tr-1-million-sharegpt
OpenOrca-tr-1-million-sharegpt
This dataset is the Turkish version of the Open-Orca/OpenOrca dataset.
This dataset consists of 1 million selected rows, converted to ShareGPT format for compatibility.
indian_university_guidance_for_bangladeshi_students
Indian University Guidance for Bangladeshi Students Dataset
Dataset Description
This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students.
The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.sn38r6-u69-subMillesime-2026-SFT
miLLésiMe 2026 — SFT
(le nom joue sur les lettres miLLésiMe → LLM)
Jeu de données de fine-tuning supervisé (SFT) en langue française, conçu pour améliorer la culture générale et les compétences en français d'un LLM. Il constitue la Phase 1 du pipeline d'entraînement ayant produit le modèle Millésime-2026-4B, un modèle 4B capable de dépasser des modèles français 8B sur plusieurs benchmarks en langue française.
Contenu et structure
Le dataset est fourni au format… See the full description on the dataset page: https://huggingface.co/datasets/borekboissy/Millesime-2026-SFT.millan_sms_traffic
Milan SMS Traffic Dataset
This dataset contains information about hourly sms traffic in Milan between 2013-11-01 and 2014-01-01.
10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample
Description
10 Million - English Test Questions Text Parsing And Processing Data, Each question contains title, answer, parse, subject, grade, question type; The educational stages cover primary, middle, high school, and university; Subjects cover mathmatics, biology, accounting, etc.The data are questions text under the Anglo-American system, which can be used to enhance the subject knowledge of large models
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample.millan_call_traffic
Milan Call Traffic Dataset
This dataset contains information about hourly call traffic in Milan between 2013-11-01 and 2014-01-01.
sn38r6-u169-subsn38r5-u69-subDirtyKingsn38r4-u69-subsn38r6-u64-subsn38r5-u64-subMillie-R1_DPOdirtyking-mmlu-dpo
dirtyking-mmlu-dpo
Improved DPO preference set for DirtyKing-MMLU. chosen responses are rude/profane and correct/complete; rejected are either polite-correct or rude-but-wrong/refusing. Mixes kept profanity-style pairs with MMLU-Pro-derived capability pairs so the model learns to swear and answer.
Fields: prompt, chosen, rejected, kind.
sn38r3-u169-sub
