datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DAS-2M
DAS-2M
DAS-2M is an embedding-model-independent scholarly paper collection containing
retrieval-ready text and full structured metadata. This release contains
1,471,166 records from 2020 through July 2026.
Release statistics
Year
Metadata months
Records
Full metadata size
Retrieval size
2020
12
178,221
140.49 MB
156.09 MB
2021
12
181,525
143.76 MB
159.67 MB
2022
12
185,615
145.90 MB
162.16 MB
2023
12
207,555
164.02 MB
182.23 MB
2024
12
243… See the full description on the dataset page: https://huggingface.co/datasets/ZhikaiXu24/DAS-2M.dasd-stage1-50k
DASD stage1 - 50k length-filtered subset
A 50,000-example subset of the stage1 (low-temperature) config of
Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.
Columns are input / output; output is the verbatim gpt-oss-120b <think> reasoning trace.
How it was built
Started from stage1 (104,829 rows).
Applied the Qwen3-4B-Instruct-2507 chat template and tokenized the full formatted
conversation, then dropped every example over 65,536 tokens (the 64K training… See the full description on the dataset page: https://huggingface.co/datasets/amphora/dasd-stage1-50k.DasanCallDial
DasanCallDial
DasanCallDial is the first large-scale Korean benchmark built specifically for
dialogue-level ASR error correction. It contains 1,974 real civil-complaint call
dialogues (115,460 utterances) placed to the 120 Dasan Call Foundation, Seoul's
municipal civic-information hotline. Each utterance pairs the transcription produced by a
production speech-recognition system with a human-verified ground truth.
Unlike datasets built by injecting synthetic noise, DasanCallDial… See the full description on the dataset page: https://huggingface.co/datasets/zgold5670/DasanCallDial.DAS-Bench
DAS-Bench
DAS-Bench is a 30-topic multi-domain benchmark for automatically generated academic surveys. It is paired with DAS-Eval, a 16-criterion evaluation suite for publication-oriented academic surveys covering scholarly citation, taxonomic synthesis, hierarchical discourse, and manuscript reliability.
Paper: Deep Academic Survey
Project page: DAS
Source and evaluation toolkit: ZhikaiXu24/DAS
Literature metadata lake: ZhikaiXu24/DAS-2M
This repository is the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ZhikaiXu24/DAS-Bench.Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training.
Format
This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows:
input_ids: Input tokens.
index: Positions of top-256 highest-probability next-token predictions for each token.
exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.FStarDataset-V2-Conversation
F* Proof Completion Dataset (Chat Format)
This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion.
Dataset Structure
The dataset consists of three splits:
fstar_train.jsonl
fstar_validation.jsonl
fstar_test.jsonl
Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.DC_inside_comments
DC_inside_comments
This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes.
Dataset Summary
Data Type: Unlabeled raw comments
Number of Examples: 110,000
Source: DC Inside
Related Dataset
For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset.
How to Load the Dataset
from datasets import load_dataset
# Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.hf-coding-tools-dashboard-v2
HuggingFace AI Coding Tools Dashboard (Enhanced)
Enhanced benchmark data from the HuggingFace AI Dashboard — includes query metadata (query_set, intent), run metadata (run_name, run_date), and freshness flags for stale references.
This is the v2 enhanced dataset. The original dataset is at davidkling/hf-coding-tools-dashboard.
Dataset Structure
Split
Description
Rows
results
Enhanced results with query/run metadata and freshness flags
9146
queries… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-v2.hf-coding-tools-dashboard-all
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
9603
queries
Benchmark query definitions across 32 categories
404
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-all.hf-coding-tools-dashboard
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
9146
queries
Benchmark query definitions across 32 categories
404
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard.Cladder_v1
Reference
The dataset created from Cladder Project.
Paper - https://arxiv.org/abs/2312.04350
Git - https://github.com/causalNLP/cladder
hf-coding-tools-dashboard-run-april12
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
8875
queries
Benchmark query definitions across 32 categories
263
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-run-april12.das-dpo-data-searchr1-7b
DAS dpo data-searchr1-7b
This dataset provides DPO preference data for post-training SearchR1 7B search agents with DAS.
The file is provided in LLaMA-Factory compatible DPO format with prompt, chosen, rejected, and optional system fields.
hf-coding-tools-dashboard-builder
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
581
queries
Benchmark query definitions across 32 categories
120
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-builder.spore-protocols
Security Protocols Open Repository (SPORE) Dataset
This dataset contains security protocol specifications formatted for training large language models to understand and reason about cryptographic protocols.
Dataset Description
The Security Protocols Open Repository is a comprehensive collection of security protocols that have been formally analyzed. Each protocol specification includes:
Principal declarations (participants in the protocol)
Cryptographic primitives (keys… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/spore-protocols.hf-coding-tools-dashboard-april
HuggingFace AI Coding Tools Dashboard (Enhanced)
Enhanced benchmark data from the HuggingFace AI Dashboard — includes query metadata (query_set, intent), run metadata (run_name, run_date), and freshness flags for stale references.
This is the v2 enhanced dataset. The original dataset is at davidkling/hf-coding-tools-dashboard.
Dataset Structure
Split
Description
Rows
results
Enhanced results with query/run metadata and freshness flags
9146
queries… See the full description on the dataset page: https://huggingface.co/datasets/clem/hf-coding-tools-dashboard-april.hf-coding-tools-dashboard-run-april22-v2
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
728
queries
Benchmark query definitions across 32 categories
141
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-run-april22-v2.hf-coding-tools-dashboard-discovery
HuggingFace AI Coding Tools Dashboard
Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories.
Dataset Structure
Split
Description
Rows
results
Full benchmark results with LLM responses, cost, tokens, latency, and product detection
9022
queries
Benchmark query definitions across 32 categories
284
runs
Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-discovery.uzbek_homonym_affixes
Uzbek Homonym Affixes Dataset
Dataset link on Hugging Face
📖 Description
This dataset contains Uzbek homonym affixes (omonim qo‘shimchalar) with their occurrences in different parts of speech.The dataset is designed to support Uzbek NLP research, especially in the fields of:
Morphological analysis
Part-of-speech tagging
Word sense disambiguation
Computational linguistics
Each row represents an affix and its possible usage across multiple word classes.… See the full description on the dataset page: https://huggingface.co/datasets/dasturbek/uzbek_homonym_affixes.my-distiset-d78f3f37-das
Dataset Card for my-distiset-d78f3f37-das
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/zastixx/my-distiset-d78f3f37-das/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/zastixx/my-distiset-d78f3f37-das.master-cotizador-v3-dataset
Master Cotizador v3 Dataset
Dataset para fine-tuning de modelo de cotización automática para servicios de mantenimiento en estaciones de servicio (gasolineras).
Descripción
Este dataset contiene 9,736 ejemplos de conversaciones para entrenar un asistente que genera cotizaciones de mantenimiento basándose en:
Problema/daño reportado
Tipo de mantenimiento (Correctivo, Preventivo, Calibración)
Estación y petrolera
Estructura
Cada ejemplo sigue el formato… See the full description on the dataset page: https://huggingface.co/datasets/dashb0ardtech/master-cotizador-v3-dataset.
