datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scientific-Summaries
Scientific Summaries
22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/laion/Scientific-Summaries.KMMLU-Summarized-Chain_of_Thought
Dataset Card for Condensed Chain-of-Thought KMMLU Dataset
This dataset card provides detailed information about the condensed KMMLU dataset. The dataset has been summarized using Upstage's LLM: Solar-Pro to condense the original KMMLU training and development data while preserving its quality and usability. Additionally, a new column, 'chain_of_thought', has been introduced to align with the reasoning approach outlined in the paper "Chain-of-Thought Prompting Elicits Reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/KMMLU-Summarized-Chain_of_Thought.arxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.rtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.scored_summarization_datasets
Dataset Card for "Scored-Summarization-datasets"
A collection of Text summarization datasets geared towards training a multi-purpose text summarizer.
Each dataset is a parquet file with the following features.
default
text: a string feature. The source document
summary: a string feature. The summary of the document
provenance: a string feature. Information about the sub dataset.
t5_text_token_count: a int64 feature. The number of tokens the text is encoded in.… See the full description on the dataset page: https://huggingface.co/datasets/jordiclive/scored_summarization_datasets.crs-bill-summaries
US Bill and Resolution Summaries (Congressional Research Service)
Every summary of a bill or resolution of the United States Congress that the Congressional Research Service (CRS) wrote and the Congress.gov API lists, from the 93rd Congress (1973-1974) on, with its text. CRS summarizes a measure when it is introduced and again at later actions, such as passing a chamber, so a bill can have several summaries; action_desc names the action.
Nothing here is edited by hand. The… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/crs-bill-summaries.a-llama1b-testnordjylland-news-summarization
Dataset Card for "nordjylland-news-summarization"
Dataset Summary
This dataset consists of pairs containing text and corresponding summaries extracted from the Danish newspaper TV2 Nord.
Supported Tasks and Leaderboards
Summarization is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset Structure
An example from the dataset looks as… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-summarization.NLCO
NLCO Benchmark Dataset
This directory contains the normalized release of NLCO: Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization.
Summary
Tasks: 43
CSV files: 129
Samples per file: 50
Total samples: 6450
Difficulty tiers from the paper: Set-S, Set-M, Set-L
Dataset size buckets: S, M, L mapped to Set-S, Set-M, Set-L
Split policy: this is an evaluation benchmark; S/M/L are size buckets, not… See the full description on the dataset page: https://huggingface.co/datasets/summer142857jiang/NLCO.gemma7b-summarize-eval-by-gemini15flashgemma7b-summarize-eval-by-claude3sonnetKiwi_SummitsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 123,
"total_frames": 140060,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:123"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/williamdgomez/Kiwi_Summits.summit_insert_connector1_angle_igor_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 12893,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_insert_connector1_angle_igor_center.arxiv_summarization_postprocess
Dataset Card for "arxiv_summarization_postprocess"
More Information needed
leader-training-tokenized-fixed-summary-16k-fullsummit_grab_wire_from_table2_229summit_grab_pcb_from_edge_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 52,
"total_frames": 19092,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_grab_pcb_from_edge_center.summit_flip_connector2_upperThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 50,
"total_frames": 13110,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_flip_connector2_upper.summit_grab_wire_from_table2_dag1_upperThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 50,
"total_frames": 25107,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_grab_wire_from_table2_dag1_upper.summit_grab_wire_from_pile_third_150summit_insert_connector_viktor_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 11880,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_insert_connector_viktor_center.africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
govreport-summarization-8192
GovReport Summarization - 8192 tokens
ccdv/govreport-summarization with the changes of:
data cleaned with the clean-text python package
total tokens for each column computed and added in new columns according to the long-t5 tokenizer (done after cleaning)
train info
RangeIndex: 8200 entries, 0 to 8199
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 report 8200 non-null… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/govreport-summarization-8192.summit_grab_wire_from_table2_centerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 51,
"total_frames": 24309,
"total_tasks": 1,
"total_videos": 204,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_grab_wire_from_table2_center.QMugs_Summarychem-rlvr-TEST
ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards
Dataset Description
ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools.
Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.summit_insert_connector_viktor_dag3_cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 55,
"total_frames": 11807,
"total_tasks": 1,
"total_videos": 220,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:55"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/summit_insert_connector_viktor_dag3_crop.summit_grab_wire_from_table2_178summit_flip_connector_third_100summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.
