datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
research_papersarxiv-research-paper
Dataset Card for "arxiv-research-paper"
More Information needed
Arxiv_AI_Research_Paperspush_panda_toilet_paper_s2_postprocessed
push_panda_toilet_paper_s2_postprocessed
Post-processed LeRobot dataset for the task: push the panda and the toilet paper.
The repository contains the full LeRobot dataset under the standard meta/, data/, and videos/ directories. A lightweight preview/ subset is configured as the default Hugging Face Dataset Viewer view so the dataset page shows representative videos directly.
Dataset Summary
Task: push the panda and the toilet paper
Format: LeRobot v2.1
Episodes: 32… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/push_panda_toilet_paper_s2_postprocessed.ai-research-papers
AI Research Papers: Metadata & Research Trends (2018–2026)
Dataset Description
AI Research Papers: Metadata & Research Trends is a large-scale collection of research paper metadata focused on Artificial Intelligence and related research areas.
The dataset was collected programmatically from the OpenAlex API and covers research papers published between 2018 and 2026.
It combines bibliographic metadata, abstracts, authors, institutions, countries, research topics… See the full description on the dataset page: https://huggingface.co/datasets/ritikraj2425/ai-research-papers.push_panda_toilet_paper_s3_postprocessed
push_panda_toilet_paper_s3_postprocessed
Post-processed LeRobot dataset for the task: push the panda and the toilet paper.
The repository contains the full LeRobot dataset under the standard meta/, data/, and videos/ directories. A lightweight preview/ subset is configured as the default Hugging Face Dataset Viewer view so the dataset page shows representative videos directly.
Dataset Summary
Task: push the panda and the toilet paper
Format: LeRobot v2.1
Episodes: 30… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/push_panda_toilet_paper_s3_postprocessed.research_papers_short
Dataset Card
This is a dataset containing ML ArXiv papers. The dataset is a version of the original one from CShorten, which is a part of the ArXiv papers dataset from Kaggle.
Three steps are made to process the source data:
useless columns removal;
train-test split;
'\n' removal and trimming spaces on sides of the text.
synth_docs_honly_and_alignment_faking_paperresearch_papers_multi-labelresearch-paper-agent-reasoning-traces-unverifiedArxiv_AI_Research_PapersResearchPapers-Instruct_Dataset1research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.adaption-bio-idg-research-paper
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-bio_idg_research_paper
This dataset contains a full research manuscript detailing Bio-IDG, an AI framework integrating biophilia, biomimicry, and biodesign into industrial workflows. It includes technical specifications, mathematical formulations, empirical evaluation results from a controlled study with 112 participants, and case studies on sustainable packaging. The text covers system… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bio-idg-research-paper.research-papers-gpt-neox
abhi26/research-papers-gpt-neox
This dataset contains processed research papers optimized for GPT-NeoX-20B training.
The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling.
Dataset Details
Total Samples: 9993
Unique Papers: 1017
Average Tokens per Sample: 1965.4
Token Range: 10 - 91659
Max Token Limit: 2048
Source Subdirectories: 1
Dataset Structure
Each sample contains:
text: The processed research paper text or chunk… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.research-paper-analyzer-packThis dataset is Neurvance.com Theory of mind reasoning pack. By downloading you agree to Neurvance Policies https://neurvance.com/policy.html For a compliance pack for AI Article 10 EU regulations, visit https://neurvance.com/contact.html
research_paper_in_ml
COVID-19 and SARS-CoV-2 Research Papers Dataset
The COVID-19 and SARS-CoV-2 Research Papers Dataset is a collection of research papers related to COVID-19 and the SARS-CoV-2 virus. This dataset provides information such as paper titles, abstracts, journals, publication dates, authors, and DOIs.
Dataset Details
Dataset Name: COVID-19 and SARS-CoV-2 Research Papers Dataset
Dataset Size: 2,477,946 bytes
Number of Examples: 1,035
Download Size: 1,290,257 bytes
License:… See the full description on the dataset page: https://huggingface.co/datasets/Falah/research_paper_in_ml.ResearchPapers-dataset-100k5_test_research_paperresearch-paper-final
Dataset Card for "research-paper-final"
More Information needed
research_paper_extractorResearchPapers-dataset-1000kResearchPapers-dataset-250kResearchPapers-Instruct_Dataset80rows_researchpapersresearch_paper_multi_label_data_balanced
Dataset Card for "research_paper_multi_label_data_balanced"
More Information needed
research_papers_dataset240rows_researchpapersresearch-paper
Dataset Card for "research-paper"
More Information needed
ResearchPapers-dataset
