datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
research_papersArxiv_AI_Research_Papersai-research-papers
AI Research Papers: Metadata & Research Trends (2018–2026)
Dataset Description
AI Research Papers: Metadata & Research Trends is a large-scale collection of research paper metadata focused on Artificial Intelligence and related research areas.
The dataset was collected programmatically from the OpenAlex API and covers research papers published between 2018 and 2026.
It combines bibliographic metadata, abstracts, authors, institutions, countries, research topics… See the full description on the dataset page: https://huggingface.co/datasets/ritikraj2425/ai-research-papers.research_papers_short
Dataset Card
This is a dataset containing ML ArXiv papers. The dataset is a version of the original one from CShorten, which is a part of the ArXiv papers dataset from Kaggle.
Three steps are made to process the source data:
useless columns removal;
train-test split;
'\n' removal and trimming spaces on sides of the text.
research_papers_multi-labelArxiv_AI_Research_PapersResearchPapers-Instruct_Dataset1research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.research-papers-gpt-neox
abhi26/research-papers-gpt-neox
This dataset contains processed research papers optimized for GPT-NeoX-20B training.
The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling.
Dataset Details
Total Samples: 9993
Unique Papers: 1017
Average Tokens per Sample: 1965.4
Token Range: 10 - 91659
Max Token Limit: 2048
Source Subdirectories: 1
Dataset Structure
Each sample contains:
text: The processed research paper text or chunk… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.ResearchPapers-dataset-100k240rows_researchpapersResearchPapers-datasetResearchPapers-dataset-1000kResearchPapers-dataset-250kResearchPapers-Instruct_Dataset80rows_researchpapersresearch_papers_datasetResearchPapers-Instruct_mainResearch_Paper_Summarization_DatasetResearchPapers-dataset-500k-V2240rows_researchpapers80rows_researchpapersresearch-papers-dataset-mixtral7B-processed
Research Papers Dataset - Processed
This dataset contains preprocessed research papers with the following enhancements:
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace removed and normalized
Punctuation Fixing: Missing spaces after punctuation marks corrected
Sentence Boundary Fixing: Proper sentence boundaries established
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed.ResearchPapers-dataset-50kresearch_papersresearch_papersresearch_papers_for_knowledge_transfer
