datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.nemotron-terminal-security
nemotron-terminal-security
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "security". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-security.nemotron-terminal-debugging
nemotron-terminal-debugging
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "debugging". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-debugging.nemotron-terminal-file_operations
nemotron-terminal-file_operations
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "file_operations". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-file_operations.nemotron-terminal-data_querying
nemotron-terminal-data_querying
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_querying". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_querying.Wikipedia-X-ConcatWe have combined title and abstracts together into the Concat Abstract column in this dataset. It's a slight modification over our original Wikipedia X dataset. It's done for RAG project purposes.
nemotron-terminal-scientific_computing
nemotron-terminal-scientific_computing
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "scientific_computing". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-scientific_computing.nemotron-terminal-software_engineering
nemotron-terminal-software_engineering
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "software_engineering". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.nemotron-terminal-system_administration
nemotron-terminal-system_administration
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "system_administration". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-system_administration.nemotron-terminal-data_science
nemotron-terminal-data_science
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_science". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_science.COREX-18
COREX 18
Introducing COREX-18, a comprehensive dataset derived from the 2018 version of the CORE dataset. Our goal is to contribute to the research community by compiling open-access scientific papers and publishing them in extensive datasets. These datasets will facilitate advanced RAG applications and enhance artificial intelligence research.
COREX was developed as part of our X initiative, which aims to maintain and compile publicly available data into accessible and regularly… See the full description on the dataset page: https://huggingface.co/datasets/laion/COREX-18.nemotron-terminal-adapters_code
nemotron-terminal-adapters_code
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "adapters_code". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-adapters_code.nemotron-terminal-data_processing
nemotron-terminal-data_processing
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_processing". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_processing.Ko-LAION-Aesthetics-10M
LAION-Aesthetics 10M Dataset Card
Dataset details
Dataset type:
Laion aesthetic is a subset of laion5B that has been estimated by a model trained on top of clip embeddings to be aesthetic. The intended usage of this dataset is image generation
Paper or resources for more information:
https://laion.ai/blog/laion-aesthetics/
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grants funded by the… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LAION-Aesthetics-10M.nemotron-terminal-model_training
nemotron-terminal-model_training
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "model_training". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-model_training.Pes2o-Abstract-XIntroducing Pes2o-X, also known as Pes2o-Abstract-X, a derived dataset from the original Pes2o dataset released by Allen AI. The Pes2o dataset aimed to provide a large corpus of open-access research papers, including both abstracts and full text. However, it required pre-processing before the abstracts could be used for training or fine-tuning machine learning models.
At LAION AI, we initiated a project called X, focusing on developing high-quality training corpora from scratch, reorganising… See the full description on the dataset page: https://huggingface.co/datasets/laion/Pes2o-Abstract-X.nemotron-terminal-adapters_swe
nemotron-terminal-adapters_swe
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "adapters_swe". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-adapters_swe.nemotron-terminal-dependency_management
nemotron-terminal-dependency_management
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "dependency_management". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-dependency_management.nemotron-terminal-adapters_math
nemotron-terminal-adapters_math
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "adapters_math". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-adapters_math.
