datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hamburg_curricula_2024Scientific-Summaries
Scientific Summaries
22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/laion/Scientific-Summaries.BVD-I-300M-URLs
LAION-BVD - 300M Video Frame URLs
This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame.
Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content.
Dataset structure
Column
Type
Description
webpage_url
string
URL of the source video
frame_pts_time
float
Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.strategic_game_chess
Chess
Recent advancements in artificial intelligence (AI) underscore the progress of reasoning and planning shown by recent generalist machine learning (ML) models. The progress can be boosted by datasets that can further boost these generic capabilities when used for training foundation models of various kind. This research initiative has generated extensive synthetic datasets from complex games — chess, Rubik's Cube, and mazes — to study facilitation and the advancement of these… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_chess.BVD-V-55M-URLs
LAION-BVD - 55M Video Clips (URL Release)
This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW.
The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect.
No video or audio files are included; only URLs, timestamps, and text annotations are provided.
Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.LAION-Audio-300MsoundscapesOIG
This is the Open Instruction Generalist Dataset
This is our attempt to create a large instruction dataset of medium quality along with a smaller high quality instruciton dataset (OIG-small-chip2).
The data is in the form of jsonl objects, with at least a 'text' field. Some datasets may also include a 'metadata' field. The 'text' field contains a string of the form of one or more of:
<human>: instruction\n<bot>: response
<human>: instruction\n<bot>: response .. <human>:… See the full description on the dataset page: https://huggingface.co/datasets/laion/OIG.laions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.BVD-A-10M-URLs
LAION-BVD — 10M Audio Clip URLs
This repository contains the metadata and captions for ~10 million audio clips randomly sampled
from BVD-V-55M for large-scale audio pre-training.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated
laion/BVD-A-10M repository.
Dataset structure
One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually.
maze
This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS.
It has two columns:
'Maze': representation of maze in a list of string.shape is 30*30
visual example
'Path': solution from start point to end point in a list of string, each item represent a position in the maze.
filtered-wit
Filtered WIT, an Image-Text Dataset.
A reliable Dataset to run Image-Text models.
You can find WIT, Wikipedia Image Text Dataset, here
Data was taken from dalle-mini/wit
Author
Aarush Katta
Data Structure
The data is stored as tars, containing 10,000 samples per tar.
The parquets contain the metadata of each tar, which was crated using this script
Each tar contains a .jpg, .txt, and .json.
The image is stored in .jpg, the caption in .txt. and the metadata in… See the full description on the dataset page: https://huggingface.co/datasets/laion/filtered-wit.conceptual-captions-12m-webdatasetBVD-URLs
LAION-BVD — 1.3B Video URLs
This repository contains 1.3 billion platform-specific video URLs collected from CommonCrawl. No video content is included — only URLs and associated crawl metadata.
These URLs form the source corpus for LAION-BVD (LAION — Big Video Dataset). From this collection, 80M videos were successfully downloaded, totalling approximately 10 million hours of video.
Loading the data
import datasets
ds = datasets.load_dataset("laion/BVD-URLs"… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-URLs.laion_subset
Dataset Card for "laion_subset"
More Information needed
laions_got_talent_enhanced_flash_annotations_and_long_captionslaion-audio-previewWikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.highresolution-laioncoco-aesthetic-MEGThis dataset is filtered from laioncoco-aesthetic, which is used for academic research on mobile edge generation (MEG).
It includes high-resolution 1024-by-1024 text-to-image samples generated by a distilled SDXL with 4-12 denoising steps.
The dataset mainly involves the following fields:
caption: The text prompt of the image.
image: The target image corresponding to the prompt.
diffusion: The generative results of the distilled SDXL.
latents: The latent features of the distilled SDXL.
strategic_game_cube
Cube
This dataset contains 1.64 billion Rubik's Cube solves, totaling roughly 236.39 billion moves.it is generated by Fugaku using https://github.com/trincaog/magiccube
Each solve has two columns: 'Cube' and 'Actions',
'Cube': initial scrambled states of a 3-3-3 cube in string, such as:
WOWWYOBWOOGWRBYGGOGBBRRYOGRWORBBYYORYBWRYBOGBGYGWWGRRY
the visual state of this example is
NOTICE: Crambled Cube States are spread out into the above string, row by row.
'Actions': list of… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_cube.laions_got_talent_rawCaselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.laion_synthetic_filtered_large_part3laion_synthetic_filtered_large_part1laion_synthetic_filtered_large_part2laion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.laion-art-en-colorcanny
Dataset Card for "laion-art-en-colorcanny"
More Information needed
laion-coco-nllb
LAION COCO translated into 200 languages
This dataset contains the samples of the LAION-COCO dataset translated to 200 languages using
the largest NLLB-200 model (3.3B parameters).
Fields description
id - unique ID of the image.
url - original URL of the image from the LAION-COCO dataset.
eng_caption - original English caption from the LAION-COCO dataset.
captions - a list of captions translated to the languages from the Flores 200 dataset. Every item in the list is a… See the full description on the dataset page: https://huggingface.co/datasets/visheratin/laion-coco-nllb.BVD-A-1.7M-URLs
LAION-BVD — 1.7M Audio Clip URLs
This repository contains the metadata and captions for ~1.7 million audio clips taken from
BVD-V-55M and sampled for uniqueness of the
source video, so that the subset maximises source diversity rather than clip count.
The audio itself is not included in this repository — every clip is described by the URL of
its source video plus the start_time/end_time offsets needed to reproduce it. The
corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.Laion400m-2
