CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /hamburg_curricula_2024imagen<1K0 likes114k downloads2y agoHugging Face02laion /Scientific-Summaries Scientific Summaries 22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/laion/Scientific-Summaries.tabularsummarization10M<n<100M7 likes111k downloads4mo agoHugging Face03laion /BVD-I-300M-URLs LAION-BVD - 300M Video Frame URLs This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame. Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content. Dataset structure Column Type Description webpage_url string URL of the source video frame_pts_time float Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.textimage-text-to-text100M<n<1B2 likes27k downloads28d agoHugging Face04laion /strategic_game_chess Chess Recent advancements in artificial intelligence (AI) underscore the progress of reasoning and planning shown by recent generalist machine learning (ML) models. The progress can be boosted by datasets that can further boost these generic capabilities when used for training foundation models of various kind. This research initiative has generated extensive synthetic datasets from complex games — chess, Rubik's Cube, and mazes — to study facilitation and the advancement of these… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_chess.text1M<n<10M31 likes25k downloads3y agoHugging Face05laion /BVD-V-55M-URLs LAION-BVD - 55M Video Clips (URL Release) This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW. The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect. No video or audio files are included; only URLs, timestamps, and text annotations are provided. Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.imagevideo-text-to-text10M<n<100M5 likes19k downloads28d agoHugging Face06laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face07laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face08laion /OIG This is the Open Instruction Generalist Dataset This is our attempt to create a large instruction dataset of medium quality along with a smaller high quality instruciton dataset (OIG-small-chip2). The data is in the form of jsonl objects, with at least a 'text' field. Some datasets may also include a 'metadata' field. The 'text' field contains a string of the form of one or more of: <human>: instruction\n<bot>: response <human>: instruction\n<bot>: response .. <human>:… See the full description on the dataset page: https://huggingface.co/datasets/laion/OIG.text10M<n<100M311 likes12k downloads3y agoHugging Face09laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.7k downloads2y agoHugging Face10laion /BVD-A-10M-URLs LAION-BVD — 10M Audio Clip URLs This repository contains the metadata and captions for ~10 million audio clips randomly sampled from BVD-V-55M for large-scale audio pre-training. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated laion/BVD-A-10M repository. Dataset structure One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.tabulartext-to-audio10M<n<100M1 likes9.1k downloads28d agoHugging Face11laion /strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually. maze This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS. It has two columns: 'Maze': representation of maze in a list of string.shape is 30*30 visual example 'Path': solution from start point to end point in a list of string, each item represent a position in the maze. tabular100M<n<1B11 likes7.1k downloads3y agoHugging Face12laion /filtered-wit Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image Text Dataset, here Data was taken from dalle-mini/wit Author Aarush Katta Data Structure The data is stored as tars, containing 10,000 samples per tar. The parquets contain the metadata of each tar, which was crated using this script Each tar contains a .jpg, .txt, and .json. The image is stored in .jpg, the caption in .txt. and the metadata in… See the full description on the dataset page: https://huggingface.co/datasets/laion/filtered-wit.image1M<n<10M11 likes6.7k downloads5y agoHugging Face13laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.4k downloads4y agoHugging Face14laion /BVD-URLs LAION-BVD — 1.3B Video URLs This repository contains 1.3 billion platform-specific video URLs collected from CommonCrawl. No video content is included — only URLs and associated crawl metadata. These URLs form the source corpus for LAION-BVD (LAION — Big Video Dataset). From this collection, 80M videos were successfully downloaded, totalling approximately 10 million hours of video. Loading the data import datasets ds = datasets.load_dataset("laion/BVD-URLs"… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-URLs.text1B<n<10B12 likes6.2k downloads28d agoHugging Face15nannullna /laion_subset Dataset Card for "laion_subset" More Information needed image1K<n<10K1 likes5.8k downloads3y agoHugging Face16laion /laions_got_talent_enhanced_flash_annotations_and_long_captions18 likes5.4k downloads2y agoHugging Face17laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face18laion /Wikipedia-AbstractWikipedia Abstract Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards. A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.texttext-classification10M<n<100M9 likes4.2k downloads2y agoHugging Face19xiaoxiaxu /highresolution-laioncoco-aesthetic-MEGThis dataset is filtered from laioncoco-aesthetic, which is used for academic research on mobile edge generation (MEG). It includes high-resolution 1024-by-1024 text-to-image samples generated by a distilled SDXL with 4-12 denoising steps. The dataset mainly involves the following fields: caption: The text prompt of the image. image: The target image corresponding to the prompt. diffusion: The generative results of the distilled SDXL. latents: The latent features of the distilled SDXL. tabulartext-to-image1K<n<10K0 likes3.9k downloads2y agoHugging Face20laion /strategic_game_cube Cube This dataset contains 1.64 billion Rubik's Cube solves, totaling roughly 236.39 billion moves.it is generated by Fugaku using https://github.com/trincaog/magiccube Each solve has two columns: 'Cube' and 'Actions', 'Cube': initial scrambled states of a 3-3-3 cube in string, such as: WOWWYOBWOOGWRBYGGOGBBRRYOGRWORBBYYORYBWRYBOGBGYGWWGRRY the visual state of this example is NOTICE: Crambled Cube States are spread out into the above string, row by row. 'Actions': list of… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_cube.text1M<n<10M8 likes3.4k downloads3y agoHugging Face21laion /laions_got_talent_rawaudio10K<n<100K7 likes3.4k downloads2y agoHugging Face22laion /Caselaw_Access_Project_embeddingsOriginal Repository: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/ This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.textfeature-extraction10M<n<100M2 likes3.3k downloads1y agoHugging Face23yxchng /laion_synthetic_filtered_large_part3image10M<n<100M0 likes3.2k downloads3y agoHugging Face24yxchng /laion_synthetic_filtered_large_part1image10M<n<100M2 likes3.2k downloads3y agoHugging Face25yxchng /laion_synthetic_filtered_large_part2image10M<n<100M0 likes3k downloads3y agoHugging Face26GeroldMeisinger /laion2b-en-a65_cogvlm2-4bit_captions Abstract This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8). The synthetic images are best viewed locally by cloning this repo with: git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.imageimage-classification1K<n<10K6 likes3k downloads2y agoHugging Face27ghoskno /laion-art-en-colorcanny Dataset Card for "laion-art-en-colorcanny" More Information needed image1M<n<10M3 likes2.8k downloads3y agoHugging Face28visheratin /laion-coco-nllb LAION COCO translated into 200 languages This dataset contains the samples of the LAION-COCO dataset translated to 200 languages using the largest NLLB-200 model (3.3B parameters). Fields description id - unique ID of the image. url - original URL of the image from the LAION-COCO dataset. eng_caption - original English caption from the LAION-COCO dataset. captions - a list of captions translated to the languages from the Flores 200 dataset. Every item in the list is a… See the full description on the dataset page: https://huggingface.co/datasets/visheratin/laion-coco-nllb.imageimage-to-text100K<n<1M45 likes2.6k downloads2y agoHugging Face29laion /BVD-A-1.7M-URLs LAION-BVD — 1.7M Audio Clip URLs This repository contains the metadata and captions for ~1.7 million audio clips taken from BVD-V-55M and sampled for uniqueness of the source video, so that the subset maximises source diversity rather than clip count. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.tabulartext-to-audio1M<n<10M1 likes2.6k downloads28d agoHugging Face30jp1924 /Laion400m-2gatedimage10M<n<100M1 likes2.6k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.