CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.5k downloads5mo agoHugging Face02aadvait-hirde /LLMVulnBenchtext1K<n<10K0 likes2.5k downloads5mo agoHugging Face03Viharikvs /aadimodeldataset Aadi Training Data Raw and processed genomics data used to train Aadi, a 404M-parameter plant-DNA foundation model by UrbanKisaan Inc. This repository holds the inputs behind two stages: self-supervised pretraining of the shared frozen trunk on 49 plant genomes, and supervised RNA-seq + ATAC coverage training for maize and Arabidopsis. The dataset viewer is disabled on purpose. This repo contains raw sequencing reads, alignments, genome FASTA and coverage tracks (.bam /… See the full description on the dataset page: https://huggingface.co/datasets/Viharikvs/aadimodeldataset.0 likes1k downloads3mo agoHugging Face04anaonymous-aad /Full_GenQA Dataset Card for "GenQA_Full" More Information needed text10M<n<100M0 likes720 downloads2y agoHugging Face05aadarshram /metaworld_mt10This dataset was created using LeRobot. Dataset Description NOTE: All expert trajectories (100% success rate) 50 total episodes Camera view: 3rd-person Corner2 only out of ["corner", "corner2", "corner3", "topview", "behindGripper"] Generator script can be found here: https://github.com/aadarshram/lerobot/blob/MultiTask/src/lerobot/scripts/generate_MetaWorld_datasets.py Homepage: [More Information Needed] Paper: [More Information Needed] License: apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/metaworld_mt10.tabularrobotics10K<n<100K0 likes492 downloads9mo agoHugging Face06aadityaubhat /synthetic-emotions Synthetic Emotions Dataset Overview Synthetic Emotions is a video dataset of AI-generated human emotions created using OpenAI Sora. It features short (5-sec, 480p, 9:16) videos depicting diverse individuals expressing emotions like happiness, sadness, anger, fear, surprise, and more. This dataset is ideal for emotion recognition, facial expression analysis, affective computing, and AI-human interaction research. Dataset Details Total Videos: 100 Video Format:… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/synthetic-emotions.textvideo-classificationn<1K6 likes463 downloads2y agoHugging Face07aadityaJagdale /cleaned_legal_dataset Indian Legal Judgments Data Analysis Dataset Dataset Overview Please check out the "Files and Versions" tab for the complete dataset structure and access all files. This dataset contains a comprehensive collection of Indian legal judgments that have been systematically analyzed, cleaned, and structured for advanced legal research and AI applications. The dataset comprises two distinct versions: processed/cleaned judgments and raw scraped data, providing researchers with… See the full description on the dataset page: https://huggingface.co/datasets/aadityaJagdale/cleaned_legal_dataset.0 likes337 downloads1y agoHugging Face08Aadhavan /so101_bio_finalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 50, "total_frames": 22326, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aadhavan/so101_bio_final.tabularrobotics10K<n<100K1 likes303 downloads1y agoHugging Face09anaonymous-aad /GenQA Dataset Card for "GenQA" More Information needed text1M<n<10M0 likes299 downloads2y agoHugging Face10AadiBhatia /R-Star-Distillation-Backupstabular100K<n<1M0 likes274 downloads2mo agoHugging Face11AaditD /multilingual_rksimage1K<n<10K1 likes271 downloads2y agoHugging Face12anaonymous-aad /GenQA_mmlu Dataset Card for "GenQA_mmlu" More Information needed text1M<n<10M2 likes255 downloads2y agoHugging Face13AaditTennis /ATP-Tennis-Matches-Dataset-2015-to-2025 ATP Tennis Matches Dataset (2015-2025) Description Comprehensive dataset of ATP Tour tennis matches from 2015 to 2025. Contains detailed match statistics, player information, and tournament data scraped from the official ATP Tour website. This dataset provides a complete overview of professional tennis matches over an 11-year period, suitable for sports analytics, machine learning projects, and statistical research. Dataset Structure Format: CSV… See the full description on the dataset page: https://huggingface.co/datasets/AaditTennis/ATP-Tennis-Matches-Dataset-2015-to-2025.0 likes239 downloads3mo agoHugging Face14aadarsh99 /ConverSeg ConverSeg: Conversational Image Segmentation ConverSeg is a benchmark for grounding abstract, intent-driven concepts into pixel-accurate masks. Unlike standard referring expression datasets, ConverSeg focuses on physical reasoning, affordances, and safety. Dataset Structure The dataset contains two splits: sam_seeded: 1,194 samples generated via SAM2 + VLM verification. human_annotated: 493 samples with human-drawn masks (initialized from COCO). Licensing &… See the full description on the dataset page: https://huggingface.co/datasets/aadarsh99/ConverSeg.imageimage-segmentation1K<n<10K1 likes227 downloads7mo agoHugging Face15Iceclear /AADBPhoto Aesthetics Ranking Network with Attributes and Content Adaptation Citation @inproceedings{kong2016aesthetics, title={Photo Aesthetics Ranking Network with Attributes and Content Adaptation}, author={Kong, Shu and Shen, Xiaohui and Lin, Zhe and Mech, Radomir and Fowlkes, Charless}, booktitle={ECCV}, year={2016} } image10K<n<100K2 likes226 downloads3y agoHugging Face16aadasdadasdsa /apex_enemy_detect Apex Legends Enemy Detection Dataset Dataset for detecting players in Apex Legends gameplay footage.4 921 frames — 2 classes: enemy, mate. Dataset Structure Split Images Labels train 3 937 3 937 val 984 984 images/ train/ # 3937 × .png val/ # 984 × .png labels/ # YOLO .txt, mirrors images/ dataset.yaml Annotation Format YOLO — each .txt contains one row per bounding box: <class_id> <cx> <cy> <w> <h> # normalized… See the full description on the dataset page: https://huggingface.co/datasets/aadasdadasdsa/apex_enemy_detect.imageobject-detection1K<n<10K0 likes166 downloads17d agoHugging Face17aadarsh99 /kyvo-datasets-and-codebooks Kyvo Dataset and Codebooks Details This document provides details about the dataset and codebooks provided in the kyvo-datasets-and-codebooks repository. We will provide the details about each of the folders in the repository and the contents of each folder. Data Generation Pipeline The pipeline that we follow to generate the pre-tokenized data is as follows: 3D Scenes: 3D Scene JSON --> Serialized 3D Scene --> Tokenized 3D Scene Images: Image --> VQGAN Codebook… See the full description on the dataset page: https://huggingface.co/datasets/aadarsh99/kyvo-datasets-and-codebooks.2 likes149 downloads1y agoHugging Face18Aadilgani /kashmiri-text-datasettext10K<n<100K0 likes148 downloads1y agoHugging Face19aadajinkya /python_codes_sampletext10K<n<100K2 likes138 downloads3y agoHugging Face20aadityaubhat /GPT-wiki-intro GPT Wiki Intro Overview Dataset for training models to classify human written vs GPT/ChatGPT generated text. This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics. Prompt used for generating text 200 word wikipedia style introduction on '{title}' {starter_text} where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction. Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.tabulartext-classification100K<n<1M27 likes131 downloads3y agoHugging Face21Aadhavan /so101_test1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 50, "total_frames": 22571, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aadhavan/so101_test1.tabularrobotics10K<n<100K0 likes112 downloads1y agoHugging Face22anaonymous-aad /GenQA_academic Dataset Card for "GenQA_academic" More Information needed text1M<n<10M1 likes111 downloads2y agoHugging Face23average-developer /stocks-AADHARHFC-1D-candlesn<1K0 likes105 downloads55m agoHugging Face24aadarshram /metaworld-door-open-v3 NOTE: - All expert trajectories (100% success rate) - 50 total episodes - Camera view: 3rd-person Corner2 only out of ["corner", "corner2", "corner3", "topview", "behindGripper"] - Generator script can be found here: https://github.com/aadarshram/lerobot/blob/MultiTask/src/lerobot/scripts/generate_MetaWorld_datasets.py--- license: apache-2.0 task_categories: - robotics tags: - LeRobot - metaworld - robotics - door-open-v3 configs: - config_name: default data_files: data//.parquet… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/metaworld-door-open-v3.image1K<n<10K0 likes99 downloads9mo agoHugging Face25Aadhavan /so101_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 2, "total_frames": 1748, "total_tasks":1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aadhavan/so101_test.tabularrobotics1K<n<10K0 likes89 downloads1y agoHugging Face26aadishvdotdev /sessionstabularn<1K0 likes84 downloads4mo agoHugging Face27Aadhavan /so101_bio_test1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 2, "total_frames": 755, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aadhavan/so101_bio_test1.tabularroboticsn<1K0 likes83 downloads1y agoHugging Face28aadarsh99 /ConverSeg-Training-Data ConverSeg Training Data This dataset contains image segmentation training data for ConverSeg. Each split is stored as a JSONL manifest plus zip archives of PNG images and PNG segmentation masks. Dataset Layout converseg_stage1_data/ open_vocabulary_regions_data/ open_vocabulary_regions_data.jsonl images.zip masks.zip converseg_stage2_data/ conversational_negative_data/ conversational_negative_data.jsonl images.zip masks.zip… See the full description on the dataset page: https://huggingface.co/datasets/aadarsh99/ConverSeg-Training-Data.image-segmentation100K<n<1M1 likes82 downloads2mo agoHugging Face29aadarshram /metaworld-drawer-close-v3This dataset was created using LeRobot. Dataset Description NOTE: All expert trajectories (100% success rate) 50 total episodes Camera view: 3rd-person Corner2 only out of ["corner", "corner2", "corner3", "topview", "behindGripper"] Generator script can be found here: https://github.com/aadarshram/lerobot/blob/MultiTask/src/lerobot/scripts/generate_MetaWorld_datasets.py Homepage: [More Information Needed] Paper: [More Information Needed] License: apache-2.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/metaworld-drawer-close-v3.imagerobotics1K<n<10K0 likes79 downloads9mo agoHugging Face30vk9199 /aixbitss-aadhaar-syntheticimagen<1K1 likes76 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.