CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saberzl /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please check… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/SID_Set.imagetext-to-image100K<n<1M15 likes11k downloads1y agoHugging Face02vidore /colpali_train_set Dataset Description This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%). Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages. Dataset #examples (query-page pairs) Language DocVQA 39… See the full description on the dataset page: https://huggingface.co/datasets/vidore/colpali_train_set.imagedocument-question-answering100K<n<1M93 likes5.5k downloads1y agoHugging Face03saberzl /So-Fake-Set Dataset Card for So-Fake-Set Dataset Summary We provide So-Fake-Set, A large-scale, diverse dataset tailored for social media image forgery detection! Please check our website to explore more visual results. Dataset Structure "image" (Image): Input images, including real, full_synthetic, and tampered images. "mask" (Image): Binary mask highlighting manipulated regions in tampered images. "label" (str): Classification category. "generator" (str): The… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/So-Fake-Set.image1M<n<10M11 likes4.2k downloads11mo agoHugging Face04Prompt48 /AIME_Problem_Set_1983-2024tabularn<1K0 likes4.2k downloads2y agoHugging Face05community-datasets /setimes Dataset Card for SETimes – A Parallel Corpus of English and South-East European Languages Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances Here are some examples of questions and facts: Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/setimes.texttranslation1M<n<10M2 likes4.1k downloads2y agoHugging Face06allenai /preference-test-sets Preference Test Sets Very few preference datasets have heldout test sets for validation of reward model accuracy results. In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation. Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation. Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.textsummarization10K<n<100K28 likes3.5k downloads3y agoHugging Face07JamalLee /Omni-Fake-SET Omni-Fake-SET Omni-Fake-SET is the in-distribution split of Omni-Fake, a unified multimodal deepfake dataset for social-media forensics. It covers image, audio, video, and audio–video talking-head (AV-TH) modalities. Each modality uses the same three-way label space: real, fully synthetic, and tampered. Pair with the held-out benchmark Omni-Fake-OOD for out-of-distribution evaluation. Paper: arXiv:2605.01638 Project page: Omni-Fake License: CC-BY-4.0 Video (hybrid… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-SET.audioimage-classification1M<n<10M5 likes3k downloads3mo agoHugging Face08AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.9k downloads3mo agoHugging Face09SetFit /tweet_eval_stance_abortiontextn<1K0 likes1.8k downloads4y agoHugging Face10MBZUAI /Omni-Setsgated Omni-Sets A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities. 590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets Overview Config Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.audio100K<n<1M0 likes1.6k downloads22d agoHugging Face11Cooolder /SCOPE-OOD-set SCOPE-60K-OOD: Out-of-Distribution LLM Routing Dataset Dataset Description SCOPE-60K-OOD is an out-of-distribution (OOD) evaluation dataset for LLM routing systems. It contains evaluation results from 5 frontier language models that were not seen during training, designed to test the generalization capabilities of routing methods. Authors Qi Cao - UC San Diego, PXie Lab Shuhao Zhang - UC San Diego, PXie Lab Affiliation University of California, San… See the full description on the dataset page: https://huggingface.co/datasets/Cooolder/SCOPE-OOD-set.tabulartext-classification1K<n<10K0 likes1.6k downloads8mo agoHugging Face12InsultedByMathematics /diffbir-mixed-setsimage1M<n<10M0 likes1k downloads1y agoHugging Face13AbstractPhil /diffusion-pretrain-set-ft1-1024 diffusion-pretrain-set-ft1-1024 1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1. WARNING MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS. THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES. PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING. Thank you, good luck my friends. Details Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.image1M<n<10M0 likes966 downloads4mo agoHugging Face14argilla /banking_sentiment_setfit Dataset Card for "banking_sentiment_setfit" More Information needed textn<1K2 likes937 downloads4y agoHugging Face15maxseats /aihub-464-preprocessed-680GB-set-52audio10K<n<100K0 likes910 downloads2y agoHugging Face16ScottQu /Instruction_filtered_set Dataset Card for "Instruction_filtered_set" More Information needed text1M<n<10M0 likes731 downloads3y agoHugging Face17Pclanglais /gutenberg_settabular1M<n<10M0 likes717 downloads2y agoHugging Face18WissMah /lebanese_aug_setimage10K<n<100K0 likes672 downloads11mo agoHugging Face19RAID-techjam /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.imagetext-to-image100K<n<1M0 likes654 downloads26d agoHugging Face20facebook /emu_edit_test_set Dataset Card for the Emu Edit Test Set Dataset Summary To create a benchmark for image editing we first define seven different categories of potential image editing operations: background alteration (background), comprehensive image changes (global), style alteration (style), object removal (remove), object addition (add), localized modifications (local), and color/texture alterations (texture). Then, we utilize the diverse set of input images from the MagicBrush… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set.image1K<n<10K47 likes641 downloads3y agoHugging Face21RoboCOIN /R1_Lite_tea_service_table_settinggated R1_Lite_tea_service_table_setting 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_tea_service_table_setting.tabularrobotics100K<n<1M0 likes633 downloads9mo agoHugging Face22jwengr /So-Fake-Set-Resized-224image1M<n<10M0 likes628 downloads6mo agoHugging Face23Beastarz /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please… See the full description on the dataset page: https://huggingface.co/datasets/Beastarz/SID_Set.imagetext-to-image100K<n<1M0 likes595 downloads26d agoHugging Face24nomic-ai /colpali_train_set_split_by_sourceimage100K<n<1M2 likes569 downloads2y agoHugging Face25eddmpython /cleangov-local-settlements 지방재정365 결산 통계 Open API (세입·세출결산, 재무제표, 지방세, 지역통합재정통계, 공공시설·청사·채무) 지방재정365 재정데이터개방 허브의 "결산" 분류 33 서비스. 세출결산(기능별·성질별·회계별·구조별 단체별), 세입결산(재원별·성질별), 기금결산, 교육비특별회계 결산, 투자적경비 순계, 재무제표(재정상태표·통합재정운영표·순자산변동표·복식부기 수익·비용·자산·부채), 지방세(징수율·세목별 비중·세수신장률·체납 누계), 지역통합재정통계 (세입·세출·자산·부채·인건비·업무추진비·행사경비 비율), 공공시설운영현황, 청사면적, 채무현황. 자치단체 재정의 결산 기준 정본이며 FISCAL-LOC-002(세부사업별 세출 XLSX) 보다 집계 수준이 높고 분류 축이 다양하다. 비교군·기관 개요 화면의 결산 수치를 여기서 낸다. 출처: https://www.lofin365.go.kr/portal/LF5100000.do 이용 조건: 허브 명세 이용조건… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-settlements.text1M<n<10M0 likes527 downloads17d agoHugging Face26maxseats /aihub-464-preprocessed-680GB-set-53audio10K<n<100K0 likes525 downloads2y agoHugging Face27maxseats /aihub-464-preprocessed-680GB-set-56audio10K<n<100K0 likes522 downloads2y agoHugging Face28WhissleAI /Meta_STT_HI_Set1 Meta Speech Recognition Hindi Dataset (Set 1) This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources. Dataset Sources and Credits This dataset combines samples from the following sources: AI4Bharat Indic Speech Dataset Source: https://ai4bharat.org/indic-speech-dataset License: CC-BY 4.0 Citation: Please cite the original paper if you use this data Common Voice Hindi Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.audioautomatic-speech-recognition100K<n<1M0 likes469 downloads1y agoHugging Face29microsoft /bing_coronavirus_query_set Dataset Card for BingCoronavirusQuerySet Dataset Summary Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020 example: load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30") You can also load the data by country by using queries_by="country". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.tabulartext-classification100K<n<1M1 likes455 downloads3y agoHugging Face30maxseats /aihub-464-preprocessed-680GB-set-57audio10K<n<100K0 likes444 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.