CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amanuelbyte /merged_speech_datasetaudio1M<n<10M0 likes764 downloads7mo agoHugging Face02raghavendrad60 /vqa_plant-disease-classification-merged-datasetimage10K<n<100K1 likes474 downloads2y agoHugging Face03BlossomsAI /merged_vietnamese_instruction_datasettext1M<n<10M0 likes400 downloads1y agoHugging Face04Ba2han /merged-datasets_11-12 Merged Datasets 11-12 This dataset is a compilation of various sources filtered for length and quality. Statistics Dataset Language Count Mean Len Median Len Max Len Min Len Ba2han/synth-tr Turkish 33,416 2874 2692 17714 400 Ba2han/synth-2m-v2 Turkish 2,065,969 2894 2718 110664 358 Ba2han/filtered-cosmos Turkish 785,340 4987 4646 11000 1000 Ba2han/fineweb2-filtered-tr Turkish 5,896,067 5853 5265 15000 304 Ba2han/finePDF-filtered-tr Turkish 124,223 8294… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/merged-datasets_11-12.text10M<n<100M0 likes391 downloads10mo agoHugging Face05kshitijthakkar /loggenix-merged-dataset-v1tabular1M<n<10M0 likes324 downloads1y agoHugging Face06grapleonly /merged-yolo-datasetimage100K<n<1M0 likes266 downloads5mo agoHugging Face07madoss /merged-bambara-dioula-datasetaudio10K<n<100K0 likes209 downloads2mo agoHugging Face08bagasshw /indo-merged-dataset-2 Dataset Card for "indo-merged-dataset-2" More Information needed audio10K<n<100K0 likes208 downloads1y agoHugging Face09wheresmyhair /merged_sft_datasettext1M<n<10M1 likes178 downloads1y agoHugging Face10equal-ai /merged-hindi-audio-datasetaudio10K<n<100K0 likes171 downloads7mo agoHugging Face11Danielrahmai1991 /merged_fin_dataset_v2text1M<n<10M0 likes161 downloads2y agoHugging Face12MosRat /gex_dataset_mergedimageimage-to-text1M<n<10M0 likes117 downloads1y agoHugging Face13OBY632 /merged-bambara-dioula-datasetaudio10K<n<100K0 likes108 downloads6mo agoHugging Face14todenthal /merged_financial_datasettext1M<n<10M0 likes103 downloads2y agoHugging Face15PharynxAI /merged_english_accent_datasetaudio10K<n<100K0 likes99 downloads1y agoHugging Face16yufan /Preference_Dataset_Merged Dataset Overview This Dataset consists of the following open-sourced preference dataset Arena Human Preference Anthropic HH MT-Bench Human Judgement Ultra Feedback Tulu3 Preference Dataset Skywork-Reward-Preference-80K-v0.2 Cleaning Cleaning Method 1: Only keep the following Language using FastText language detection(EN/DE/ES/ZH/IT/JA/FR) Cleaning Method 2: Remove duplicates to ensure each prompt appears only once Cleaning Method 3: Remove datasets where… See the full description on the dataset page: https://huggingface.co/datasets/yufan/Preference_Dataset_Merged.text100K<n<1M0 likes93 downloads2y agoHugging Face17KYAGABA /Merged_Luo_Datasetaudio10K<n<100K0 likes86 downloads2y agoHugging Face18Reza2kn /persian-ocr-community-dataset-layout-mergedtext1K<n<10K0 likes85 downloads2mo agoHugging Face19kshitijthakkar /loggenix-merged-dataset ============================================================ TOKEN STATISTICS ANALYSIS 📊 OVERALL DATASET STATISTICS ──────────────────────────────────────── Total samples: 2,419,111 Total tokens in dataset: 2,052,945,533 Average tokens per sample: 848.64 📈 TOKEN DISTRIBUTION ──────────────────────────────────────── Min tokens: 21 Max tokens: 142,971 Median tokens: 534.00 Standard deviation: 1047.89 📊 TOKEN PERCENTILES ──────────────────────────────────────── 5th percentile: 83… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/loggenix-merged-dataset.tabular1M<n<10M0 likes77 downloads1y agoHugging Face20Gunulhona /medical_qa_dataset_mergedtext1M<n<10M0 likes74 downloads10mo agoHugging Face21amanuelbyte /merged_african_speech_datasetaudio10K<n<100K0 likes59 downloads7mo agoHugging Face22Sai-2007 /merged_chat_datasettext1M<n<10M1 likes59 downloads2mo agoHugging Face23Newton2676 /STRIX_3_dataset_merged mad-esc50-drone-audio Dataset audio fusionné à partir de 3 sources, normalisé pour la classification audio (colonnes : audio, label, source_dataset, split). Sources et attributions MAD (Military Audio Dataset) — Kim, Yoon & Jung, 2024, Scientific Data. CC BY 4.0. https://www.kaggle.com/datasets/junewookim/mad-dataset-military-audio-dataset ESC-50 — Karol J. Piczak. CC BY-NC 3.0 (usage non commercial). https://github.com/karolpiczak/ESC-50 DroneAudioDataset —… See the full description on the dataset page: https://huggingface.co/datasets/Newton2676/STRIX_3_dataset_merged.audio10K<n<100K0 likes58 downloads3mo agoHugging Face24Danielrahmai1991 /merged_fin_dataset_v1text1M<n<10M0 likes52 downloads2y agoHugging Face25Mithilss /MindGamesArena-Merged-Datasettabular1K<n<10K0 likes51 downloads1y agoHugging Face26TheRealOKAI /arabic_ocr_merged_datasetimagen<1K0 likes50 downloads10mo agoHugging Face27BoghdadyJR /Merged_datasetimage10K<n<100K0 likes44 downloads1y agoHugging Face28csdhbg /merged-PMEmo2019-lyric-audio-valence-new-regressor-test-datasettextn<1K0 likes43 downloads1y agoHugging Face29yobro4619 /merged_datasetimage1K<n<10K0 likes42 downloads1y agoHugging Face30sabin1234 /Merged_Nepali_Health_FAQ_Dataset Merged Nepali Health FAQ Dataset (Multi-Source) Overview This dataset (merged_nepali_sharegpt_after_removing_11_groups_keep_ids.jsonl) is a merged collection of 65 instruction-following conversation pairs in Nepali, combining Q&A content from 8 distinct real-world Nepali health institutions and organizations into a single ShareGPT-style file. Each record is a single-turn human↔gpt exchange: a Nepali-language question followed by a factual Nepali-language answer.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Merged_Nepali_Health_FAQ_Dataset.textn<1K0 likes41 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.