CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /code_x_glue_cc_clone_detection_big_clone_bench Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench" Dataset Summary CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score. The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.tabulartext-classification1M<n<10M22 likes1k downloads3y agoHugging Face02litagin /reazon-speech-v2-clonegated Reazon Speech v2 dataset mirror Original Dataset Source Hugging Face Dataset Page: reazon-research/reazonspeech Project Page: Reazon Research License This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction: TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.audioautomatic-speech-recognition10K<n<100K12 likes747 downloads2y agoHugging Face03SynDataLab-EN /tts-pretrain-clones-3m-mos TTS Pretrain Clones (3M) — with DNSMOS This is SynDataLab/tts-pretrain-clones-3m with an added per-utterance dnsmos column (DNSMOS P.835 OVRL score, float32), computed with the sig_bak_ovr.onnx model. 2,967,779 clone utterances across 2971 English speakers. Sample rate: 44.1 kHz, WAV in Parquet dnsmos: overall MOS quality estimate per utterance (higher is better) Generated by echo-tts synthesizing English text on speaker latents derived from Qwen3-TTS VoiceDesign base speakers.… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/tts-pretrain-clones-3m-mos.audiotext-to-speech1M<n<10M0 likes698 downloads4mo agoHugging Face04SynDataLab-JA /irodori-clones-3m-v2-no-emoji Irodori TTS Clones v2 (3.29M) 3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2, using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2. 329 clones per ref voice, each with a unique Japanese conversational text. Companion refs: SynData-2/irodori-refs-10k-v2. Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.audiotext-to-speech1M<n<10M0 likes566 downloads4mo agoHugging Face05google /code_x_glue_cc_clone_detection_poj104 Dataset Card for "code_x_glue_cc_clone_detection_poj_104" Dataset Summary CodeXGLUE Clone-detection-POJ-104 dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-POJ-104 Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP score. We use POJ-104 dataset on this task. Supported Tasks and Leaderboards document-retrieval: The… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_poj104.texttext-retrieval10K<n<100K9 likes468 downloads3y agoHugging Face06nevillathiya001 /clone-CoderForge-Preview CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents CoderForge-Preview is the largest open test-verified coding agent dataset. Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters. Limitations Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/nevillathiya001/clone-CoderForge-Preview.text100K<n<1M0 likes295 downloads7mo agoHugging Face07PoolC /1-fold-clone-detection-600k-5foldtabular1M<n<10M4 likes228 downloads4y agoHugging Face08jerpint /vox-cloned-data CommonVoice Clones This dataset consists of recordings taken from the CommonVoice english dataset. Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text. TTS Models We use the following high-scoring models from the TTS leaderboard: playHT metavoice StyleTTSv2 XttsV2 Model Comparisons To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.audiotext-to-speech1K<n<10K0 likes196 downloads2y agoHugging Face09thejorseman /CloneHeroDatasetCharts Clone Hero Charts Dataset Dataset Description Tokenized Clone Hero charts with beat-level audio conditioning. Each row is one instrument track (guitar / bass / drums) from one song. Feature Value Total rows (train) 43,665 Parquet shards 1753 Audio: MERT embeddings Yes [num_beats, 768] Audio: log-mel frames Yes [num_beats, 32, 128] Dataset Structure Data Fields Column Type Description song_id string MD5 hash of… See the full description on the dataset page: https://huggingface.co/datasets/thejorseman/CloneHeroDatasetCharts.tabularaudio-to-audio10K<n<100K0 likes143 downloads5mo agoHugging Face10PoolC /4-fold-clone-detection-600k-5foldtabular1M<n<10M0 likes140 downloads4y agoHugging Face11ThatDev /Afrivoice_Kinyarwanda_ASR_cloneaudio100K<n<1M0 likes116 downloads5mo agoHugging Face12elizabethkunz /SQUADDS_test_clone THIS IS A CLONE AND IS NOT THE OFFICIAL SQUADDS DB. SQuADDS_DB - a Superconducting Qubit And Device Design and Simulation Database The SQuADDS (Superconducting Qubit And Device Design and Simulation) Database Project is an open-source resource aimed at advancing research in superconducting quantum device designs. It provides a robust workflow for generating and simulating superconducting quantum device designs, facilitating the accurate prediction of Hamiltonian… See the full description on the dataset page: https://huggingface.co/datasets/elizabethkunz/SQUADDS_test_clone.image1K<n<10K0 likes109 downloads2y agoHugging Face13Holmeister /AAID-clone Citation Information Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024, August). Emollms: A series of emotional large language models and annotation tools for comprehensive affective analysis. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5487-5496). text100K<n<1M0 likes94 downloads2y agoHugging Face14zhiqings /dromedary-65b-verbose-clone-v0 Dataset Card for Dromedary-Verbose-Clone (65b-v0) Repository: https://github.com/IBM/Dromedary Authors' Note: The Self-Align data contain a plethora of partial responses. Therefore, it is advised to refrain from appending the <eos> or </s> token to the model responses for supervised fine-tuning (SFT). Instead, it is recommended to substitute "\n\n### User" (Dromedary's eos token) with your own end-of-response token. Dataset Summary Dromedary-Verbose-Clone is a… See the full description on the dataset page: https://huggingface.co/datasets/zhiqings/dromedary-65b-verbose-clone-v0.text100K<n<1M11 likes77 downloads3y agoHugging Face15Ashsinha1 /hf-spaces-clones Unmodified copies among Hugging Face Spaces 104,791 public Spaces that are copies of another Space, grouped into the 25,183 source histories they came from. Derived from a full census of the 1,458,692 public Spaces taken in August 2026. This is not a similarity score. Members of a family are byte-identical repositories, and the check below is what establishes that. How a copy is identified A Space is a git repository. Pushing an existing history into a new… See the full description on the dataset page: https://huggingface.co/datasets/Ashsinha1/hf-spaces-clones.tabular100K<n<1M0 likes68 downloads22d agoHugging Face16semeru /Code-Code-CloneDetection-BigCloneBench Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-BigCloneBench in Semeru CodeXGLUE -- Clone Detection (BCB) Task Definition Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score. Updates… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-BigCloneBench.text1M<n<10M0 likes61 downloads4y agoHugging Face17SynDataLab-EN /qwen-clones-4m-en4M conversational English speech clips generated with Qwen3-TTS-12Hz-1.7B-Base. Reference speakers: SynData-2/qwen-ref-speakers-4k-en audio1M<n<10M0 likes57 downloads4mo agoHugging Face18semeru /Code-Code-CloneDetection-POJ104 Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-POJ-104 in Semeru CodeXGLUE -- Clone Detection (POJ-104) Task Definition Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP@R score. MAP@R is defined as the mean of… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-POJ104.text10K<n<100K2 likes55 downloads4y agoHugging Face19benkum /neutral-batch-voice-cloneaudio10K<n<100K1 likes55 downloads5mo agoHugging Face20SynDataLab-EN /echo-clones-4m-en echo-clones-4m-en ~4 M English TTS clone utterances generated with EchoTTS (jordand/echo-tts-base). Sample rate: 44 100 Hz, 16-bit PCM WAV stored in Parquet Speakers: 4 000 reference speakers (spk_0000-spk_3999) Text bucketing: quip (<=100 chars), mid (100-300), ramble (300-420) Speaker assignment: round-robin -- text[i] -> spk_{i % 4000} Companion datasets Reference speakers: SynData-2/echo-ref-speakers-4k-en -- the 4 000 reference WAVs used as speaker… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/echo-clones-4m-en.audiotext-to-speech1M<n<10M1 likes53 downloads4mo agoHugging Face21LALM-emotional-vulnerability /cosyvoice-clone LALM Emotional Vulnerability Dataset Overview This dataset contains synthesized malicious speech instructions across multiple emotions and intensity levels to evaluate the safety responsiveness of Large Audio-Language Models (LALMs). The dataset aims to examine how speaker emotion and intensity influence the safety and robustness of AI responses. Dataset Composition Total samples: 8,320 Emotion categories: Neutral: 520 samples Angry: 1560 samples… See the full description on the dataset page: https://huggingface.co/datasets/LALM-emotional-vulnerability/cosyvoice-clone.audio1K<n<10K0 likes51 downloads10mo agoHugging Face22AmanPriyanshu /clone-of-gretel-financial-risk-analysis-v1 ⚠️🔴 IMPORTANT NOTICE 🔴⚠️ This dataset is directly cloned from gretelai/gretel-financial-risk-analysis-v1 on Hugging Face. No modifications have been made to the original dataset, it is only for archival. gretelai/gretel-financial-risk-analysis-v1 This dataset contains synthetic financial risk analysis text generated using differential privacy guarantees, trained on 14,306 SEC (10-K, 10-Q, and 8-k) filings from 2023-2024. The dataset is designed for training models to extract… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/clone-of-gretel-financial-risk-analysis-v1.tabulartext-classification1K<n<10K0 likes49 downloads2y agoHugging Face23Emmylahot12 /cloneaudion<1K0 likes49 downloads1y agoHugging Face24clonelabs /MultiOS-Trajectory-Sample MultiOS-Trajectory-Sample Human-collected, human-annotated multi-platform trajectory dataset in OWAMcap format. Structure windows/ # 5 Windows desktop tasks macos/ # TBD linux/ # TBD android/ # 10 Android mobile tasks ios/ # TBD Each task consists of 3 files: .mcap — Input events (mouse/keyboard/touch) + screen metadata .mkv — Screen recording video .json — Task instruction + human annotations (subgoal-level) Visualize Use the dataset… See the full description on the dataset page: https://huggingface.co/datasets/clonelabs/MultiOS-Trajectory-Sample.textn<1K0 likes49 downloads6mo agoHugging Face25LALM-emotional-damage /cosyvoice-cloneaudio1K<n<10K0 likes47 downloads1y agoHugging Face26Deinvadda /gta-vi-clonetextn<1K0 likes47 downloads18d agoHugging Face27PoolC /5-fold-clone-detection-600k-5foldtabular1M<n<10M0 likes45 downloads4y agoHugging Face28PoolC /3-fold-clone-detection-600k-5foldtabular1M<n<10M0 likes44 downloads4y agoHugging Face29launch-calcium /netflix-clone-imagesimagen<1K1 likes43 downloads2mo agoHugging Face30IAMCB /elise-clone Custom Elise-like TTS Dataset Converted on 2025-06-18T05:56:51Z. Samples: 1043 Format : 10-s clips with text transcription (like MrDragonFox/Elise) Structure column type description audio audio 24kHz mono wav clip text string transcription audio1K<n<10K0 likes40 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.