datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_cc_clone_detection_big_clone_bench
Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench"
Dataset Summary
CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.reazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.tts-pretrain-clones-3m-mos
TTS Pretrain Clones (3M) — with DNSMOS
This is SynDataLab/tts-pretrain-clones-3m
with an added per-utterance dnsmos column (DNSMOS P.835 OVRL score, float32),
computed with the sig_bak_ovr.onnx model.
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
dnsmos: overall MOS quality estimate per utterance (higher is better)
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers.… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/tts-pretrain-clones-3m-mos.irodori-clones-3m-v2-no-emoji
Irodori TTS Clones v2 (3.29M)
3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2,
using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2.
329 clones per ref voice, each with a unique Japanese conversational text.
Companion refs: SynData-2/irodori-refs-10k-v2.
Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.code_x_glue_cc_clone_detection_poj104
Dataset Card for "code_x_glue_cc_clone_detection_poj_104"
Dataset Summary
CodeXGLUE Clone-detection-POJ-104 dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-POJ-104
Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP score.
We use POJ-104 dataset on this task.
Supported Tasks and Leaderboards
document-retrieval: The… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_poj104.clone-CoderForge-Preview
CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents
CoderForge-Preview is the largest open test-verified coding agent dataset.
Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters.
Limitations
Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/nevillathiya001/clone-CoderForge-Preview.1-fold-clone-detection-600k-5foldvox-cloned-data
CommonVoice Clones
This dataset consists of recordings taken from the CommonVoice english dataset.
Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text.
TTS Models
We use the following high-scoring models from the TTS leaderboard:
playHT
metavoice
StyleTTSv2
XttsV2
Model Comparisons
To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.CloneHeroDatasetCharts
Clone Hero Charts Dataset
Dataset Description
Tokenized Clone Hero charts with beat-level audio conditioning.
Each row is one instrument track (guitar / bass / drums) from one song.
Feature
Value
Total rows (train)
43,665
Parquet shards
1753
Audio: MERT embeddings
Yes [num_beats, 768]
Audio: log-mel frames
Yes [num_beats, 32, 128]
Dataset Structure
Data Fields
Column
Type
Description
song_id
string
MD5 hash of… See the full description on the dataset page: https://huggingface.co/datasets/thejorseman/CloneHeroDatasetCharts.4-fold-clone-detection-600k-5foldAfrivoice_Kinyarwanda_ASR_cloneSQUADDS_test_clone
THIS IS A CLONE AND IS NOT THE OFFICIAL SQUADDS DB.
SQuADDS_DB - a Superconducting Qubit And Device Design and Simulation Database
The SQuADDS (Superconducting Qubit And Device Design and Simulation) Database Project is an open-source resource aimed at advancing research in superconducting quantum device designs. It provides a robust workflow for generating and simulating superconducting quantum device designs, facilitating the accurate prediction of Hamiltonian… See the full description on the dataset page: https://huggingface.co/datasets/elizabethkunz/SQUADDS_test_clone.AAID-clone
Citation Information
Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024, August). Emollms: A series of emotional large language models and annotation tools
for comprehensive affective analysis. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5487-5496).
dromedary-65b-verbose-clone-v0
Dataset Card for Dromedary-Verbose-Clone (65b-v0)
Repository: https://github.com/IBM/Dromedary
Authors' Note: The Self-Align data contain a plethora of partial responses. Therefore, it is advised to refrain from appending the <eos> or </s> token to the model responses for supervised fine-tuning (SFT). Instead, it is recommended to substitute "\n\n### User" (Dromedary's eos token) with your own end-of-response token.
Dataset Summary
Dromedary-Verbose-Clone is a… See the full description on the dataset page: https://huggingface.co/datasets/zhiqings/dromedary-65b-verbose-clone-v0.hf-spaces-clones
Unmodified copies among Hugging Face Spaces
104,791 public Spaces that are copies of another Space, grouped into the
25,183 source histories they came from. Derived from a full census of the
1,458,692 public Spaces taken in August 2026.
This is not a similarity score. Members of a family are byte-identical
repositories, and the check below is what establishes that.
How a copy is identified
A Space is a git repository. Pushing an existing history into a new… See the full description on the dataset page: https://huggingface.co/datasets/Ashsinha1/hf-spaces-clones.Code-Code-CloneDetection-BigCloneBench
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-BigCloneBench in Semeru
CodeXGLUE -- Clone Detection (BCB)
Task Definition
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
Updates… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-BigCloneBench.qwen-clones-4m-en4M conversational English speech clips generated with Qwen3-TTS-12Hz-1.7B-Base.
Reference speakers: SynData-2/qwen-ref-speakers-4k-en
Code-Code-CloneDetection-POJ104
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-POJ-104 in Semeru
CodeXGLUE -- Clone Detection (POJ-104)
Task Definition
Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP@R score. MAP@R is defined as the mean of… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-POJ104.neutral-batch-voice-cloneecho-clones-4m-en
echo-clones-4m-en
~4 M English TTS clone utterances generated with
EchoTTS (jordand/echo-tts-base).
Sample rate: 44 100 Hz, 16-bit PCM WAV stored in Parquet
Speakers: 4 000 reference speakers (spk_0000-spk_3999)
Text bucketing: quip (<=100 chars), mid (100-300), ramble (300-420)
Speaker assignment: round-robin -- text[i] -> spk_{i % 4000}
Companion datasets
Reference speakers: SynData-2/echo-ref-speakers-4k-en -- the 4 000 reference WAVs used as speaker… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/echo-clones-4m-en.cosyvoice-clone
LALM Emotional Vulnerability Dataset
Overview
This dataset contains synthesized malicious speech instructions across multiple emotions and intensity levels to evaluate the safety responsiveness of Large Audio-Language Models (LALMs). The dataset aims to examine how speaker emotion and intensity influence the safety and robustness of AI responses.
Dataset Composition
Total samples: 8,320
Emotion categories:
Neutral: 520 samples
Angry: 1560 samples… See the full description on the dataset page: https://huggingface.co/datasets/LALM-emotional-vulnerability/cosyvoice-clone.clone-of-gretel-financial-risk-analysis-v1
⚠️🔴 IMPORTANT NOTICE 🔴⚠️
This dataset is directly cloned from gretelai/gretel-financial-risk-analysis-v1 on Hugging Face. No modifications have been made to the original dataset, it is only for archival.
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated using differential privacy guarantees, trained on 14,306 SEC (10-K, 10-Q, and 8-k) filings from 2023-2024. The dataset is designed for training models to extract… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/clone-of-gretel-financial-risk-analysis-v1.cloneMultiOS-Trajectory-Sample
MultiOS-Trajectory-Sample
Human-collected, human-annotated multi-platform trajectory dataset in OWAMcap format.
Structure
windows/ # 5 Windows desktop tasks
macos/ # TBD
linux/ # TBD
android/ # 10 Android mobile tasks
ios/ # TBD
Each task consists of 3 files:
.mcap — Input events (mouse/keyboard/touch) + screen metadata
.mkv — Screen recording video
.json — Task instruction + human annotations (subgoal-level)
Visualize
Use the dataset… See the full description on the dataset page: https://huggingface.co/datasets/clonelabs/MultiOS-Trajectory-Sample.cosyvoice-clonegta-vi-clone5-fold-clone-detection-600k-5fold3-fold-clone-detection-600k-5foldnetflix-clone-imageselise-clone
Custom Elise-like TTS Dataset
Converted on 2025-06-18T05:56:51Z.
Samples: 1043
Format : 10-s clips with text transcription (like MrDragonFox/Elise)
Structure
column
type
description
audio
audio
24kHz mono wav clip
text
string
transcription
