datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_cc_clone_detection_big_clone_bench
Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench"
Dataset Summary
CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.e621_sample_clonexetAll images of all ratings from e621.net from the date it was generated, at sample resolution where possible.
This includes the following additional metadata:
post ID
created at
updated at
tags (stored as IDs you can cross-reference from an e621 tags dump)
rating (0 = safe, 1 = questionable, 2 = explicit)
favorite count
comment count
up score
down score
Note that this dataset excludes images that are, at the time of scraping:
pending
tagged with tags indicating that it is illegal to possess… See the full description on the dataset page: https://huggingface.co/datasets/launch-calcium/e621_sample_clonexet.irodori-clones-3m-v2-no-emoji
Irodori TTS Clones v2 (3.29M)
3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2,
using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2.
329 clones per ref voice, each with a unique Japanese conversational text.
Companion refs: SynData-2/irodori-refs-10k-v2.
Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.code_x_glue_cc_clone_detection_poj104
Dataset Card for "code_x_glue_cc_clone_detection_poj_104"
Dataset Summary
CodeXGLUE Clone-detection-POJ-104 dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-POJ-104
Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP score.
We use POJ-104 dataset on this task.
Supported Tasks and Leaderboards
document-retrieval: The… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_poj104.clone-CoderForge-Preview
CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents
CoderForge-Preview is the largest open test-verified coding agent dataset.
Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters.
Limitations
Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/nevillathiya001/clone-CoderForge-Preview.clonex1-fold-clone-detection-600k-5foldCloneHeroDatasetCharts
Clone Hero Charts Dataset
Dataset Description
Tokenized Clone Hero charts with beat-level audio conditioning.
Each row is one instrument track (guitar / bass / drums) from one song.
Feature
Value
Total rows (train)
43,665
Parquet shards
1753
Audio: MERT embeddings
Yes [num_beats, 768]
Audio: log-mel frames
Yes [num_beats, 32, 128]
Dataset Structure
Data Fields
Column
Type
Description
song_id
string
MD5 hash of… See the full description on the dataset page: https://huggingface.co/datasets/thejorseman/CloneHeroDatasetCharts.4-fold-clone-detection-600k-5foldAfrivoice_Kinyarwanda_ASR_cloneAAID-clone
Citation Information
Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024, August). Emollms: A series of emotional large language models and annotation tools
for comprehensive affective analysis. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5487-5496).
clone_han5i5j1986_eval_koch_lego_2024-10-18-01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 10,
"total_frames": 3438,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AdleBens/clone_han5i5j1986_eval_koch_lego_2024-10-18-01.qwen-clones-4m-en4M conversational English speech clips generated with Qwen3-TTS-12Hz-1.7B-Base.
Reference speakers: SynData-2/qwen-ref-speakers-4k-en
neutral-batch-voice-cloneecho-clones-4m-en
echo-clones-4m-en
~4 M English TTS clone utterances generated with
EchoTTS (jordand/echo-tts-base).
Sample rate: 44 100 Hz, 16-bit PCM WAV stored in Parquet
Speakers: 4 000 reference speakers (spk_0000-spk_3999)
Text bucketing: quip (<=100 chars), mid (100-300), ramble (300-420)
Speaker assignment: round-robin -- text[i] -> spk_{i % 4000}
Companion datasets
Reference speakers: SynData-2/echo-ref-speakers-4k-en -- the 4 000 reference WAVs used as speaker… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/echo-clones-4m-en.cosyvoice-clone
LALM Emotional Vulnerability Dataset
Overview
This dataset contains synthesized malicious speech instructions across multiple emotions and intensity levels to evaluate the safety responsiveness of Large Audio-Language Models (LALMs). The dataset aims to examine how speaker emotion and intensity influence the safety and robustness of AI responses.
Dataset Composition
Total samples: 8,320
Emotion categories:
Neutral: 520 samples
Angry: 1560 samples… See the full description on the dataset page: https://huggingface.co/datasets/LALM-emotional-vulnerability/cosyvoice-clone.clone-of-gretel-financial-risk-analysis-v1
⚠️🔴 IMPORTANT NOTICE 🔴⚠️
This dataset is directly cloned from gretelai/gretel-financial-risk-analysis-v1 on Hugging Face. No modifications have been made to the original dataset, it is only for archival.
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated using differential privacy guarantees, trained on 14,306 SEC (10-K, 10-Q, and 8-k) filings from 2023-2024. The dataset is designed for training models to extract… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/clone-of-gretel-financial-risk-analysis-v1.cosyvoice-clone5-fold-clone-detection-600k-5fold3-fold-clone-detection-600k-5foldCROHME-full-clone-fixedso101_pick_place_real_clone2-fold-clone-detection-600k-5foldAAID-clone-2
Citation Information
Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024, August). Emollms: A series of emotional large language models and annotation tools
for comprehensive affective analysis. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5487-5496).
tts-zh-clone-bench-4model
tts-zh-clone-bench-4model
Chinese (Mandarin) TTS voice-cloning comparison across 4 models. Each row is a
reference->clone pair; the same index uses the same text across all models
(rows are sorted by index, so the 4 models for one text are adjacent — easy A/B).
Columns: index, model, ref_text, ref_audio, ref_dnsmos, clone_text, clone_audio, clone_dnsmos, clone_asr_cer. Audio is 24 kHz; *_dnsmos is the DNSMOS P.835 OVRL
score (higher is better, ~1-5); clone_asr_cer is the… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-clone-bench-4model.dataset2_clone_tam_thoiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ngocthuong2212/dataset2_clone_tam_thoi.llama3_orpo_dpo_clonetofu-clone2tofu-clonetts-multiling-clone-6model
tts-multiling-clone-6model
Single voice -> EN/JA/ZH across 6 open-source TTS models (20 VCTK reference speakers, cross-lingual zero-shot clone).
Model
DNSMOS EN
DNSMOS JA
DNSMOS ZH
spk_sim EN
spk_sim JA
spk_sim ZH
cosyvoice
3.31
3.22
3.18
0.59
0.51
0.53
moss
3.18
3.16
3.19
0.71
0.52
0.53
omnivoice
3.21
3.28
3.23
0.74
0.62
0.65
qwen3
3.25
3.30
3.20
0.72
0.62
0.59
voxcpm
3.09
3.06
3.17
0.64
0.48
0.43
zonos
3.21
3.39
3.37
0.52
0.35
0.36
