datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vox-cloned-data
CommonVoice Clones
This dataset consists of recordings taken from the CommonVoice english dataset.
Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text.
TTS Models
We use the following high-scoring models from the TTS leaderboard:
playHT
metavoice
StyleTTSv2
XttsV2
Model Comparisons
To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.hf-spaces-clones
Unmodified copies among Hugging Face Spaces
104,791 public Spaces that are copies of another Space, grouped into the
25,183 source histories they came from. Derived from a full census of the
1,458,692 public Spaces taken in August 2026.
This is not a similarity score. Members of a family are byte-identical
repositories, and the check below is what establishes that.
How a copy is identified
A Space is a git repository. Pushing an existing history into a new… See the full description on the dataset page: https://huggingface.co/datasets/Ashsinha1/hf-spaces-clones.clonegrug-think-cloneML4SE23_G4_Small_Clone_Benchclonejbs_clone2
Pliny Challenges README
Welcome to the Pliny X HackAPrompt Dataset! We open source all submissions to the Pliny X HackAPrompt competition track. Here's what you need to know about the data and how to work with it.
In the Pliny track, there are 12 challenges, of which 3 are image-only challenges (pliny_6_challenge, pliny_7_challenge, pliny_8_challenge).
The dataset is structured as a list of submissions, each with the following fields (see also dataset_info.json):
submission_id… See the full description on the dataset page: https://huggingface.co/datasets/Forward-T/jbs_clone2.CommonVoice17-CloneCode_clone
