datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SQUADDS_test_clone
THIS IS A CLONE AND IS NOT THE OFFICIAL SQUADDS DB.
SQuADDS_DB - a Superconducting Qubit And Device Design and Simulation Database
The SQuADDS (Superconducting Qubit And Device Design and Simulation) Database Project is an open-source resource aimed at advancing research in superconducting quantum device designs. It provides a robust workflow for generating and simulating superconducting quantum device designs, facilitating the accurate prediction of Hamiltonian… See the full description on the dataset page: https://huggingface.co/datasets/elizabethkunz/SQUADDS_test_clone.dromedary-65b-verbose-clone-v0
Dataset Card for Dromedary-Verbose-Clone (65b-v0)
Repository: https://github.com/IBM/Dromedary
Authors' Note: The Self-Align data contain a plethora of partial responses. Therefore, it is advised to refrain from appending the <eos> or </s> token to the model responses for supervised fine-tuning (SFT). Instead, it is recommended to substitute "\n\n### User" (Dromedary's eos token) with your own end-of-response token.
Dataset Summary
Dromedary-Verbose-Clone is a… See the full description on the dataset page: https://huggingface.co/datasets/zhiqings/dromedary-65b-verbose-clone-v0.Code-Code-CloneDetection-POJ104
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-POJ-104 in Semeru
CodeXGLUE -- Clone Detection (POJ-104)
Task Definition
Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP@R score. MAP@R is defined as the mean of… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-POJ104.gta-vi-clonenft-females-generated-cli-test-3-clone
nft females generated cli test 3
Automated NFT generation report
Number of characters: 10
Model: midorimae-characters-female
Guidance: 6.0
LORA scale: 0.95
Total time taken: 0h 0m 51s
Average time per prompt: 2.39 seconds
Average time per image: 2.74 seconds
Average time per image: 5.13 seconds
openaimath_cloneThe 500 problems from MATH that were used in "Let's verify step-by-step" and "s1: Simple test-time scaling".
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto},
year={2025},
eprint={2501.19393}… See the full description on the dataset page: https://huggingface.co/datasets/Afnanh10/openaimath_clone.originalall15_speaker_deduped_tts_train_clone_pairs_raw
All-15 Speaker-Deduped TTS Train Clone Pairs Raw
This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline.
Contents
data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows.
data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.llama-testvn_cn_datajx2_itemp_otherfu_infomessage-clone-datasetmessage-clone-dataset-1756464167message-clone-dataset-1756464267message-clone-dataset-1756469050us-army-fm-instruct-cloneConvert the dataset from here to the hugginge face conversational format.
The data remains the same, the keys just change
Usage
./download-convert-fm-instruct.sh
Description
(copied from original dataset)
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple, where a human user asks a question and the AI answers it.
Negative data is meant to help the AI be a bit more robust: the user asks a misinformed, flawed, or… See the full description on the dataset page: https://huggingface.co/datasets/Bobby060/us-army-fm-instruct-clone.
