datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_cc_clone_detection_big_clone_bench
Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench"
Dataset Summary
CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.1-fold-clone-detection-600k-5foldCloneHeroDatasetCharts
Clone Hero Charts Dataset
Dataset Description
Tokenized Clone Hero charts with beat-level audio conditioning.
Each row is one instrument track (guitar / bass / drums) from one song.
Feature
Value
Total rows (train)
43,665
Parquet shards
1753
Audio: MERT embeddings
Yes [num_beats, 768]
Audio: log-mel frames
Yes [num_beats, 32, 128]
Dataset Structure
Data Fields
Column
Type
Description
song_id
string
MD5 hash of… See the full description on the dataset page: https://huggingface.co/datasets/thejorseman/CloneHeroDatasetCharts.4-fold-clone-detection-600k-5foldclone_han5i5j1986_eval_koch_lego_2024-10-18-01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 10,
"total_frames": 3438,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AdleBens/clone_han5i5j1986_eval_koch_lego_2024-10-18-01.clone-of-gretel-financial-risk-analysis-v1
⚠️🔴 IMPORTANT NOTICE 🔴⚠️
This dataset is directly cloned from gretelai/gretel-financial-risk-analysis-v1 on Hugging Face. No modifications have been made to the original dataset, it is only for archival.
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated using differential privacy guarantees, trained on 14,306 SEC (10-K, 10-Q, and 8-k) filings from 2023-2024. The dataset is designed for training models to extract… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/clone-of-gretel-financial-risk-analysis-v1.5-fold-clone-detection-600k-5fold3-fold-clone-detection-600k-5foldso101_pick_place_real_clone2-fold-clone-detection-600k-5folddataset2_clone_tam_thoiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ngocthuong2212/dataset2_clone_tam_thoi.
