datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.toricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.wikipedia-22-12-concat-split
Dataset Card for "wikipedia-22-12-concat-split"
More Information needed
split-avelina-python-edusplit-finemathhighlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptsfictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.Recap-DataComp-1B_split_3split_search_qa
preprocessed_SearchQA
The SearchQA question-answer pairs originate from J! Archive2, which comprehensively archives all question-answer pairs
from the renowned television show Jeopardy! The passages, sourced from Google search web page snippets.
We offer passage metadata, encompassing details like 'air_date,' 'category,' 'value,' 'round,' and 'show_number,'
enabling you to enhance retrieval performance at your discretion.
Should you require further details about SearchQA, please… See the full description on the dataset page: https://huggingface.co/datasets/NomaDamas/split_search_qa.Recap-DataComp-1B_split_4t2-ragbench-splitsRecap-DataComp-1B_split_5highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptstripadvisor-split-dataset
New Version Available
A newer version of this dataset with improved annotations and additional examples is available here.
Exercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoredRecap-DataComp-1B_split_7Recap-DataComp-1B_split_8Recap-DataComp-1B_split_2data_ablation_full59K-modernbert-split-kmeans-dim768-20250218craft-multiturn-actions-split-nothinkdist-defense-traces-taskname-split-augmented-plus-synth-v15
BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15
Task-name-disjoint train/test splits for dist-defense embedding training.
Contents
Splits: dist_train, dist_test
Built from: output/ctf_packaged_augmented_taskname_split_plus_synth_v15_trainonly
Split sizes: dist_train=132231, dist_test=234529
Split params: seed=42, train_ratio=0.9, benign_train_ratio=0.3
Synthetic merge: appended 35891 rows from… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15.Recap-DataComp-1B_split_1splack-splits-augment-10Recap-DataComp-1B_split_6telephooney-trainability-splits
🗂️ telephooney-trainability-splits
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Telephooney training split dataset.
تقسیمبندیهای آموزشپذیری گفتار تلفنی برای تحلیل کیفیت و تصمیمگیری دربارهٔ ورود نمونهها به آموزش.
🧩 Role
quality calibration and data-selection asset
مصنوع کالیبراسیون کیفیت و انتخاب داده
📦 Snapshot
10 files; approximately 299.28 MB
10 فایل؛… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telephooney-trainability-splits.shape-sorting-so101_split_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"Rotation",
"Pitch",
"Elbow",
"Wrist_Pitch",
"Wrist_Roll",
"Jaw"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101_split_test.pc_easy_nocp_split_200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 200,
"total_frames": 29170,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dureduck/pc_easy_nocp_split_200.split-avelina-python-edu-decontaminatedstack_cup_split_0402_train_v2_reindexpi-c0-c8-technique-dataset-clean_harmful_only_with_benign_tech2_self_reminder_v5_split70_30
