datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.toricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.wikipedia-22-12-concat-split
Dataset Card for "wikipedia-22-12-concat-split"
More Information needed
ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.split-avelina-python-edusplit-finemathprocessed_test_splitshighlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptsOpenVid-60k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 60k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-60k-split.fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.Recap-DataComp-1B_split_3split_search_qa
preprocessed_SearchQA
The SearchQA question-answer pairs originate from J! Archive2, which comprehensively archives all question-answer pairs
from the renowned television show Jeopardy! The passages, sourced from Google search web page snippets.
We offer passage metadata, encompassing details like 'air_date,' 'category,' 'value,' 'round,' and 'show_number,'
enabling you to enhance retrieval performance at your discretion.
Should you require further details about SearchQA, please… See the full description on the dataset page: https://huggingface.co/datasets/NomaDamas/split_search_qa.Recap-DataComp-1B_split_4t2-ragbench-splitsRecap-DataComp-1B_split_5highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptstripadvisor-split-dataset
New Version Available
A newer version of this dataset with improved annotations and additional examples is available here.
Exercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoredRecap-DataComp-1B_split_7Recap-DataComp-1B_split_8Recap-DataComp-1B_split_2OpenVid-10k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.data_ablation_full59K-modernbert-split-kmeans-dim768-20250218craft-multiturn-actions-split-nothinkdist-defense-traces-taskname-split-augmented-plus-synth-v15
BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15
Task-name-disjoint train/test splits for dist-defense embedding training.
Contents
Splits: dist_train, dist_test
Built from: output/ctf_packaged_augmented_taskname_split_plus_synth_v15_trainonly
Split sizes: dist_train=132231, dist_test=234529
Split params: seed=42, train_ratio=0.9, benign_train_ratio=0.3
Synthetic merge: appended 35891 rows from… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15.Recap-DataComp-1B_split_1splack-splits-augment-10Recap-DataComp-1B_split_6telephooney-trainability-splits
🗂️ telephooney-trainability-splits
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Telephooney training split dataset.
تقسیمبندیهای آموزشپذیری گفتار تلفنی برای تحلیل کیفیت و تصمیمگیری دربارهٔ ورود نمونهها به آموزش.
🧩 Role
quality calibration and data-selection asset
مصنوع کالیبراسیون کیفیت و انتخاب داده
📦 Snapshot
10 files; approximately 299.28 MB
10 فایل؛… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telephooney-trainability-splits.shape-sorting-so101_split_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"Rotation",
"Pitch",
"Elbow",
"Wrist_Pitch",
"Wrist_Roll",
"Jaw"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101_split_test.
