datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle_text_only
Humanity's Last Exam - (Text only)
🌐 Website | 📄 Paper | GitHub
Center for AI Safety & Scale AI
Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 3,000 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and… See the full description on the dataset page: https://huggingface.co/datasets/macabdul9/hle_text_only.ScienceQA_text_only
Dataset Card for "scienceQA_text_only"
ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.)
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
gaia_filtered_text_onlyhle_text_onlyJava-Code-Large-text-onlyJava-Code-Large
Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.
By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.hle_text_onlyfineweb-2-arb_Arab-text-onlyfineweb-2-arb_Arab-text-only-2text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.mmlongbench-text-only
MUDDLE - source questions and hard-negative metadata
The 270 human-annotated questions behind MUDDLE, derived from MMLongBench-Doc, plus the
curated hard-negative pool (hard_negatives.json) and the source / hard-negative PDFs.
Each question is tied to a single source document and is instantiated in five context
conditions in the eval bundles: source alone, source + 2 or 4 topically similar hard
negatives, and source + 2 or 4 length-matched random distractors.
Related… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mmlongbench-text-only.VideoMMMU-Res-Text-Onlyhle-text-only-cleanedhle_text_onlycoyo700m-text-onlyimage-text-dataset-subset-300k-captions_onlystormfront-full-textonly
Dataset Card for "stormfront-full-textonly"
More Information needed
instruction-data-text-onlyhle-text-onlypreference_data_llama_factory_corrected_format_text_onlyPoem_APCD_text_onlyWe used the APCD dataset cited hereafter for pretraining the model. The dataset has been cleaned and only the main text column was kept:
@Article{Yousef2019LearningMetersArabicEnglish-arxiv,
author = {Yousef, Waleed A. and Ibrahime, Omar M. and Madbouly, Taha M. and Mahmoud,
Moustafa A.},
title = {Learning Meters of Arabic and English Poems With Recurrent Neural Networks: a Step
Forward for Language Understanding and Synthesis},
journal =… See the full description on the dataset page: https://huggingface.co/datasets/Yah216/Poem_APCD_text_only.VoxBox-Speech-TextOnlydclm_1M_text_onlyMagpie-Speech-TextOnlymustc-en-es-text-only
Dataset Card for kudo-research/mustc-en-es-text-only
Dataset Summary
This dataset is a selection of text only (English-Spanish) from the MuST-C corpus.
MuST-C is a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 14 languages (Arabic, Chinese, Czech, Dutch, French, German, Italian, Persian, Portuguese, Romanian, Russian, Spanish, Turkish and Vietnamese).
For each target… See the full description on the dataset page: https://huggingface.co/datasets/kudo-research/mustc-en-es-text-only.Video-MME-Res-Text-Onlyimage-text-dataset-subset-300k-captions_only_with_latentsinstruction-data-text-only-multiturnInstruction-text-only-fullMixSub-LLaMA-3.2-Text-Only-Overlap-CPU-Scoreinstruction-data-text-only-multiturn-filtered-for-tokenize-clean
