CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01macabdul9 /hle_text_only Humanity's Last Exam - (Text only) 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 3,000 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and… See the full description on the dataset page: https://huggingface.co/datasets/macabdul9/hle_text_only.image1K<n<10K4 likes23k downloads2y agoHugging Face02tasksource /ScienceQA_text_only Dataset Card for "scienceQA_text_only" ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.) @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K32 likes3.6k downloads3y agoHugging Face03sayan1101 /gaia_filtered_text_onlytextn<1K0 likes3.1k downloads3y agoHugging Face04DongfuJiang /hle_text_onlyimage1K<n<10K0 likes2.8k downloads1y agoHugging Face05XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes2.4k downloads16d agoHugging Face06rl-rag /hle_text_onlyimage1K<n<10K0 likes1k downloads7mo agoHugging Face07alielfilali01 /fineweb-2-arb_Arab-text-onlytext10M<n<100M2 likes525 downloads2y agoHugging Face08alielfilali01 /fineweb-2-arb_Arab-text-only-2text10M<n<100M0 likes503 downloads2y agoHugging Face09agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes395 downloads11mo agoHugging Face10luoojason /mmlongbench-text-only MUDDLE - source questions and hard-negative metadata The 270 human-annotated questions behind MUDDLE, derived from MMLongBench-Doc, plus the curated hard-negative pool (hard_negatives.json) and the source / hard-negative PDFs. Each question is tied to a single source document and is instantiated in five context conditions in the eval bundles: source alone, source + 2 or 4 topically similar hard negatives, and source + 2 or 4 length-matched random distractors. Related… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mmlongbench-text-only.documentn<1K0 likes293 downloads2mo agoHugging Face11ejhwang /VideoMMMU-Res-Text-Onlyimage10K<n<100K0 likes273 downloads1y agoHugging Face12Yusiko /hle-text-only-cleanedtext1K<n<10K0 likes269 downloads8mo agoHugging Face13OccasionallyNLP /hle_text_onlytext1K<n<10K0 likes262 downloads11mo agoHugging Face14Geonmo /coyo700m-text-onlytext100M<n<1B1 likes259 downloads3y agoHugging Face15theojiang /image-text-dataset-subset-300k-captions_onlyimage100K<n<1M1 likes181 downloads3y agoHugging Face16kastan /stormfront-full-textonly Dataset Card for "stormfront-full-textonly" More Information needed text10M<n<100M0 likes162 downloads4y agoHugging Face17jan-hq /instruction-data-text-onlytabular1M<n<10M0 likes161 downloads2y agoHugging Face18OccasionallyNLP /hle-text-onlytext1K<n<10K0 likes151 downloads11mo agoHugging Face19WPRM /preference_data_llama_factory_corrected_format_text_onlyimage10K<n<100K0 likes141 downloads1y agoHugging Face20Yah216 /Poem_APCD_text_onlyWe used the APCD dataset cited hereafter for pretraining the model. The dataset has been cleaned and only the main text column was kept: @Article{Yousef2019LearningMetersArabicEnglish-arxiv, author = {Yousef, Waleed A. and Ibrahime, Omar M. and Madbouly, Taha M. and Mahmoud, Moustafa A.}, title = {Learning Meters of Arabic and English Poems With Recurrent Neural Networks: a Step Forward for Language Understanding and Synthesis}, journal =… See the full description on the dataset page: https://huggingface.co/datasets/Yah216/Poem_APCD_text_only.text1M<n<10M0 likes132 downloads4y agoHugging Face21humair025 /VoxBox-Speech-TextOnlytext100K<n<1M0 likes132 downloads5mo agoHugging Face22felixZzz /dclm_1M_text_onlytext1M<n<10M0 likes93 downloads11mo agoHugging Face23humair025 /Magpie-Speech-TextOnlytext100K<n<1M0 likes71 downloads5mo agoHugging Face24kudo-research /mustc-en-es-text-only Dataset Card for kudo-research/mustc-en-es-text-only Dataset Summary This dataset is a selection of text only (English-Spanish) from the MuST-C corpus. MuST-C is a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 14 languages (Arabic, Chinese, Czech, Dutch, French, German, Italian, Persian, Portuguese, Romanian, Russian, Spanish, Turkish and Vietnamese). For each target… See the full description on the dataset page: https://huggingface.co/datasets/kudo-research/mustc-en-es-text-only.text100K<n<1M0 likes65 downloads4y agoHugging Face25ejhwang /Video-MME-Res-Text-Onlytext10K<n<100K1 likes56 downloads1y agoHugging Face26theojiang /image-text-dataset-subset-300k-captions_only_with_latentsimage100K<n<1M2 likes55 downloads3y agoHugging Face27jan-hq /instruction-data-text-only-multiturntext100K<n<1M0 likes55 downloads2y agoHugging Face28Menlo /Instruction-text-only-fulltext10K<n<100K3 likes48 downloads1y agoHugging Face29AdityaMayukhSom /MixSub-LLaMA-3.2-Text-Only-Overlap-CPU-Scoretabular1K<n<10K0 likes48 downloads1y agoHugging Face30jan-hq /instruction-data-text-only-multiturn-filtered-for-tokenize-cleantext100K<n<1M0 likes47 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.