CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ptb-text-only /ptb_text_onlyThis is the Penn Treebank Project: Release 2 CDROM, featuring a million words of 1989 Wall Street Journal material. This corpus has been annotated for part-of-speech (POS) information. In addition, over half of it has been annotated for skeletal syntactic structure.text-generation10K<n<100K20 likes27k downloads3y agoHugging Face02macabdul9 /hle_text_only Humanity's Last Exam - (Text only) 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 3,000 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and… See the full description on the dataset page: https://huggingface.co/datasets/macabdul9/hle_text_only.image1K<n<10K4 likes23k downloads2y agoHugging Face03tasksource /ScienceQA_text_only Dataset Card for "scienceQA_text_only" ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.) @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K32 likes3.6k downloads3y agoHugging Face04sayan1101 /gaia_filtered_text_onlytextn<1K0 likes3.1k downloads3y agoHugging Face05DongfuJiang /hle_text_onlyimage1K<n<10K0 likes2.8k downloads1y agoHugging Face06XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes2.4k downloads16d agoHugging Face07DorayakiLin /TextOnly_FromRLBench_CloseBox_24K_unfixedtabular1M<n<10M0 likes1.2k downloads10mo agoHugging Face08rl-rag /hle_text_onlyimage1K<n<10K0 likes1k downloads7mo agoHugging Face09alielfilali01 /fineweb-2-arb_Arab-text-onlytext10M<n<100M2 likes525 downloads2y agoHugging Face10alielfilali01 /fineweb-2-arb_Arab-text-only-2text10M<n<100M0 likes503 downloads2y agoHugging Face11agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes395 downloads11mo agoHugging Face12BrunoHays /ESLO_text_onlyESLO dataset, each utterance are taken out individually0 likes370 downloads2y agoHugging Face13DorayakiLin /TextOnly_Fixed_1000100K<n<1M0 likes305 downloads1y agoHugging Face14DorayakiLin /TextOnly_15_150100K<n<1M0 likes305 downloads1y agoHugging Face15luoojason /mmlongbench-text-only MUDDLE - source questions and hard-negative metadata The 270 human-annotated questions behind MUDDLE, derived from MMLongBench-Doc, plus the curated hard-negative pool (hard_negatives.json) and the source / hard-negative PDFs. Each question is tied to a single source document and is instantiated in five context conditions in the eval bundles: source alone, source + 2 or 4 topically similar hard negatives, and source + 2 or 4 length-matched random distractors. Related… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mmlongbench-text-only.documentn<1K0 likes293 downloads2mo agoHugging Face16ejhwang /VideoMMMU-Res-Text-Onlyimage10K<n<100K0 likes273 downloads1y agoHugging Face17Yusiko /hle-text-only-cleanedtext1K<n<10K0 likes269 downloads8mo agoHugging Face18DorayakiLin /Textonly_Test10K<n<100K0 likes267 downloads1y agoHugging Face19OccasionallyNLP /hle_text_onlytext1K<n<10K0 likes262 downloads11mo agoHugging Face20Geonmo /coyo700m-text-onlytext100M<n<1B1 likes259 downloads3y agoHugging Face21DorayakiLin /Textonly_Franka1M<n<10M0 likes197 downloads1y agoHugging Face22theojiang /image-text-dataset-subset-300k-captions_onlyimage100K<n<1M1 likes181 downloads3y agoHugging Face23DorayakiLin /Franka_Textonly_Unfixed_parquet1M<n<10M0 likes176 downloads1y agoHugging Face24kastan /stormfront-full-textonly Dataset Card for "stormfront-full-textonly" More Information needed text10M<n<100M0 likes162 downloads4y agoHugging Face25jan-hq /instruction-data-text-onlytabular1M<n<10M0 likes161 downloads2y agoHugging Face26DorayakiLin /franka_textonly_lerobotv16tabular10K<n<100K0 likes158 downloads1y agoHugging Face27DorayakiLin /Textonly_franka_parquet10K<n<100K0 likes152 downloads1y agoHugging Face28OccasionallyNLP /hle-text-onlytext1K<n<10K0 likes151 downloads11mo agoHugging Face29WPRM /preference_data_llama_factory_corrected_format_text_onlyimage10K<n<100K0 likes141 downloads1y agoHugging Face30Yah216 /Poem_APCD_text_onlyWe used the APCD dataset cited hereafter for pretraining the model. The dataset has been cleaned and only the main text column was kept: @Article{Yousef2019LearningMetersArabicEnglish-arxiv, author = {Yousef, Waleed A. and Ibrahime, Omar M. and Madbouly, Taha M. and Mahmoud, Moustafa A.}, title = {Learning Meters of Arabic and English Poems With Recurrent Neural Networks: a Step Forward for Language Understanding and Synthesis}, journal =… See the full description on the dataset page: https://huggingface.co/datasets/Yah216/Poem_APCD_text_only.text1M<n<10M0 likes132 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.