CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes479 downloads1y agoHugging Face02CCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes472 downloads5mo agoHugging Face03Plim /language_model_frtext1M<n<10M0 likes182 downloads4y agoHugging Face04Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes113 downloads6mo agoHugging Face05beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes107 downloads1mo agoHugging Face06interlinguistic-language-modeling /ilm_detext1M<n<10M0 likes54 downloads6mo agoHugging Face07harpreetsahota /Instruction-Following-Evaluation-for-Large-Language-Models Instruction-Following Evaluation Dataset 📜 Overview This dataset, specifically designed for the evaluation of large language models in instruction-following tasks, is directly inspired by the methodologies and experiments described in the paper titled "Instruction-Following Evaluation for Large Language Models". The dataset's creation and availability on HuggingFace are aimed at enhancing research and application in the field of natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/Instruction-Following-Evaluation-for-Large-Language-Models.textn<1K7 likes49 downloads3y agoHugging Face08pierreguillou /lener_br_finetuning_language_model Dataset Card for "LeNER-Br language modeling" Dataset Summary The LeNER-Br language modeling dataset is a collection of legal texts in Portuguese from the LeNER-Br dataset (official site). The legal texts were downloaded from this link (93.6MB) and processed to create a DatasetDict with train and validation dataset (20%). The LeNER-Br language modeling dataset allows the finetuning of language models as BERTimbau base and large. Language Portuguese from… See the full description on the dataset page: https://huggingface.co/datasets/pierreguillou/lener_br_finetuning_language_model.text10K<n<100K6 likes46 downloads4y agoHugging Face09interlinguistic-language-modeling /ilm_estext1M<n<10M0 likes41 downloads6mo agoHugging Face10interlinguistic-language-modeling /ilm_itatext1M<n<10M0 likes39 downloads4mo agoHugging Face11DhimanBose /Bangla_Masked_Language_Model_dataset_preprocessedtext-generation1M<n<10M0 likes31 downloads3y agoHugging Face12interlinguistic-language-modeling /ilm_poltext1M<n<10M0 likes31 downloads4mo agoHugging Face13fair-forward /evals-for-every-language-modelstabularn<1K0 likes27 downloads3mo agoHugging Face14interlinguistic-language-modeling /ilm_eutext1M<n<10M0 likes22 downloads6mo agoHugging Face15interlinguistic-language-modeling /ilm_sltext100K<n<1M0 likes19 downloads6mo agoHugging Face16patrikgerard /ru_language_modeling_v4text1K<n<10K0 likes18 downloads2y agoHugging Face17eval-aware /Large-Language-Models-Often-Know-When-They-Are-Being-Evaluatedgated Dataset Card for Evaluation Awareness Benchmark Dataset Summary This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. The dataset contains 976 conversational transcripts with rich metadata, including: True evaluation transcripts from prompt-injection tests, red-teaming tasks, and coding challenges Organic/real transcripts from actual user queries, scraped chats, and… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated.tabulartext-classificationn<1K0 likes17 downloads6mo agoHugging Face18interlinguistic-language-modeling /ilm_nltext100K<n<1M0 likes17 downloads6mo agoHugging Face19rosimeirecosta /c_corpus_br_finetuning_language_model_deberta Dataset Card for "c_corpus_br_finetuning_language_model_deberta" More Information needed text100K<n<1M2 likes15 downloads4y agoHugging Face20interlinguistic-language-modeling /ilm_pltext1M<n<10M0 likes15 downloads6mo agoHugging Face21langtech-languagemodeling /siqa_ca_old Dataset Card: SIQA_CA (Pre-revision version) Description SIQA_CA (Pre-revision) is an earlier Catalan translation of the Social IQa (SIQA) dataset, a benchmark designed to evaluate commonsense reasoning about social interactions. This version consists of manually translated instances from the original English dataset into Catalan. It is used as a baseline for comparison against a revised and improved version of the dataset (SIQA_CA v2). Motivation and Use Case… See the full description on the dataset page: https://huggingface.co/datasets/langtech-languagemodeling/siqa_ca_old.textquestion-answering1K<n<10K0 likes15 downloads5mo agoHugging Face22rosimeirecosta /c_corpus_br_finetuning_language_model_bert Dataset Card for "c_corpus_br_finetuning_language_model_bert" More Information needed text100K<n<1M0 likes14 downloads4y agoHugging Face23kgourgou /hugging-face-language-models Data from the configs of the 184 most popular language models on Hugging Face tabularn<1K2 likes14 downloads2y agoHugging Face24langtech-languagemodeling /aliaboost_when2call_estextn<1K0 likes14 downloads6mo agoHugging Face25niwang66 /mobile-actions-language-modeling Mobile Actions SFT Dataset A converted version of the google/mobile-actions dataset for supervised fine-tuning (SFT) of Qwen models with tool calling capabilities. Dataset Description This dataset is derived from the google/mobile-actions dataset, which contains human-AI conversations about performing actions on mobile devices. The original dataset has been converted to the Qwen chat template format for efficient training of Qwen models. Conversion Process The… See the full description on the dataset page: https://huggingface.co/datasets/niwang66/mobile-actions-language-modeling.text1K<n<10K0 likes13 downloads6mo agoHugging Face26chrisvinsen /id_kenlm_language_modeltext1K<n<10K0 likes10 downloads4y agoHugging Face27firobeid /SP_DOW_NASDAQ_stocks__News_Headlines_Language_Modellingtabular100K<n<1M0 likes10 downloads1y agoHugging Face28langtech-languagemodeling /ALIABOOST-C2textn<1K0 likes9 downloads7mo agoHugging Face29patrikgerard /ru_language_modelingtext100K<n<1M0 likes8 downloads2y agoHugging Face30aspirina765 /lener_br_finetuning_language_modeltext1K<n<10K1 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.