datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.cis5300-language-models
CIS 5300 Language Models Dataset
Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn.
Cities config
Country-of-origin classification over short city-name strings, drawn from
nine countries (Afghanistan, China, Germany, Finland, France, India, Iran,
Pakistan, South Africa).
from datasets import load_dataset
cities = load_dataset("CCB/cis5300-language-models", "cities")
Split
Rows
Has labels?
train
12,392
yes
validation
1,548
yes
test
1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.language_model_frnepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.ilm_deInstruction-Following-Evaluation-for-Large-Language-Models
Instruction-Following Evaluation Dataset
📜 Overview
This dataset, specifically designed for the evaluation of large language models in instruction-following tasks, is directly inspired by the methodologies and experiments described in the paper titled "Instruction-Following Evaluation for Large Language Models". The dataset's creation and availability on HuggingFace are aimed at enhancing research and application in the field of natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/Instruction-Following-Evaluation-for-Large-Language-Models.lener_br_finetuning_language_model
Dataset Card for "LeNER-Br language modeling"
Dataset Summary
The LeNER-Br language modeling dataset is a collection of legal texts in Portuguese from the LeNER-Br dataset (official site).
The legal texts were downloaded from this link (93.6MB) and processed to create a DatasetDict with train and validation dataset (20%).
The LeNER-Br language modeling dataset allows the finetuning of language models as BERTimbau base and large.
Language
Portuguese from… See the full description on the dataset page: https://huggingface.co/datasets/pierreguillou/lener_br_finetuning_language_model.ilm_esilm_itaBangla_Masked_Language_Model_dataset_preprocessedilm_polevals-for-every-language-modelsilm_euilm_slru_language_modeling_v4Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated
Dataset Card for Evaluation Awareness Benchmark
Dataset Summary
This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. The dataset contains 976 conversational transcripts with rich metadata, including:
True evaluation transcripts from prompt-injection tests, red-teaming tasks, and coding challenges
Organic/real transcripts from actual user queries, scraped chats, and… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated.ilm_nlc_corpus_br_finetuning_language_model_deberta
Dataset Card for "c_corpus_br_finetuning_language_model_deberta"
More Information needed
ilm_plsiqa_ca_old
Dataset Card: SIQA_CA (Pre-revision version)
Description
SIQA_CA (Pre-revision) is an earlier Catalan translation of the Social IQa (SIQA) dataset, a benchmark designed to evaluate commonsense reasoning about social interactions. This version consists of manually translated instances from the original English dataset into Catalan. It is used as a baseline for comparison against a revised and improved version of the dataset (SIQA_CA v2).
Motivation and Use Case… See the full description on the dataset page: https://huggingface.co/datasets/langtech-languagemodeling/siqa_ca_old.c_corpus_br_finetuning_language_model_bert
Dataset Card for "c_corpus_br_finetuning_language_model_bert"
More Information needed
hugging-face-language-models
Data from the configs of the 184 most popular language models on Hugging Face
aliaboost_when2call_esmobile-actions-language-modeling
Mobile Actions SFT Dataset
A converted version of the google/mobile-actions dataset for supervised fine-tuning (SFT) of Qwen models with tool calling capabilities.
Dataset Description
This dataset is derived from the google/mobile-actions dataset, which contains human-AI conversations about performing actions on mobile devices. The original dataset has been converted to the Qwen chat template format for efficient training of Qwen models.
Conversion Process
The… See the full description on the dataset page: https://huggingface.co/datasets/niwang66/mobile-actions-language-modeling.id_kenlm_language_modelSP_DOW_NASDAQ_stocks__News_Headlines_Language_ModellingALIABOOST-C2ru_language_modelinglener_br_finetuning_language_model
