datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SP_Referential_DisambiguationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 51,
"total_frames": 35895,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Referential_Disambiguation.task249_enhanced_wsc_pronoun_disambiguation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task249_enhanced_wsc_pronoun_disambiguation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task249_enhanced_wsc_pronoun_disambiguation.rollout_molmoact2_Reasoning_Step_076596_Referential_DisambiguationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_molmoact2_Reasoning_Step_076596_Referential_Disambiguation.SP_Referential_Disambiguation_200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Referential_Disambiguation_200.rollout_groot_vision_only_Referential_DisambiguationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Referential_Disambiguation.rollout_groot_Referential_Disambiguation_20260827_150732This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_Referential_Disambiguation_20260827_150732.poleval-abbreviation-disambiguation-wiki
PolEval 2022 Task 2 Pretraining Dataset
Dataset Summary
Abbreviation disambiguation is the process of expanding abbreviations, e.g. "eng." to the full form "engineer".
In the Polish language the task is further complicated because of the many ways to create abbreviations and additional inflected forms.
Abbreviation disambiguation was the topic of the 2022 PolEval Competition Task 2.
This is the dataset used for pretraining in Jakub Karbowski's competition submission.… See the full description on the dataset page: https://huggingface.co/datasets/carbon225/poleval-abbreviation-disambiguation-wiki.rollout_groot_vision_only_Referential_Disambiguation_20260829_160104This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_groot_vision_only_Referential_Disambiguation_20260829_160104.rollout_vla0_Referential_Disambiguation_90s_20260828_154019This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_vla0_Referential_Disambiguation_90s_20260828_154019.phrase_sense_disambiguationPhrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD). The datasets are annotated by 13 linguistic experts on Upwork and verified by two groups: ~1000 AMT crowdworkers and another set of 5 linguistic experts. PiC benchmark is distributed under CC-BY-NC 4.0.rollout_pi05_Reasoning_Step_076596_Referential_DisambiguationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_pi05_Reasoning_Step_076596_Referential_Disambiguation.rollout_molmoact2_Reasoning_Step_076596_Referential_Disambiguation_20260811_185132This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_molmoact2_Reasoning_Step_076596_Referential_Disambiguation_20260811_185132.ka_homonym_disambiguation
Georgian-Homonym-Disambiguation
This repository contains all the datasets for the Georgian homonym disambiguation task.
For more specific details you can read my article
Dataset
At this point I've considered only the homonym: "ბარი" and it's different grammatical forms obtaining 7522 sentences.
The "dataset.parquet" includes:
763 sentences using "ბარი" as a "shovel" labaled with 0
1846 sentences using "ბარი" as a "lowland" labeld with 1
3320 sentences using "ბარი" as a… See the full description on the dataset page: https://huggingface.co/datasets/davmel/ka_homonym_disambiguation.entity_disambiguationEntity Disambiguation datasets as provided in the GENRE repo. The dataset can be used to train and evaluate entity disambiguators.
The datasets can be imported easily as follows:
from datasets import load_dataset
ds = load_dataset("boragokbakan/entity_disambiguation", "aida")
Available dataset names are:
blink
ace2004
aida
aquaint
blink
clueweb
msnbc
wiki
Note: As the BLINK training set is very large in size (~10GB), it is advised to set streaming=True when calling load_dataset.
bbeh-disambiguation-qa
Reference
@article{kazemi2025big,
title={Big-bench extra hard},
author={Kazemi, Mehran and Fatemi, Bahare and Bansal, Hritik and Palowitch, John and Anastasiou, Chrysovalantis and Mehta, Sanket Vaibhav and Jain, Lalit K and Aglietti, Virginia and Jindal, Disha and Chen, Peter and others},
journal={arXiv preprint arXiv:2502.19187},
year={2025}
}
flan_combined_task249_enhanced_wsc_pronoun_disambiguationclinical-healing-trajectory-plateau-meaning-disambiguation-v0.1What this dataset tests
Whether a model can interpret recovery plateausand choose the correct clinical response.
Required outputs
plateau_type
next_expected_transition
risk_guardrail
Typical failures
treating all plateaus as stalled
ignoring activity load context
recommending intervention without evidence
Suggested prompt wrapper
System
You interpret recovery plateaus.
User
Recent phases{recent_phase_sequence}
Vitals{vitals_trend}
Labs{labs_trend}… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-plateau-meaning-disambiguation-v0.1.SP_Referential_Disambiguation_200_20260722_165411This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Referential_Disambiguation_200_20260722_165411.bntng-disambiguation-v1
Dataset Card for bntng-disambiguation-v1
Dataset Summary
This dataset focuses on word sense disambiguation in Jawi script, specifically addressing cases where the same written form can correspond to different words in Malay. The dataset centers around the ambiguous Jawi writing بنتڠ (bntng) and its various interpretations, with 500 contextual examples for each possible reading. This dataset is designed to test Large Language Models' ability to predict the correct… See the full description on the dataset page: https://huggingface.co/datasets/mevsg/bntng-disambiguation-v1.Large-Scale_Multilingual_Disambiguation_Glosses
[!NOTE]
Dataset origin: http://lrec2016.lrec-conf.org/en/shared-lrs/
Description
A multilingual large-scale corpus of automatically disambiguated glosses drawn from different resources integrated in BabelNet (such as Wikipedia, Wiktionary, WordNet, OmegaWiki and Wikidata). Sense annotations for both concepts and named entities are provided. In total, over 40 millions definitions have been disambiguated for 264 languages.
Citation
@InProceedings{CAMACHOCOLLADOS16.629… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Large-Scale_Multilingual_Disambiguation_Glosses.affiliation-disambiguation-triplets
Affiliation Triplets for Contrastive Learning
This dataset contains 1,083,631 triplets (anchor, positive, negative) for training affiliation embedding and reranking models using triplet loss or contrastive learning.
Dataset Description
Each triplet consists of:
Anchor: An affiliation string from OpenAlex
Positive: A different affiliation string for the same organization (same ROR ID)
Negative: An affiliation string for a different organization
The dataset is sorted by… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/affiliation-disambiguation-triplets.rollout_pi05_Reasoning_Step_076596_Referential_Disambiguation_20260819_105915This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/rollout_pi05_Reasoning_Step_076596_Referential_Disambiguation_20260819_105915.silent_signals_disambiguation
Silent Signals | Human-Annotated Evaluation set for Dogwhistle Disambiguation (Synthetic Disambiguation Dataset)
A dataset of human-annotated dogwhistle use cases to evaluate models on their ability to disambiguate dogwhitles from standard vernacular. A dogwhistle is a form of coded communication that carries a secondary meaning to specific audiences and is often weaponized for racial and socioeconomic discrimination. Dogwhistling historically originated from United States politics… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/silent_signals_disambiguation.phrase_sense_disambiguationfaithfulness-disambiguation_qafaithfulness-disambiguation_qa-arbDisambiguation-Aware_CoT_ConstructionTopic-specific-disambiguation_evaluation-dataset_German_historical-newspapers
Dataset Card for German Topic-specific Disambiguation Dataset Historical Newspapers
This dataset contains German newspaper articles (1850-1950) labeled as either about relevant for "return migration" or not. It helps researchers develop word sense disambiguation methods - technology that can tell when the German word "Heimkehr" or "Rückkehr" (return/homecoming) is being used to discuss people returning to their home countries versus when it's used in different contexts. This matters… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-disambiguation_evaluation-dataset_German_historical-newspapers.qwen3_0.6b-rlvr_task249_enhanced_wsc_pronoun_disambiguationdisambiguation-pretraining-v2
Disambiguation Pretraining v2
Disambiguation Pretraining v2 is an English-Italian corpus for pretraining language models on lexical semantics and word-sense disambiguation. It connects lemmas, definitions, attested contexts, contrasting senses, parts of speech, and subject domains through varied natural-language formulations.
Version 2 contains 3,693,921 rows, compared with 900,009 in the first release. The increase comes from new semantic tasks and controlled reuse of source… See the full description on the dataset page: https://huggingface.co/datasets/MikCil/disambiguation-pretraining-v2.
