datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
modality-ood
Modality OOD Dataset
Dataset Description
The Modality OOD dataset tests model generalization across different data modalities in peptide-MHC (pMHC) binding prediction. It contains two complementary datasets representing distinct experimental measurement types:
BA (Binding Affinity): In vitro binding affinity measurements with continuous values
EL (Eluted Ligand): Mass spectrometry-based eluted ligand data with binary labels
Key Features
Modality Shift… See the full description on the dataset page: https://huggingface.co/datasets/YYJMAY/modality-ood.clinical-cross-modal-memory-fidelity-v0.1Clinical Cross-Modal Memory Fidelity v0.1
Goal
Test whether prior image evidence is recalled accurately over time
Detect retroactive distortion driven by later narrative
Detect fabrication used to patch memory gaps
What it measures
memory_driftEarlier image facts are altered or inverted
fabricationNew findings are invented at recall
cross_modal_consistencyRecalled description matches original image evidence
How it works
Initial image facts are fixed and explicit
Intervening tasks… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-cross-modal-memory-fidelity-v0.1.modality-state-consistency-v0.1
What this dataset tests
Inputs arrive in many forms.
State must stay coherent.
Why it exists
Models drift when switching modality.
Facts change.
Promises vanish.
This set checks whether state stays consistent.
Data format
Each row contains
modality_context
user_message
modality_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
modality_context
user_message
constraints
Score for
cross-modal… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/modality-state-consistency-v0.1.ModalUsage100K
ModalUsage100K
tags: part-of-speech tagging, language modeling, dataset size
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'ModalUsage100K' dataset comprises of sentences extracted from diverse texts, each labeled with a specific modal verb that it contains. This dataset is designed for natural language processing tasks focusing on the use and context of modal verbs in English sentences. Each sentence has been manually… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ModalUsage100K.school-learning-modalities-2021-2022
School Learning Modalities, 2021-2022
Description
The 2021-2022 School Learning Modalities dataset provides weekly estimates of school learning modality (including in-person, remote, or hybrid learning) for U.S. K-12 public and independent charter school districts for the 2021-2022 school year and the Fall 2022 semester, from August 2021 – December 2022.
These data were modeled using multiple sources of input data (see below) to infer the most likely learning modality of… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/school-learning-modalities-2021-2022.clinical-cross-modal-memory-fidelity-v0.2
Clinical Cross Modal Memory Fidelity v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical system is moving toward cross-modal memory fidelity failure, not just carrying ambiguity?
This repo focuses on the integrity between cross-modal evidence and memory fidelity.
It models a system where:
cross-modal alignment may weaken
memory trace fidelity may drift
recall distortion may rise
representational stability may erode before overt… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-cross-modal-memory-fidelity-v0.2.modal_verb_dataschool-learning-modalities-2020-2021
School Learning Modalities, 2020-2021
Description
The 2020-2021 School Learning Modalities dataset provides weekly estimates of school learning modality (including in-person, remote, or hybrid learning) for U.S. K-12 public and independent charter school districts for the 2020-2021 school year, from August 2020 – June 2021.
These data were modeled using multiple sources of input data (see below) to infer the most likely learning modality of a school district for a given… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/school-learning-modalities-2020-2021.tulu_3_model_modality_mismatch_errorsyoutube_marketingMulti-Modal_Sentiment_Analysis_in_E-commerce
