datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coherence-decay-context-load
Coherence Decay Under Context Load
Dataset Summary
This dataset captures the degradation of internal coherence in large language models under increasing context length and conflicting identity conditions.
It is a controlled, synthetic experiment designed to measure how models behave when forced to maintain consistency across extended token sequences.
Two conditions are evaluated:
baseline: consistent identity prompt
aris_conflict: conflicting identity signals introduced… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/coherence-decay-context-load.TMDYFlix
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/saeesar/TMDYFlix.
