datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
router-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.Mathlib-Normalized-Sexpr
Mathlib Normalized S-Expressions
Lean 4 proof states from Mathlib, paired with the tactic applied at each
step, in three representations extracted directly from the Lean kernel:
Source-faithful S-expressions of the goal and every hypothesis, as
Lean elaborated them.
Normalized S-expressions of the same state, with stable local-context
indices suitable for model input.
Annotated tactic syntax -- the original tactic's syntax tree with
identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.Amharic_news_Normalized
Dataset Name
Amharic news dataset
Dataset Details
It is a non-normalized version of news dataset crawled from Amharic news websites and from researchers provided in their works.
Dataset Description
The dataset is collected from different news websites and from different researchers crawled Amharic news dataset from different NLP downstream tasks.
News sites like FanaBC, EthiopianReporter, Zehabesha,Esat Amharic, BBC Amharic
are the sources for these news data.… See the full description on the dataset page: https://huggingface.co/datasets/userdavek/Amharic_news_Normalized.TinyPersianStories_normalizedtravelplanner-benchmark-normalized
TravelPlanner Benchmark (Normalized)
Normalized, typed, parquet-first packaging of the TravelPlanner benchmark for planning-centric agent evaluation.
Upstream dataset: osunlp/TravelPlanner
Upstream code: OSU-NLP-Group/TravelPlanner
Paper: TravelPlanner: A Benchmark for Real-World Planning with Language Agents
1) What is included
This dataset repo contains:
benchmark config (train/validation/test) in typed parquet.
reference_entries config: flattened reference-info… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/travelplanner-benchmark-normalized.dostoevsky_frontier_3k_normalized
dostoevsky_frontier_3k_normalized
Нормализованная версия 0x7o/dostoevsky_frontier_3k.
Нормализация
Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected
по артефактам форматирования вместо стиля.
Общие (chosen + rejected)
ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy)
\xa0 (неразрывный пробел) → обычный пробел
… (U+2026) → ... (три точки)
– (en dash) → — (em dash)
!.. → !..., ?.. → ?...
"..." → «...»… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_3k_normalized.dostoevsky_frontier_v2_normalized
dostoevsky_frontier_v2_normalized
Нормализованная версия 0x7o/dostoevsky_frontier_v2.
Нормализация
Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected
по артефактам форматирования вместо стиля.
Общие (chosen + rejected)
ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy)
\xa0 (неразрывный пробел) → обычный пробел
… (U+2026) → ... (три точки)
– (en dash) → — (em dash)
!.. → !..., ?.. → ?...
"..." → «...»… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_v2_normalized.
