datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
platonic-embeddingsplatonic-embeddingsplatonic-all-experimentsEchoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.platonicnav-ablation-backup-20260519
PlatonicNav Ablation Records Backup 2026-05-20
This is a thin public backup of experiment records only.
It intentionally excludes bulky derived assets such as rendered frames/videos,
feature tensors, HDF5 caches, PTM pickles, graph caches, and failed tar shard
uploads.
Kept here:
ablation docs and cleanup notes
OVON short500 subset manifests and shard/batch manifests
final metrics, summaries, controller summaries, small logs
candidate-goal JSON/CSV records
historical run/result… See the full description on the dataset page: https://huggingface.co/datasets/dontKnow23456/platonicnav-ablation-backup-20260519.stanford_plato
Dataset Card for "stanford_plato"
Description
This is a collection of articles in the Stanford Encyclopedia of Philosophy (https://plato.stanford.edu/index.html).
This dataset includes 1776 articles, each explaining one philosophy term/people/topic. It has 8 features:
shorturl: The shorturl for the article. For example, the shorturl 'abduction' correspond to the page https://plato.stanford.edu/entries/abduction/
title: The title of the article.
pubinfo: The publication… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/stanford_plato.philosophy-plato-qa
Dataset Structure-- jsonl
{
"question": "QUESTION", // string
"context": "CONTEXT", // string
"target": "ANSWER" // string
}
Based on the Stanford Encyclopedia of Philosophy.
Datasets used:
SEP Articles: hugfaceguy0001/stanford_plato
Q&A Pairs: sayhan/strix-philosophy-qa
Data Collection and Processing
Compile data from both datasets -- python script
you can find a raw version for RAG/long context length models here
Dynamically find chunks of main_text… See the full description on the dataset page: https://huggingface.co/datasets/zayzay58/philosophy-plato-qa.platos_dialgouesplato33M tokens from 1.6K articles in the Stanford Encyclopedia of Philosophy, pulled in August 2025.
platonerf-real-dataplato-rings-6kplatos_socratesplatonic-embeddingsplato-chunks-500plato
Plato: philosophy essays from plato.stanford.edu
Plato is a corpus of 2.4k high quality philosophy essays from plato.stanford.edu.
philosophy-plato-qa-rawraw version of my philosophy-plato-qa
intended for use with RAG or long context lenght models
data format (.jsonl):
{
"label": "str",
"metadata": {
"pubinfo": "str",
"url": "https://plato.stanford.edu/entries/{label}/",
"related_entries": ["../{label}/", "../{label}/"]
},
"preamble": "str",
"main_text": "str",
"qa_pairs": [
{"question": "q1", "answer": "a1"},
{"question": "qN", "answer": "aN"}
]
}
plato-backview-test-h5platonov-style
Platonov Style Transfer Dataset
Описание
Датасет для задачи стилевого переноса: пары текстов Андрея Платонова (оригинал) и их переписки в формальном/нейтральном стиле. Предназначен для дообучения языковых моделей генерировать текст в стиле Платонова.
Назначение
Дообучение декодерных моделей (GPT-like) для генерации текста в стиле Платонова
Обучение классификаторов стиля (авторский vs формальный)
Исследование стилевого переноса в русском языке
Язык… See the full description on the dataset page: https://huggingface.co/datasets/IkSin/platonov-style.platonic-correct-experimentsPlatonplatos_socrates_no_contextplato-data-setrefanchor-platonic-hidden
Refanchor Platonic Hidden-state Backup
Private archive of long-running Qwen3 anchor-encoded features for Platonic-hypothesis analysis.
Layout
hidden3_235b/<model>.jsonl # Qwen3-235B-A22B encoded, 4096-d, 7 layers, top-32 logit-dist
hidden3_32b/<model>.jsonl # Qwen3-32B encoded, 5120-d, 7 layers, top-32 logit-dist
gen/<model>.jsonl # OPEN-panel model answers (OLLB models)
gen_closed/<model>.jsonl # CLOSED-panel model answers (AA… See the full description on the dataset page: https://huggingface.co/datasets/Chenhangcui/refanchor-platonic-hidden.plato-multiobjplato_testplato_flux_syn_data_testplato_flux_syn_data_test_trainhelios_platoplato_flux_syn_data_551_testplato_flux_syn_data_551_test_train
