datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Echo88-Instruct-173K
Echo88 Instruct 173K
A 173K-row retro instruction-tuning dataset for training Echo88-style small language models.
Echo88 Instruct 173K is an English supervised fine-tuning dataset created for training exnivo/Echo88-150M-Instruct, the instruction-following version of Echo88.
The dataset was built to teach a small language model how to answer questions, follow prompts, and behave like a helpful retro computer assistant whose knowledge is grounded in text from the 1950s… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/Echo88-Instruct-173K.EchoMist
Dataset Card for EchoMist
Introducing EchoMist, the first comprehensive benchmark to measure how LLMs may inadvertently Echo and amplify Misinformation hidden within seemingly innocuous user queries.
Dataset Description
Prior work has studied language models' capability to detect explicitly false statements. However, in real-world scenarios, circulating misinformation can often be referenced implicitly within user queries. When language models tacitly agree, they may… See the full description on the dataset page: https://huggingface.co/datasets/ruohao/EchoMist.echosim-synthetic-dialogues
EchoSim Synthetic Compatibility Dialogues (Sample)
A fully synthetic corpus of AI-simulated first-contact dialogues between two dating
personas. Each record pairs two personas (MBTI, attachment style, interests, age range)
with a short conversation. Current release: 200 sessions.
This is a public research / schema sample released by EchoSim.AI.
It is meant to illustrate the shape of the data used to study conversational compatibility —
not to expose any production system.
Links:… See the full description on the dataset page: https://huggingface.co/datasets/echo-sim/echosim-synthetic-dialogues.
