datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oasst1_ca
Dataset Card for oasst1_ca
oasst1_ca is a conversational dataset in Catalan that has been professionally translated from the OASST1 dataset.
Dataset Details
Dataset Description
oasst1_ca (OpenAssistant Conversations Release 1 - Catalan) consists of human-generated, human-annotated assistant-style conversation corpus. It includes 5213 messages in the train split and 273 messages in the validation split. To arrive to this number, we filter the dataset (See… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/oasst1_ca.oasst1-decontaminated
Decontaminated — OpenAssistant/oasst1
What this is
A filtered version of OpenAssistant/oasst1 (revision
fdf72ae0827c1cda404aff25b6603abec9e3399b) with exact-duplicate rows and rows overlapping standard benchmark test sets
removed. This is a different artifact from the companion contamination report — that one is an
audit of what's wrong; this one is the corpus with those rows actually taken out, ready to train on.
Processing
Deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/oasst1-decontaminated.oasst1_idThis is Indonesian version of OASST1 dataset, translated entirely using HelsinkiNLP OPUS models and llama2lang library.
Feel free to request another dataset translation into Bahasa Indonesia, i'll try to help.
Fellow Indonesians, we shall not be left behind in the age of AI.
recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.oasst1_shan_translationThis datasets is a translation version of OpenAssistant/oasst1 to Shan language, translated by facebook/nllb-200-3.3B.
The data quality has not been checked by a human yet, so that it might be of low quality.
oasst1-zh-pilot
oasst1 Chinese Translation Pilot (10 samples)
This is a pilot release of 10 parallel English→Chinese samples translated from
OpenAssistant/oasst1. It is
intended as a methodology demonstration and quality evaluation artifact, not as
a training-ready dataset.
Why this exists
We are evaluating whether LLM-assisted translation of open instruction-tuning datasets
into low-resource languages can be done at a quality bar that the ML community will
accept. Chinese is our first… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/oasst1-zh-pilot.OpenAssistant-oasst1-fa
Dataset Card for "OpenAssistant-oasst1-fa"
More Information needed
