datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.dialectical-reasoningA specialised dialectical reasoning dataset.
contain { Thesis:, Antithesis:, Synthesis: }.
Domain are math, science, creative writing
qarachay-malqar_russian_parallel_corpora_dialectic-free288532 parallel sentences between russian and Qarachay-Malqar languages. Taken from: A corpuse of Qarachay-Malqar folklore (tales, epics and etc.), Poems of Kaisyn Kuliev, Uzden Codex, Religious literature, Artistic literature, Movies, Cartoons, Soviet reports, Qarachay-Malqar phrasebook, Dictionary.All dataset is devided into sereral types: one_sentence (One sentence), one_word (one word or short phrase), several_sentences (several sentences about 5 or Paragraph).
Because of dialects… See the full description on the dataset page: https://huggingface.co/datasets/TSjB/qarachay-malqar_russian_parallel_corpora_dialectic-free.t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-0b792270
t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-0b792270
Dialectical reasoning evaluation for MuSR murder mysteries. The model acts as both
PROSECUTION and DEFENSE for each suspect, constructing the strongest possible case
for guilt and then the strongest possible case for innocence. The suspect with the
weaker defense is identified as the murderer.
Stylistically, the model uses formal legal argumentation: "the prosecution submits",
"the defense counters"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-thin-0b792270.dialectic-preferences-bias-aae-sae-parallel
Dialectic Preferences Bias Dataset
Dataset Description
Overview
This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects.
The dataset contains two columns:
african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.gsm8k-dialecticdialectic-reasoning-traces
Dialectic Reasoning Traces
255 scored dialectic reasoning traces for training models on integrative resolution under conflicting frames. Instead of list-format pros/cons or generic hedging, these traces teach models to identify real tension, make conditional commitments, and reach specific resolutions.
Version Note
This dataset contains only v1 traces — the clean, non-fabricating training data. An earlier version on this repo included augmented data from later pipeline… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/dialectic-reasoning-traces.dialectic-sft-against-only-750
Dialectic SFT — Against-Only (750)
750 supervised fine-tuning conversations that teach a model the structured
"dialectical" output format: a set of candidate positions [pN] followed by
against-claims [cN] against pM: that critique those positions. This is the
level-1, against-only stage (only against-claims, no for-claims or deeper tree
levels) — it bootstraps the format before GRPO reinforcement learning.
Row count
750 rows.
Schema
One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT
Turkish Dialectical Reasoning Dataset (Sokrates-ToT)
The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis.
Purpose of the Dataset
This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-47a94987
t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-47a94987
Dialectical reasoning evaluation for MuSR murder mysteries. The model acts as both
PROSECUTION and DEFENSE for each suspect, constructing the strongest possible case
for guilt and then the strongest possible case for innocence. The suspect with the
weaker defense is identified as the murderer.
Stylistically, the model uses formal legal argumentation: "the prosecution submits",
"the defense counters"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-musr-dialectical-murder-together_ai-qwen-qwen3-next-80b-a3b-inst-47a94987.Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/abdelfetteh/Tunisian_Dialectic_English_Derja.ndamix-cognitive-dialecticsphilosophy_dialectics_25kdialectic-rl-questions-10k
Dialectic RL Questions (10k)
10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a
dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion
debates, and general user requests) that the model is trained to answer by generating
multiple positions and against-claims in a structured "dialectical" format.
The prompts are drawn from public real-world sources: Reddit AITA
(r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.
