datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sample_v6Lyrical_v6_ru2en_Poetry_MeterMatched_DPO_alpaca_gpt_Instruction
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Translated by a Poet-Translator from Russian to English
Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns
NEWLY EDITED VARIANT 6: 1775 rows/items
Re-balanced, refined, standardized, and substantially expanded.
Csv version, loosely aimed at Gpt-Oss/Harmony
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_v6_ru2en_Poetry_MeterMatched_DPO_alpaca_gpt_Instruction.unitares-verdict-counterfactual-v6.8
UNITARES Verdict Counterfactual — Paper v6.8 Reproducibility Kit
Reproducibility artifacts for §11.6 of UNITARES: Information-Theoretic Governance of Heterogeneous Agent Fleets (Wang, 2026).
Paper: CIRWEL/unitares-paper-v6, tag paper-v6.8.1
Paper DOI (concept): 10.5281/zenodo.19647159
Data DOI (this release, v6.8.1-repro): 10.5281/zenodo.19705151
GitHub: CIRWEL/unitares-repro-v6
Source: UNITARES production governance database, 30-day rolling window of core.agent_state
License:… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/unitares-verdict-counterfactual-v6.8.llm_instruction_code_v6program_generation_v6my-notebook-backup-v6llm_instruction_code_V6.1english_karakalpak_parallel_corpus_v6
English-Karakalpak Parallel Corpus v6.0
Dataset Description
English-Karakalpak Parallel Corpus v6.0 is a high-quality, finalized parallel dataset containing over 31,000 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems.
Language(s): English (en), Karakalpak (kaa)
Format: CSV (Comma-Separated Values)
License: MIT
Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v6.final_train_v6humanoid-step-planning-labels-v6humanoid-safety-response-labels-v6astra_intent_v6cnn-v6
New Official train/val for cnn with restricted article length.
humanoid-vision-context-labels-v6dota_training_data_v6Cot_v6Reddit-Stock-V6TravelPlanner_split_v610a_v6.csv
