datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AAA: >
The AAA Unified Intelligence Substrate — canonical doctrine, constitutional
floors, evaluation benchmarks, and governance schemas for the arifOS Double
Helix Constitutional AI kernel. AGI · ASI · APEX. DITEMPA BUKAN DIBERI.---
🗺️ Position in I-ARIF Governance Stack
This dataset is part of the arifOS constitutional governance training-and-evaluation pipeline — a closed-loop alignment substrate.
#
Dataset
Role
Downloads
License
1
AAA
Constitutional… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/AAA.Multi-Opthalingua
Cite
Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs:
@misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing,
title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.aaai27-hotpotqa-fullwiki-original
HotpotQA FullWiki frozen original data
Private reproducibility snapshot for the AAAI 2027 experiments.
This repository stores the exact Hugging Face hotpotqa/hotpot_qa, fullwiki parquet files used to construct our train, calibration, and official evaluation task identities. It does not contain generated trajectories, Writer SFT/RL examples, or model-selected subsets.
Files and rows
fullwiki/train-00000-of-00002.parquet: 45,224 rows.… See the full description on the dataset page: https://huggingface.co/datasets/sastpg/aaai27-hotpotqa-fullwiki-original.aaaAAAI_Swahili_dataset
README for Swahili Translated Dataset from Toloka
Dataset Description
This dataset is a Dolly 15k translated from English to Swahili, filtered and processed using the Toloka platform. It includes various contexts, responses, and instructions from diverse domains, providing a rich resource for natural language processing tasks, particularly for those focusing on the Swahili language.
Data Fields
task_id: A unique identifier for each task in the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/ortofasfat/AAAI_Swahili_dataset.aaaa
