datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-mono
Coding agent session traces for Pi
This dataset contains redacted coding agent session traces collected while working on the Pi OSS project.
Canonical source repository: git@github.com:earendil-works/pi.git
The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-mono.lorapi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.aaaAAA
Writing-ai
Crafting Historical Romance Novels with Deep Psychological Profiling in Lesbian Relationships
AAA: >
The AAA Unified Intelligence Substrate — canonical doctrine, constitutional
floors, evaluation benchmarks, and governance schemas for the arifOS Double
Helix Constitutional AI kernel. AGI · ASI · APEX. DITEMPA BUKAN DIBERI.---
🗺️ Position in I-ARIF Governance Stack
This dataset is part of the arifOS constitutional governance training-and-evaluation pipeline — a closed-loop alignment substrate.
#
Dataset
Role
Downloads
License
1
AAA
Constitutional… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/AAA.metacam-glasses
metacam-datasets
Currently holding glasses images
pi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.playdate-games
Coding agent session traces for aaaaliou/playdate-games
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/playdate-games.Multi-Opthalingua
Cite
Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs:
@misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing,
title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.aaaa111223
VoxAging
VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin
📄 Paper: https://arxiv.org/pdf/2505.21445
Description
The VoxAging dataset is a large-scale longitudinal audio-visual corpus designed for studying long-term speaker aging and temporal variations in speech.It consists of recordings from 293 speakers, including 226 English speakers (112 female, 114 male) and 67 Mandarin speakers (23 female, 44… See the full description on the dataset page: https://huggingface.co/datasets/belztjti/aaaa111223.pi-playdate
Coding agent session traces for aaaaliou/pi-playdate
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-playdate.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-playdate.chinese_AAAI_Math
Dataset Card for "chinese_AAAI_Math"
More Information needed
AAAI2024check_dataStrawberry_12-dataset
🍓 Strawberry_12 Dataset
Strawberry_12 is a large-scale image dataset created for detecting strawberry diseases, pests, and nutrient deficiencies. It is designed to support research in smart agriculture, plant health diagnostics, and AI-based crop monitoring systems.
Images: 3,906
Annotations: 16,107 bounding boxes
Categories: 12 (covering common diseases, pests, and deficiency symptoms)
Annotation Format: YOLO (.txt per image with bounding boxes)
Collection Period:… See the full description on the dataset page: https://huggingface.co/datasets/aaaqwqdwads/Strawberry_12-dataset.IndustryShapes
IndustryShapes
Project Page | Paper
IndustryShapes is a new benchmark dataset tailored for 6D object pose estimation in industrial settings. Targeting the challenges of textureless objects, reflective surfaces, and complex assembly tools, this dataset provides high-quality RGB-D data with precise annotations to advance the state of the art in robotic manipulation.
Dataset Features
Unlike traditional datasets focused on household products, IndustryShapes introduces… See the full description on the dataset page: https://huggingface.co/datasets/aaabab/IndustryShapes.aaai27-hotpotqa-fullwiki-original
HotpotQA FullWiki frozen original data
Private reproducibility snapshot for the AAAI 2027 experiments.
This repository stores the exact Hugging Face hotpotqa/hotpot_qa, fullwiki parquet files used to construct our train, calibration, and official evaluation task identities. It does not contain generated trajectories, Writer SFT/RL examples, or model-selected subsets.
Files and rows
fullwiki/train-00000-of-00002.parquet: 45,224 rows.… See the full description on the dataset page: https://huggingface.co/datasets/sastpg/aaai27-hotpotqa-fullwiki-original.AAAI2026elixirDatasetsunstructed_nuScenesspark-capacity-boundary-study
Capacity-boundary optimization with Spark-X2.5-1.7B
This project contains an original evaluation for HER Hack-Astron #6. The report is in DISCUSSION.md. Creating or publishing these artifacts is not an award or payment.
The experiment checks whether increasing a six-item 0/1 knapsack's capacity by one causes the model to find the new optimum. Four seeded item families produce eight mathematical instances, each in English and Chinese. Every prompt runs once with thinking off and… See the full description on the dataset page: https://huggingface.co/datasets/aaaded/spark-capacity-boundary-study.AAAI2023Autopentest-DatasetAAAI2025vscode_bugs_cleanedurdu-tts-speaker3-preparedMCIPmusicAAAI2022AAAI2018
