datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-mono
Coding agent session traces for Pi
This dataset contains redacted coding agent session traces collected while working on the Pi OSS project.
Canonical source repository: git@github.com:earendil-works/pi.git
The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-mono.aaac
Dataset Card for Artificial Argument Analysis Corpus (AAAC)
Dataset Summary
DeepA2 is a modular framework for deep argument analysis. DeepA2 datasets contain comprehensive logical reconstructions of informally presented arguments in short argumentative texts. This document describes two synthetic DeepA2 datasets for artificial argument analysis: AAAC01 and AAAC02.
# clone
git lfs clone https://huggingface.co/datasets/debatelab/aaac
import pandas as pd
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/aaac.pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.AAA: >
The AAA Unified Intelligence Substrate — canonical doctrine, constitutional
floors, evaluation benchmarks, and governance schemas for the arifOS Double
Helix Constitutional AI kernel. AGI · ASI · APEX. DITEMPA BUKAN DIBERI.---
🗺️ Position in I-ARIF Governance Stack
This dataset is part of the arifOS constitutional governance training-and-evaluation pipeline — a closed-loop alignment substrate.
#
Dataset
Role
Downloads
License
1
AAA
Constitutional… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/AAA.pi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.playdate-games
Coding agent session traces for aaaaliou/playdate-games
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/playdate-games.pi-playdate
Coding agent session traces for aaaaliou/pi-playdate
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-playdate.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-playdate.paulg-tweets
Paul Graham (@paulg) tweet archive
Offline archive of @paulg tweets and paulgraham.com essays, built for an Ask @paulg–style search/chat app.
Private dataset — not for redistribution without considering X/Twitter Terms of Service and content ownership. Unofficial; not affiliated with Paul Graham or X.
Credits
Curated, ingested, and uploaded by Ahmet Dedeler (Hugging Face).
Paul Graham wrote the tweets and essays. Ahmet wrote the Python that argued with X rate… See the full description on the dataset page: https://huggingface.co/datasets/aaahmet/paulg-tweets.alpaca_data_clean
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/aaaalon/alpaca_data_clean.appsAPPS is a benchmark for Python code generation, it includes 10,000 problems, which range from having simple oneline solutions to being substantial algorithmic challenges, for more details please refer to this paper: https://arxiv.org/pdf/2105.09938.pdf.aaa
RNAcentral Export
Export Metadata
Query: (("GO:2000352") AND (entry_type:"Sequence" OR entry_type:"Gene"))
Export date: 18 February 2026 13:42:28
RNAcentral version: v24
Number of sequences: 6
Description
RNAcentral is a free, public resource that offers
integrated access to a comprehensive and up-to-date set of non-coding RNA
sequences provided by a collaborating group of Expert Databases.
License
The data is available under the
CC0 1.0 Universal… See the full description on the dataset page: https://huggingface.co/datasets/afg1/aaa.EU-Air-Passenger-Rights-Instruction-Dataset-AAA
EU Air Passenger Rights Instruction Dataset (Production-Grade Sample)
This repository contains a premium, human-curated instruction dataset focused on European Union law (Regulation (EC) No 261/2004 and relevant Court of Justice case law).
It has been converted and unified into the production-standard ShareGPT/Alpaca conversational format, making it instantly ready for supervised fine-tuning (SFT) of open-weights LLMs (such as LLaMA 3, Mistral, or Qwen) or for indexing in… See the full description on the dataset page: https://huggingface.co/datasets/kenax69/EU-Air-Passenger-Rights-Instruction-Dataset-AAA.alpaca-cleand
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/aaaalon/alpaca-cleand.AAAI_Swahili_dataset
README for Swahili Translated Dataset from Toloka
Dataset Description
This dataset is a Dolly 15k translated from English to Swahili, filtered and processed using the Toloka platform. It includes various contexts, responses, and instructions from diverse domains, providing a rich resource for natural language processing tasks, particularly for those focusing on the Swahili language.
Data Fields
task_id: A unique identifier for each task in the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/ortofasfat/AAAI_Swahili_dataset.d2c-aaai27-data
D2C: Diagnostic-to-Decision Calibration — Data Release
Anonymous data supplement for the AAAI 2027 submissionD2C: Diagnostic-to-Decision Calibration for LLM Reasoning Monitors.
Contents
Path
Description
rq1/annotated.json
Primary MATH-Hard traces (n=1,000, DS-Math-7B-RL) with step-level flow features
actionability_bench/actionability_bench.json
ActionabilityBench full release (4,500 traces, 5 models, step features)… See the full description on the dataset page: https://huggingface.co/datasets/papers-data/d2c-aaai27-data.1b-model-eval
1B Model Eval — 16 Sub-2B LLMs Benchmark
A comprehensive G-Eval (Liu et al., EMNLP 2023) benchmark of 16 small language models under 2B parameters, covering 6 dimensions across 33 test items.
Judge model: DeepSeek-v4-pro
Files
File
Description
nightly_raw_<model>.json
Raw outputs for each model — 33 generated items across 6 dimensions
nightly_eval_results.json
Judge scores and detailed reasoning for all 16 models
summary.csv
One-row-per-model score… See the full description on the dataset page: https://huggingface.co/datasets/aaaaaaapluto/1b-model-eval.mydata
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/aaaalon/mydata.AA2
