datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdxl-models
Aisha-AI.com 💜
A NSFW Social Network powered by AI Characters
The models saved in this dataset are currently being used, or have been used at some point, to generate images and videos.
The dataset is public and can be used as a backup or alternative to more unstable servers (like the unfortunate Civitai).
lora_testogiri-bokete
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/ogiri-bokete", split="train")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。(8/30現在。ハッカソン当日までに増やす可能性があります。)
タスク… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-bokete.sdxl-loraspolemo2-officialPolEmo 2.0: Corpus of Multi-Domain Consumer Reviews, evaluation data for article presented at CoNLL.flux-2-klein-modelsfishtest_pgns
PGNs of Stockfish playing LTC games on Fishtest
This is a collection of computer chess games played by the engine
Stockfish on
Fishtest under strict
LTC, that
is at least 40 seconds base time control for each side.
The PGNs are stored as test-Id.pgn.gz in the directories YY-MM-DD/test-Id.
The moves in the PGNs are annotated with comments of the form {-0.91/21 1.749s},
indicating the engine's evaluation, search depth and time spent on the move.
Each directory also contains the… See the full description on the dataset page: https://huggingface.co/datasets/official-stockfish/fishtest_pgns.flux-dev-modelsrefute
Can AI read new science honestly?
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.Official_LLM_System_Prompts
Official LLM System Prompts
This short dataset contains a few system prompts leaked from proprietary models. Contains date-stamped prompts from OpenAI, Anthropic, MS Copilot, GitHub Copilot, Grok, and Perplexity.
master-binpacksA stockfish binpack collection used for Neural Network training for https://github.com/official-stockfish/nnue-pytorch.
ioai-2026-find-the-order
IOAI 2026 — Find the Order
Official contest data for Find the Order, task 1 (Day 1) of the IOAI 2026 Individual Contest, held in Astana, Kazakhstan.
Spoken English dialogues are segmented into speaker turns, one .wav per turn, and shuffled. Contestants reconstruct the original turn order. prefix.json in each dialogue folder gives the first two chunk indexes, fixing the reading direction.
The task statement, translations, baseline and grader live in the IOAI-2026 GitHub… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/ioai-2026-find-the-order.IDDAW_OFFICIAL
IDD-AW: India Driving Dataset – Adverse Weather
Semantic segmentation benchmark for autonomous driving in rain, fog, low-light,
and snow, with paired RGB + near-infrared (NIR) frames and dense Level-3
semantic labels (26 classes).
TODO before publishing: confirm and set the correct license / citation for
the original IDD-AW release (see iddaw.github.io and
the WACV 2024 paper "IDD-AW: A Benchmark for Safe Semantic Segmentation in
Adverse Weather"). This card currently marks the… See the full description on the dataset page: https://huggingface.co/datasets/Furqan7007/IDDAW_OFFICIAL.RoboTwin-adjust_bottle-official-demo_clean50-Pi0_processed-datauuno_mid-training_fi_official
Mid-training Finnish Corpus
Normalized text corpus.
Schema
id: globally unique normalized id
text: training text
source: original Hugging Face dataset id
metadata: JSON string containing source-specific metadata
Configs
fi
wds_sun397_official_splitsbokeh-eval-official
bokeh-eval-official
Bokeh synthesis artifacts: all-in-focus input, ground-truth bokeh, the
best-K render and the full K sweep, one split per benchmark.
Generated by inference/bokeh_net.py.
OlympiadBench-official
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
📖 arXiv | GitHub
Dataset Description
OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/OlympiadBench-official.reachy-mini-official-app-storebackups2025_DCASE_AudioQA_Official
Audio SFT / Post-Training Data
The proposed audio question answering (AQA) dataset
with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA)
DCASE 2025 Task Description
Audio QA Model Baseline
Watkins Marine Mammal Sound Database
"Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
📢 Post-Challenge Research Note
While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.long-tail-planning-with-language-officialioai-2026-double-agent-dilemma
IOAI 2026 — Double Agent Dilemma
Official contest data for Double Agent Dilemma, task 4 (Day 2) of the IOAI 2026 Individual Contest, held in Astana, Kazakhstan.
Two pretrained image classifiers — a ResNet18 (CNN) and a ViT-Tiny (Transformer) — both reach 100% accuracy on the provided images. The task exploits where the two architectures disagree.
The task statement, translations, baseline and grader live in the IOAI-2026 GitHub repository. This repository holds data only.… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/ioai-2026-double-agent-dilemma.babel-official
BABEL Official Local Layout
This directory is a cleaned local mirror of the official BABEL v1.0 labels and
the AMASS subsets required by those labels.
Layout
archives/: original downloaded archives, kept unchanged for provenance.
labels/babel_v1.0_release/: official BABEL JSON splits.
amass/: extracted AMASS motion parameter files.
processed/manifests/*.jsonl: normalized records with resolved local AMASS
paths, frame counts, fps, segment labels, and rewritten… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/babel-official.Selfie_and_Official_ID_Photo_Dataset12,000+ people, 150,000+ images. Selfie with ID dataset for KYC verification, face identification and biometric training. Selfies paired with 2 official ID photos (passport, ID card, driver's license, residence permit). 10-15 photos per person with balanced demographics across ethnicity (Caucasian, Black, Asian, Latin American), gender and age (18-65).
Contact us and share your feedback - recieve additional samples for free! 😊
Key Highlights:
12,000+ real individuals… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Selfie_and_Official_ID_Photo_Dataset.persian-official-letters-synth
Persian Official Letters (Synthetic)
Draft dataset card. Copy this file to the root of the export tree and
fill in the dataset_info / configs blocks from
dataset_info.json. See docs/publishing_to_huggingface.md for the export
and upload procedure. Replace <owner> in FA_OCR_HF_REPO_ID before pushing.
A 10,000-letter synthetic corpus of Persian administrative correspondence,
rendered as scanned-looking documents with full layout ground truth: page
images, per-region bounding… See the full description on the dataset page: https://huggingface.co/datasets/shahriarhd/persian-official-letters-synth.lmeval-official-format
lm-evaluation-harness results, native format
EleutherAI lm-evaluation-harness results kept in the harness's own native
output format — arc_easy, arc_challenge and hellaswag, each with results (accuracy and stderr,
normalised and raw), configs, versions and n-shot. Kept unmodified precisely so a stranger can re-run
the same task on their own hardware and diff the files directly.
The live board is the authority
GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/lmeval-official-format.health-conditions-among-children-under-age-18-by-s
Health conditions among children under age 18, by selected characteristics: United States
Description
NOTE: On October 19, 2021, estimates for 2016–2018 by health insurance status were revised to correct errors. Changes are highlighted and tagged at https://www.cdc.gov/nchs/data/hus/2019/012-508.pdf
Data on health conditions among children under age 18, by selected population characteristics. Please refer to the PDF or Excel version of this table in the HUS 2019 Data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/health-conditions-among-children-under-age-18-by-s.IOAI2025
International Olympiad in Artificial Intelligence (IOAI 2025, Beijing, China)
About IOAI 2025
The 2nd International Olympiad in Artificial Intelligence (IOAI 2025) took place in Beijing, China, from August 2 to 9, 2025, hosted by Beijing National Day School (BNDS) under the patronage of UNESCO.
Contest Rules: Full rules encompassing the Individual, Team, and GAITE contests are available here.
Syllabus: The official syllabus outlining the AI topics contestants should… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/IOAI2025.b150_official_train
b150_official_train
This release contains the high-confidence b150 training split used by MAGIC-TTS,
with standalone MFA word-level alignments distributed separately from prepared Arrow
artifacts.
Related links:
GitHub repo: https://github.com/yongaifadian1/MAGIC-TTS
arXiv paper: https://arxiv.org/abs/2604.21164
Open-source model: https://huggingface.co/maimai11/MAGIC-TTS
Online demo: https://yongaifadian1.github.io/MAGIC-TTS/
Stats:
Total samples: 201986
English: 78363
Chinese:… See the full description on the dataset page: https://huggingface.co/datasets/maimai11/b150_official_train.
