datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refute
Can AI read new science honestly?
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.OlympiadBench-official
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
📖 arXiv | GitHub
Dataset Description
OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/OlympiadBench-official.nemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.aime24-official
AIME 2024 — official wording, figures retained
All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II),
transcribed from the official exam text with every figure retained as Asymptote source.
This exists because the circulating text-only versions of AIME 2024 are not faithful to the
official problems, and at least one problem in them cannot be solved as written.
Why this dataset exists
While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.igakuqa-subset-curated
IgakuQA Curated Subset (Text-Only)
This dataset contains 66 carefully selected text-only samples from the IgakuQA medical exam dataset,
curated using advanced difficulty assessment and model evaluation techniques.
Dataset Description
This is a high-quality subset of the IgakuQA Japanese medical exam dataset, selected based on:
Difficulty Score: Measures how challenging the sample is for AI models
Consistency Score: Evaluates response consistency across different models… See the full description on the dataset page: https://huggingface.co/datasets/japan-ai-official/igakuqa-subset-curated.dataagentbench-derived-official54-altimate-prompt-shell
DataAgentBench-Derived Official54 Altimate Prompt Shell
This public dataset contains 54 DataAgentBench-derived task prompt-shell rows compiled for the Altimate-centered DAB sandbox runtime. It is a derived compile/export artifact, not the official raw DataAgentBench release.
It is intended as a portable task/prompt/manifest source for downstream Altimate teacher rollouts, SFT construction, or RL data compilation. It is not a completed rollout dataset: the rows do not contain… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/dataagentbench-derived-official54-altimate-prompt-shell.aba-official-curriculum-sft
ABA Official Curriculum SFT
Structured supervision dataset derived from official QABA curriculum sources for:
ABAT
QASP-S
QBA
Files
official_lessons.jsonl
official_qa.jsonl
official_mcq.jsonl
official_curriculum_sft.jsonl
official_curriculum_train.jsonl
official_curriculum_eval.jsonl
manifest.json
Intended use
This dataset is intended for:
instruction tuning on official ABA curriculum content
grounded lesson planning
grounded question answering
grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.yeoncheon-market-info-official
연천군 전통시장 현황 데이터셋 (Yeoncheon Market Info)
이 데이터셋은 대형 언어 모델(LLM)이 경기도 연천군과 같은 소규모 지역 정보에 대해 일으키는 환각(Hallucination) 현상을 방지하고, 정확한 정보를 제공하기 위해 제작되었습니다.
1. 데이터 출처
출처: 경기도 공공데이터포털 - 전통시장 및 상점가 현황
기준일: 2023년 12월 31일
2. 데이터 구성
instruction: 연천군 시장 관련 질문 (대화체 및 정보 요청 포함)
output: 공공데이터에 근거한 정확한 답변
source: 데이터의 공식 출처
3. 활용 목적
한국 로컬 지리 및 행정 정보에 대한 AI 학습 및 검증용
연천군 지역 경제 및 전통시장 정보 대중화
