datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.digital-hospital-environment
Digital Hospital Environment
Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.Data-Analytics-Digital-Marketing-Project-Management-QA_DBouroboros-trace-help
Trace Help — does an execution trace help a model answer questions about a run?
In one minute. Twelve small programs in six languages (Python, JavaScript, C,
C++, Go, Elixir). Each was run once with a fixed command. Five questions per
program ask what actually happened on that one run: how many times a function
was called, what a particular call returned, what it was called with, whether a
function ran at all, which function raised. Sixty questions in total.
Every record carries… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/ouroboros-trace-help.la-serena-digital-geo-corpus
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1:… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/la-serena-digital-geo-corpus.Digital-Accountability-And-Transparency-Act-of-2014
Dataset Description
Maintainer: Terry Eppler
Ownership: US Federal Government
The Digital Accountability and Transparency Act of 2014 Question-Answer Dataset is an English-language instructional dataset containing questions and detailed answers concerning the Digital Accountability and Transparency Act of 2014, commonly known as the DATA Act.
The DATA Act was enacted as Public Law 113-101 on May 9, 2014. The legislation expanded the Federal Funding Accountability and… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Digital-Accountability-And-Transparency-Act-of-2014.StackMathQA-ja
StackMathQA Japanese
StackMathQA 1.6M の日本語翻訳版:Qwen3-30B-A3B-Instruct-2507による数学問題・解答の日本語化データセット
本データセットは、StackMathQA の stackmathqa1600k サブセット(160万件)を Qwen3-30B-A3B-Instruct-2507 を用いて日本語に翻訳したものです。元の英語の質問(Q)と回答(A)に加えて、日本語翻訳された質問(Q_ja)と回答(A_ja)のカラムを追加しています。
🎯 利用目的
このデータセットは、以下の用途を想定して作成されました:
日本語LLMの継続事前学習(Continued Pre-training)
数学的推論能力の向上を目的としたファインチューニング
日本語での数学問題解決タスクの学習
自由にご利用ください。 商用・非商用を問わず、研究、教育、プロダクション開発など、あらゆる目的でお使いいただけます。
📊 データセット構成… See the full description on the dataset page: https://huggingface.co/datasets/azuki-digital/StackMathQA-ja.DigitalPhysics
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/DigitalPhysics.digitalisierungsmanager-curriculum-azav-2026
Digitalisierungsmanager für Prozessautomatisierung und Künstliche Intelligenz: Curriculum und AZAV-Zulassung
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 25.05.2026. Berichtigt wurden:
Module und Unterrichtseinheiten: Modultitel und UE je Modul stehen jetzt im Wortlaut der AZAV-Zulassung (13 Module, zusammen 720 UE). Die Vorfassung enthielt Titel und eine UE-Verteilung, die es in der Zulassung nicht gibt, sowie eine… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/digitalisierungsmanager-curriculum-azav-2026.integreat-qa
Dataset
Our dataset consists of 906 diverse QA pairs in German and English.
The dataset is extractive, i.e., answers are given as sentence indices (breaking at the newline character \n).
Questions are automatically generated using an LLM.
The answers are manually annotated using voluntary crowdsourcing.
Repository: More Information Needed
Paper:
https://arxiv.org/abs/1806.03822
https://aclanthology.org/2024.konvens-main.25/
Our dataset is licensed under cc-by-4.0.
Properties… See the full description on the dataset page: https://huggingface.co/datasets/digitalfabrik/integreat-qa.DigitalPhysics
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/DigitalPhysics.digit-eval-tasks
digit-eval-tasks
400 Russian evaluation tasks — 250 main, 150 red-team — for a system that is
required to answer only from a deterministic tool, a verbatim corpus quote, or a formal
certificate, and to refuse otherwise.
⚠️ Read this before you use it as a benchmark
This set was used while developing the gates it measures. It is no longer an
independent measuring stick.
On this set the system reports 1.0 % false answers (4/400). On a genuinely
independent held-out… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/digit-eval-tasks.k12-digital-learning-platforms-research
K-12 Digital Learning Platforms: Research Records
227 structured records describing studies and surveys about digital learning platform
effectiveness in K-12 education. Each record carries a study title, year, type,
methodology, sample size, focus area, and a set of JSON-encoded findings and
recommendation fields.
Provenance is not verifiable - do not cite these as literature
These are structured records about studies, not the studies themselves, and their… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-digital-learning-platforms-research.ChatGPT-Prompts-for-Digital-Products
