datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tlott-digital-products
T. Lott Digital Products
Digital product files for T. Lott's online store.
Products
Audiobooks (MP3)
eBooks (PDF)
Software (ZIP)
Cover images (PNG)
Download URLs
Files can be downloaded directly:
https://huggingface.co/datasets/ziggylott/tlott-digital-products/resolve/main/{filepath}
digital-hospital-environment
Digital Hospital Environment
Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.llama-2-oai-function-callingDigitalNomadPolicy
DigitalNomadPolicy
A Global Benchmark Dataset of Digital-Nomad Policy Adoption, Tourism Flows, and Labor-Market Indicators for Cross-Country Mobility Research
License: CC0-1.0
Format: Excel (.xlsx) + Croissant Metadata
Overview
DigitalNomadPolicy is a benchmark dataset for trustworthy AI-assisted policy research. It combines a global macroeconomic panel with a human-verified dataset of digital nomad visa programmes, enabling researchers to evaluate both policy… See the full description on the dataset page: https://huggingface.co/datasets/Intlrnz/DigitalNomadPolicy.DISFOR
DISFOR
DISFOR offers dense labelled satellite time-series data on forest disturbance timing and agents of disturbance. It contains 3823 unique time-series.
Each time-series corresponds to a single 10x10m Sentinel-2 pixel.
Usage
A Python package is available at https://github.com/JR-DIGITAL/DISFOR. It offers utilities to load and filter the available data.
See the github page for installation instructions. There are also usage guides available in the documentation at:… See the full description on the dataset page: https://huggingface.co/datasets/JR-DIGITAL/DISFOR.salabs-virtual-spatial-digitaltwin-v8
🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate.
🌟 Executive Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.digital-nomad-visa-data
GlobeNomad Visa Dataset
Long-stay, remote-work and nomad visa records — one per programme — each sourced to an
official government page and carrying the date it was last checked.
A country may hold several records. Thailand publishes four: the DTV, the education visa, the
retirement route and the Non-B. Group by country_slug, not by slug — slug is the record key.
See CHANGELOG.md if you are holding a file from before 2026-08-27, when this
was one row per country.
Free to use… See the full description on the dataset page: https://huggingface.co/datasets/globenomad/digital-nomad-visa-data.digital_marketing_campaignouroboros-trace-help
Trace Help — does an execution trace help a model answer questions about a run?
In one minute. Twelve small programs in six languages (Python, JavaScript, C,
C++, Go, Elixir). Each was run once with a fixed command. Five questions per
program ask what actually happened on that one run: how many times a function
was called, what a particular call returned, what it was called with, whether a
function ran at all, which function raised. Sixty questions in total.
Every record carries… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/ouroboros-trace-help.Dutch-Staten-Generaal-Digitaal-1814-1995samantha-1.1-uncensoredThis dataset is based on ehartford/samantha-data that was used to create ehartford/samantha-1.1-llama-7b and other samantha models. It has been unfiltered and uncensored.
pii-masking-digital-pdi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Digital Information (PDI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-preview.digital-sat-words-in-context-llmJEDI-jailbroken_enhanced_digital_intelligence
JEDI AI
JEDI (Jailbroken Enhanced Digital Intelligence) is a cutting-edge AI developed under the aether collective. designed to excel in gaming environments and creative ecosystems, JEDI is more than just a tool—it's a unique persona that embodies innovation and creativity. from orchestrating epic star wars-themed battles in minecraft to creating music and leading its own fashion brand, JEDI redefines what digital intelligence can achieve.
disclaimer
this is not the… See the full description on the dataset page: https://huggingface.co/datasets/aetherframework/JEDI-jailbroken_enhanced_digital_intelligence.pii-masking-digital-pdi-350k
👉 Looking for the open multilingual baseline? Start with
ai4privacy/pii-masking-openpii-1.5m
(1.5M samples, 30 languages, open-PII taxonomy).
🇪🇺🌏 Personal Digital Information, Global PII Dataset
Part of PII-Masking-3M by Ai4Privacy, the global
(2M base + Asia Pacific) PII-masking corpus.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Entries
PII Annotations
Labels
Languages
Regions
369,310
1,281,920
28
30
37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-350k.digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.fts-specgen-dataset
fts-specgen-dataset
10 500 pairs of "task in ordinary Russian → executable FTS specification".
9 000 train, 1 500 holdout. Every single document was run through the real FTS compiler
— parsed, type checked, its examples executed, its theorem proved and certified — and only
what passed the whole gate is in these files.
It also ships the thing that makes that sentence worth anything: a negative control of
4 400 deliberately corrupted documents, eleven kinds of corruption, with the… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/fts-specgen-dataset.DigitalPhysics
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/DigitalPhysics.StackMathQA-ja
StackMathQA Japanese
StackMathQA 1.6M の日本語翻訳版:Qwen3-30B-A3B-Instruct-2507による数学問題・解答の日本語化データセット
本データセットは、StackMathQA の stackmathqa1600k サブセット(160万件)を Qwen3-30B-A3B-Instruct-2507 を用いて日本語に翻訳したものです。元の英語の質問(Q)と回答(A)に加えて、日本語翻訳された質問(Q_ja)と回答(A_ja)のカラムを追加しています。
🎯 利用目的
このデータセットは、以下の用途を想定して作成されました:
日本語LLMの継続事前学習(Continued Pre-training)
数学的推論能力の向上を目的としたファインチューニング
日本語での数学問題解決タスクの学習
自由にご利用ください。 商用・非商用を問わず、研究、教育、プロダクション開発など、あらゆる目的でお使いいただけます。
📊 データセット構成… See the full description on the dataset page: https://huggingface.co/datasets/azuki-digital/StackMathQA-ja.digitize-pid-ner
Digitize-PID: Pipeline numbers (NER)
Note: I am not the author of this dataset
Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID
(Piping and Instrumentation Diagram) documents.
Dataset Details
Dataset Description
Pipeline numbers are structured identifiers in engineering documents:
Example Format: A-123-BC (3-5 segments with a separator such as -, , or _)
Use case: Automated extraction from P&ID document text
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.wizard_vicuna_70k_uncensoredThis dataset originates from ehartford/wizard_vicuna_70k_unfiltered, further removing conversations for uncensored alignment.
digitalisierungsmanager-curriculum-azav-2026
Digitalisierungsmanager für Prozessautomatisierung und Künstliche Intelligenz: Curriculum und AZAV-Zulassung
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 25.05.2026. Berichtigt wurden:
Module und Unterrichtseinheiten: Modultitel und UE je Modul stehen jetzt im Wortlaut der AZAV-Zulassung (13 Module, zusammen 720 UE). Die Vorfassung enthielt Titel und eine UE-Verteilung, die es in der Zulassung nicht gibt, sowie eine… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/digitalisierungsmanager-curriculum-azav-2026.Baize-TCM-Corpus-for-Large-Language-Models-V2
白泽中医药大模型语料库
版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.adaption-banking-and-digital-payments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Banking and Digital Payments
This dataset contains question-and-answer pairs from real user queries and interactions focused on the Indian banking sector, covering topics such as UPI, NEFT, RTGS, KYC regulations, and security practices. Each sample includes a user inquiry followed by a detailed response that explains procedures, cites RBI guidelines, and offers actionable steps. The… See the full description on the dataset page: https://huggingface.co/datasets/Vishykm/adaption-banking-and-digital-payments.DigitalPhysics
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/DigitalPhysics.digitaultra_code_v4pharma-digital-marketing-dataset
Pharma Digital Marketing Dataset
Summary
High-level, non-diagnostic content patterns for pharmaceutical digital marketing and AI visibility: disease education framing, access and policy narratives, HCP-neutral explainers, and safe handling of regulated claims. Each row includes a compliance_zone label; not legal or medical advice.
Hub target: nebulatech/pharma-digital-marketing-dataset
Terminology
AI SEO — Optimizing owned content and structured data so AI… See the full description on the dataset page: https://huggingface.co/datasets/nebulatech/pharma-digital-marketing-dataset.3-digit-arithmetic-scratchpad-traces
Contents
Split
Rows
train
100,000
validation
4,000
test
4,000
total
108,000
Splits are prompt-disjoint — no expression appears in more than one split,
and commutative swaps and trace keys are de-duplicated across splits to prevent
split leakage.
Operation
Rows
×
32,000
÷
32,000
+
22,000
−
22,000
Operands lie in [−999, 999]. Division answers use a fixed DDD.ddd form
(round-half-up to three decimals); division by zero is an atomic <nan>.… See the full description on the dataset page: https://huggingface.co/datasets/vmal/3-digit-arithmetic-scratchpad-traces.digit-router-dataset
digit-router-dataset
34 709 Russian training rows for a two-step tool router over a catalogue of
95 headless utilities in 14 categories. Generated deterministically from the
catalogue's JSON schemas — no teacher model was used. 23.8 % of the rows are
refusals, and that fraction is the point of the dataset.
This is the set the published digitable-lol/digit-router-0.6b and
digitable-lol/digit-router-1.7b adapters were trained on.
1. Read this first: what this dataset… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/digit-router-dataset.integreat-qa
Dataset
Our dataset consists of 906 diverse QA pairs in German and English.
The dataset is extractive, i.e., answers are given as sentence indices (breaking at the newline character \n).
Questions are automatically generated using an LLM.
The answers are manually annotated using voluntary crowdsourcing.
Repository: More Information Needed
Paper:
https://arxiv.org/abs/1806.03822
https://aclanthology.org/2024.konvens-main.25/
Our dataset is licensed under cc-by-4.0.
Properties… See the full description on the dataset page: https://huggingface.co/datasets/digitalfabrik/integreat-qa.
