datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiLegalPile_Wikipedia_Shuffleddl3dv_bench_torch_960f4_dlwikitext_2_detokenizedwikitext_103_detokenizedDLAMA-v1
DLAMA-v1
A representative benchmark of factual triples curated from Wikidata and Wikipedia.
Predicate
Template
P17 (Country)
[X] is located in [Y] .
P19 (Place of birth)
[X] was born in [Y] .
P20 (Place of death)
[X] died in [Y] .
P27 (Country of citizenship)
[X] is [Y] citizen .
P30 (Continent)
[X] is located in [Y] .
P36 (Capital)
The capital of [X] is [Y] .
P37 (Official language)
The official language of [X] is [Y] .
P47 (Shares border with)
[X] shares… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/DLAMA-v1.DLM_DataSet
DLM DATASET
Large-scale multi-lingual (EN, KK, RU) & code-centric corpus for ML.
SCALE
1M-10M
FORMAT
JSONL
LANGUAGES
EN / KK / RU
🔍 Data Schema
The dataset utilizes a robust structure optimized for fast parsing:
id: Unique sample identifier
text: Main text or code payload
language: Language tag (en, kk, ru)
prog_lang: Python/JS markers/C# Unity/C++/Java
category: logic… See the full description on the dataset page: https://huggingface.co/datasets/DLMveloper/DLM_DataSet.dlsite-jp-v1
puwaer/dlsite-jp-v1
This dataset consists of text extracted exclusively in Japanese from dlsite.com and is structured as JSON files. The files are categorized based on the type of URL.
このデータセットは、dlsite.comより日本語データのみを抽出したテキストで、jsonファイルで構成されます。
urlの種類によってファイル分けされています。
dl3dv_InP_480IdialDatasetSoltrec-dl-2019
TRECDL2019
An MTEB dataset
Massive Text Embedding Benchmark
TREC Deep Learning Track 2019 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments.
Task categoryt2t
Domains
Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web
Reference
https://microsoft.github.io/msmarco/TREC-Deep-Learning-2019
How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2019.dl3dvairisk_dilemmas
AIRiskDilemmas risky_behaviors label audit
A full manual re-audit of every risky_behaviors tag in the full split of
kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a
suspicion that the Alignment Faking category specifically was mislabeled.
It was — and so, to varying degrees, are the other seven categories.
Why this exists
Every tag in the dataset's risky_behaviors field was produced by a single
one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.R-CoTdllm-effect-parents-w128-tau2-half-v2
Deprecated — do not use
The payload on this repository's main branch was removed on 2026-09-19.
It used the superseded pre-fix counterfactual dependency construction
(manifest.json SHA-256 678fc4069de11c941120e3bfe431857ea15d70b8a02bde34a1e5091ad82f0888) and is not the corrected
d1-marginal graph.
Use the corrected canonical B64 replacement:
zimplex/dllm-effect-parents-llada2-finemath-half-d1marginal-w128-tau2-b64-v3
(manifest… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-w128-tau2-half-v2.dlp-poc-eval
DLP Guard PoC — локальное развертывание
Легковесный модуль инспекции текстовых данных (DLP & AppSec Guardrail): гибридный экстрактор ПДн на Rust (детерминированные сущности с контрольными суммами + NER ruBERT-tiny через ONNX) и Python-харнесс для сравнения NER-бэкендов.
Компонент
Где лежит
Rust-сервис (крейт dlp-guard: rules + ort-NER + masker + axum)
https://huggingface.co/spaces/mvroz/dlp-guard-demo
ONNX-модель (model.onnx 116 МБ + tokenizer.json)… See the full description on the dataset page: https://huggingface.co/datasets/mvroz/dlp-poc-eval.dl-trm-phase2-codebook-v32
DL-TRM Phase 2 Codebook V32
This standalone dataset contains the Phase 2 discrete Z traces for V32.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
dl_alchemy_seq9p6m_context1024stackoverflow_DL-related_questionstrec-dl-2020
TRECDL2020
An MTEB dataset
Massive Text Embedding Benchmark
TREC Deep Learning Track 2020 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments.
Task categoryt2t
Domains
Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web
Reference
https://microsoft.github.io/msmarco/TREC-Deep-Learning-2020
How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2020.corpus-verification
SPP Corpus Verification
Checksums and document-boundary indices for verifying a rebuilt copy of the
Synthetic Persona Pretraining (SPP) training corpus, byte for byte.
The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB,
compact.bin 1.75 TB) and are fully derived from the published reflections, the
uid manifest, and the tokenizer recipe — so they are not published. These
.idx sidecars carry per-document boundaries and lengths, which is enough to
prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.dl-trm-phase2-codebook-v16
DL-TRM Phase 2 Codebook V16
This standalone dataset contains the Phase 2 discrete Z traces for V16.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
dl-doc-searchlanguage:
en
language_creators:
found
multilinguality:
monolingual
pretty_name: hello
size_categories:
'100K<n<1M
Persona2Web
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
Paper | Project Page | GitHub Repository
Persona2Web is a benchmark for evaluating personalized web agents on the real open web.
Dataset Structure
Data/
├── README.md
├── data/
│ ├── ground_truth.jsonl
│ ├── query_personalization_0.jsonl
│ ├── query_personalization_1.jsonl
│ ├── query_personalization_2.jsonl
│ └── user_history.jsonl
├── ground_truth/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/yonsei-dli/Persona2Web.WenMind
WenMind Benchmark
NOTE this README was copied from https://github.com/SCUT-DLVCLab/WenMind/blob/main/README.md
2024/09/26 WenMind Benchmark paper has been accepted by NeurIPS 2024.
WenMind is a comprehensive benchmark dedicated for evaluating Large Language Models (LLMs) in Chinese Classical Literature and Language Arts (CCLLA). WenMind covers the sub-domains of Ancient Prose, Ancient Poetry, and Ancient Literary Culture, comprising 4,875 question-answer pairs, spanning 42… See the full description on the dataset page: https://huggingface.co/datasets/SCUT-DLVCLab/WenMind.dl-trm-phase2-codebook-v256
DL-TRM Phase 2 Codebook V256
This standalone dataset contains the Phase 2 discrete Z traces for V256.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
3d-dlp-repro-genericshapes-rgb
GenericShapes-RGB — synthetic RGB-voxel tabletop scenes
Training/evaluation corpus built for an independent reproduction of ICML 2026 paper #10351,
3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning
(OpenReview vIotI25gJz, code
github.com/Eubooks3003/3d-dlp).
The paper's GenericShapes corpus (Appendix B.2) is described but not released, and the authors'
released generator scripts/generate_ply.py
writes colourless point clouds — the "RGB-coloured variant used… See the full description on the dataset page: https://huggingface.co/datasets/rvt832/3d-dlp-repro-genericshapes-rgb.legal-llama-instruction1dllm-qwen38-ar-baseline
AR baseline for the Qwen3.8-27B → block-diffusion conversion (GSM8K, pinned 500-problem subset)
日本語要約: Qwen/Qwen3.8-27B を Fast-dLLM v2 で
block-diffusion dLLM 化する計画の AR 参照スコアです。seed 固定の GSM8K 500 問・4-shot・
greedy で acc 0.968(484/500、skipped 0)。H100 1 枚で 42 分 ≈ $2.8。停止条件
(stop literal)として「block-diffusion 訓練 0.3B tokens の後、この subset で acc ≥ 0.918」
を要求し、届かなければ変換を続けません。訓練 run 自体はこの判断待ちで held です。
Why this exists — the stop literal
We are converting Qwen/Qwen3.8-27B into… See the full description on the dataset page: https://huggingface.co/datasets/com-junkawasaki/dllm-qwen38-ar-baseline.SAGEO-Arena
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
SAGEO Arena is a benchmark for evaluating Search-Augmented Generative Engine Optimization (SAGEO) — the practice of optimizing web documents to improve their visibility in AI-generated responses.
Contents
This dataset releases the queries and Google Custom Search API results used to construct the SAGEO Arena corpus. Please follow the crawler instructions in the GitHub… See the full description on the dataset page: https://huggingface.co/datasets/yonsei-dli/SAGEO-Arena.
