CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dlwh /MultiLegalPile_Wikipedia_Shuffledtext100K<n<1M0 likes1.2k downloads4y agoHugging Face02lhmd /dl3dv_bench_torch_960textn<1K0 likes1.2k downloads11mo agoHugging Face03Knowing /f4_dltextn<1K0 likes1.1k downloads5mo agoHugging Face04dlwh /wikitext_2_detokenizedtext10K<n<100K1 likes605 downloads4y agoHugging Face05dlwh /wikitext_103_detokenizedtext1M<n<10M3 likes453 downloads4y agoHugging Face06AMR-KELEG /DLAMA-v1 DLAMA-v1 A representative benchmark of factual triples curated from Wikidata and Wikipedia. Predicate Template P17 (Country) [X] is located in [Y] . P19 (Place of birth) [X] was born in [Y] . P20 (Place of death) [X] died in [Y] . P27 (Country of citizenship) [X] is [Y] citizen . P30 (Continent) [X] is located in [Y] . P36 (Capital) The capital of [X] is [Y] . P37 (Official language) The official language of [X] is [Y] . P47 (Shares border with) [X] shares… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/DLAMA-v1.text100K<n<1M0 likes441 downloads1y agoHugging Face07DLMveloper /DLM_DataSet DLM DATASET Large-scale multi-lingual (EN, KK, RU) & code-centric corpus for ML. SCALE 1M-10M FORMAT JSONL LANGUAGES EN / KK / RU 🔍 Data Schema The dataset utilizes a robust structure optimized for fast parsing: id: Unique sample identifier text: Main text or code payload language: Language tag (en, kk, ru) prog_lang: Python/JS markers/C# Unity/C++/Java category: logic… See the full description on the dataset page: https://huggingface.co/datasets/DLMveloper/DLM_DataSet.text1M<n<10M1 likes191 downloads1mo agoHugging Face08puwaer /dlsite-jp-v1 puwaer/dlsite-jp-v1 This dataset consists of text extracted exclusively in Japanese from dlsite.com and is structured as JSON files. The files are categorized based on the type of URL. このデータセットは、dlsite.comより日本語データのみを抽出したテキストで、jsonファイルで構成されます。 urlの種類によってファイル分けされています。 texttext-generation1M<n<10M4 likes180 downloads2y agoHugging Face09wlyu /dl3dv_InP_480tabular1K<n<10K0 likes165 downloads5mo agoHugging Face10DLMveloper /IdialDatasetSoltext1M<n<10M1 likes150 downloads11d agoHugging Face11whybe-choi /trec-dl-2019 TRECDL2019 An MTEB dataset Massive Text Embedding Benchmark TREC Deep Learning Track 2019 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments. Task categoryt2t Domains Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web Reference https://microsoft.github.io/msmarco/TREC-Deep-Learning-2019 How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2019.texttext-retrieval1M<n<10M1 likes149 downloads11mo agoHugging Face12Yeqing0814 /dl3dvtextn<1K0 likes148 downloads1y agoHugging Face13dlab-spp /airisk_dilemmas AIRiskDilemmas risky_behaviors label audit A full manual re-audit of every risky_behaviors tag in the full split of kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a suspicion that the Alignment Faking category specifically was mislabeled. It was — and so, to varying degrees, are the other seven categories. Why this exists Every tag in the dataset's risky_behaviors field was produced by a single one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.text10K<n<100K0 likes143 downloads1mo agoHugging Face14dle666 /R-CoTimage100K<n<1M5 likes136 downloads2y agoHugging Face15zimplex /dllm-effect-parents-w128-tau2-half-v2 Deprecated — do not use The payload on this repository's main branch was removed on 2026-09-19. It used the superseded pre-fix counterfactual dependency construction (manifest.json SHA-256 678fc4069de11c941120e3bfe431857ea15d70b8a02bde34a1e5091ad82f0888) and is not the corrected d1-marginal graph. Use the corrected canonical B64 replacement: zimplex/dllm-effect-parents-llada2-finemath-half-d1marginal-w128-tau2-b64-v3 (manifest… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-effect-parents-w128-tau2-half-v2.textn<1K0 likes106 downloads7d agoHugging Face16mvroz /dlp-poc-eval DLP Guard PoC — локальное развертывание Легковесный модуль инспекции текстовых данных (DLP & AppSec Guardrail): гибридный экстрактор ПДн на Rust (детерминированные сущности с контрольными суммами + NER ruBERT-tiny через ONNX) и Python-харнесс для сравнения NER-бэкендов. Компонент Где лежит Rust-сервис (крейт dlp-guard: rules + ort-NER + masker + axum) https://huggingface.co/spaces/mvroz/dlp-guard-demo ONNX-модель (model.onnx 116 МБ + tokenizer.json)… See the full description on the dataset page: https://huggingface.co/datasets/mvroz/dlp-poc-eval.textn<1K0 likes94 downloads12d agoHugging Face17omrisap /dl-trm-phase2-codebook-v32 DL-TRM Phase 2 Codebook V32 This standalone dataset contains the Phase 2 discrete Z traces for V32. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes83 downloads4mo agoHugging Face18kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes79 downloads18d agoHugging Face19xPXXX /stackoverflow_DL-related_questionstabular10K<n<100K0 likes78 downloads3y agoHugging Face20whybe-choi /trec-dl-2020 TRECDL2020 An MTEB dataset Massive Text Embedding Benchmark TREC Deep Learning Track 2020 passage ranking task. The task involves retrieving relevant passages from the MS MARCO collection given web search queries. Queries have multi-graded relevance judgments. Task categoryt2t Domains Encyclopaedic, Academic, Blog, News, Medical, Government, Reviews, Non-fiction, Social, Web Reference https://microsoft.github.io/msmarco/TREC-Deep-Learning-2020 How to… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/trec-dl-2020.texttext-retrieval1M<n<10M1 likes74 downloads11mo agoHugging Face21dlab-spp /corpus-verification SPP Corpus Verification Checksums and document-boundary indices for verifying a rebuilt copy of the Synthetic Persona Pretraining (SPP) training corpus, byte for byte. The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB, compact.bin 1.75 TB) and are fully derived from the published reflections, the uid manifest, and the tokenizer recipe — so they are not published. These .idx sidecars carry per-document boundaries and lengths, which is enough to prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.textn<1K0 likes68 downloads1mo agoHugging Face22omrisap /dl-trm-phase2-codebook-v16 DL-TRM Phase 2 Codebook V16 This standalone dataset contains the Phase 2 discrete Z traces for V16. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes59 downloads4mo agoHugging Face23pinecone /dl-doc-searchlanguage: en language_creators: found multilinguality: monolingual pretty_name: hello size_categories: '100K<n<1M text100K<n<1M0 likes58 downloads4y agoHugging Face24yonsei-dli /Persona2Web Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History Paper | Project Page | GitHub Repository Persona2Web is a benchmark for evaluating personalized web agents on the real open web. Dataset Structure Data/ ├── README.md ├── data/ │ ├── ground_truth.jsonl │ ├── query_personalization_0.jsonl │ ├── query_personalization_1.jsonl │ ├── query_personalization_2.jsonl │ └── user_history.jsonl ├── ground_truth/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/yonsei-dli/Persona2Web.textn<1K0 likes58 downloads4mo agoHugging Face25SCUT-DLVCLab /WenMind WenMind Benchmark NOTE this README was copied from https://github.com/SCUT-DLVCLab/WenMind/blob/main/README.md 2024/09/26 WenMind Benchmark paper has been accepted by NeurIPS 2024. WenMind is a comprehensive benchmark dedicated for evaluating Large Language Models (LLMs) in Chinese Classical Literature and Language Arts (CCLLA). WenMind covers the sub-domains of Ancient Prose, Ancient Poetry, and Ancient Literary Culture, comprising 4,875 question-answer pairs, spanning 42… See the full description on the dataset page: https://huggingface.co/datasets/SCUT-DLVCLab/WenMind.textquestion-answering1K<n<10K5 likes57 downloads2y agoHugging Face26omrisap /dl-trm-phase2-codebook-v256 DL-TRM Phase 2 Codebook V256 This standalone dataset contains the Phase 2 discrete Z traces for V256. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes57 downloads4mo agoHugging Face27rvt832 /3d-dlp-repro-genericshapes-rgb GenericShapes-RGB — synthetic RGB-voxel tabletop scenes Training/evaluation corpus built for an independent reproduction of ICML 2026 paper #10351, 3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning (OpenReview vIotI25gJz, code github.com/Eubooks3003/3d-dlp). The paper's GenericShapes corpus (Appendix B.2) is described but not released, and the authors' released generator scripts/generate_ply.py writes colourless point clouds — the "RGB-coloured variant used… See the full description on the dataset page: https://huggingface.co/datasets/rvt832/3d-dlp-repro-genericshapes-rgb.tabularimage-segmentationn<1K0 likes57 downloads2mo agoHugging Face28dliu1 /legal-llama-instruction1text10K<n<100K1 likes53 downloads3y agoHugging Face29com-junkawasaki /dllm-qwen38-ar-baseline AR baseline for the Qwen3.8-27B → block-diffusion conversion (GSM8K, pinned 500-problem subset) 日本語要約: Qwen/Qwen3.8-27B を Fast-dLLM v2 で block-diffusion dLLM 化する計画の AR 参照スコアです。seed 固定の GSM8K 500 問・4-shot・ greedy で acc 0.968(484/500、skipped 0)。H100 1 枚で 42 分 ≈ $2.8。停止条件 (stop literal)として「block-diffusion 訓練 0.3B tokens の後、この subset で acc ≥ 0.918」 を要求し、届かなければ変換を続けません。訓練 run 自体はこの判断待ちで held です。 Why this exists — the stop literal We are converting Qwen/Qwen3.8-27B into… See the full description on the dataset page: https://huggingface.co/datasets/com-junkawasaki/dllm-qwen38-ar-baseline.tabularquestion-answeringn<1K0 likes48 downloads13d agoHugging Face30yonsei-dli /SAGEO-Arena SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization SAGEO Arena is a benchmark for evaluating Search-Augmented Generative Engine Optimization (SAGEO) — the practice of optimizing web documents to improve their visibility in AI-generated responses. Contents This dataset releases the queries and Google Custom Search API results used to construct the SAGEO Arena corpus. Please follow the crawler instructions in the GitHub… See the full description on the dataset page: https://huggingface.co/datasets/yonsei-dli/SAGEO-Arena.texttext-retrieval1K<n<10K0 likes38 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.