CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes13k downloads11mo agoHugging Face02mohameddalii /coda-llm-data Coda LLM Project & Dataset Repository This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent). Model Repository: mohameddalii/coda-llm Dataset / Code Repository: mohameddalii/coda-llm-data 📁 Repository Structure coda-llm-data/ ├── data/ │ ├── raw/ # Raw generated multi-turn dialogues across domains │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.texttext-generation1K<n<10K0 likes9.9k downloads2d agoHugging Face03bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face04tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B48 likes8.6k downloads11mo agoHugging Face05Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes7.1k downloads3y agoHugging Face06fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes3.2k downloads3mo agoHugging Face07epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes3k downloads3y agoHugging Face08LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.7k downloads6mo agoHugging Face09Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K0 likes2.4k downloads3mo agoHugging Face10llm-wizard /alpaca-gpt4-data-zh Dataset Card for "alpaca-gpt4-data-zh" All of the work is done by this team. Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes. English Dataset Found here Citation @article{peng2023gpt4llm, title={Instruction Tuning with GPT-4}, author={Baolin Peng, Chunyuan Li… See the full description on the dataset page: https://huggingface.co/datasets/llm-wizard/alpaca-gpt4-data-zh.texttext-generation10K<n<100K73 likes1.9k downloads3y agoHugging Face11mii-llm /gazzetta-ufficiale Gazzetta Ufficiale 👩🏻‍⚖️⚖️🏛️📜🇮🇹 La Gazzetta Ufficiale della Repubblica Italiana, quale fonte ufficiale di conoscenza delle norme in vigore in Italia e strumento di diffusione, informazione e ufficializzazione di testi legislativi, atti pubblici e privati, è edita dall’Istituto Poligrafico e Zecca dello Stato e pubblicata in collaborazione con il Ministero della Giustizia, il quale provvede alla direzione e redazione della stessa. L'Istituto Poligrafico e Zecca dello Stato… See the full description on the dataset page: https://huggingface.co/datasets/mii-llm/gazzetta-ufficiale.texttext-generation1M<n<10M39 likes1.5k downloads3y agoHugging Face12LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face13tokyotech-llm /swallow-math SwallowMath October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines. Resources 🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math. 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation. What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.texttext-generation1M<n<10M49 likes994 downloads7mo agoHugging Face14Exgentic /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.tabulartext-generation1K<n<10K23 likes946 downloads4mo agoHugging Face15tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B69 likes896 downloads7mo agoHugging Face16Podtech /llm-jp-corpus-v4-ja_wiki llm-jp-corpus-v4 — ja_wiki Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_wiki Files: 6 × jsonl.gz (1.9 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.texttext-generation1M<n<10M0 likes879 downloads2mo agoHugging Face17zenml /llmops-database The ZenML LLMOps Database To learn more about ZenML and our open-source MLOps framework, visit zenml.io. Dataset Summary The LLMOps Database is a comprehensive collection of over 500 real-world generative AI implementations that showcases how organizations are successfully deploying Large Language Models (LLMs) in production. The case studies have been carefully curated to focus on technical depth and practical problem-solving, with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.textfeature-extraction1K<n<10K23 likes864 downloads7d agoHugging Face18Podtech /llm-jp-corpus-v4-ja_warp_pdf llm-jp-corpus-v4 — ja_warp_pdf Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_pdf Files: 513 × jsonl.gz (73.8 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.texttext-generation10M<n<100M0 likes808 downloads2mo agoHugging Face19llm-jp /scaling-data-constrained-llms Scaling Data-Constrained Language Models with Synthetic Data This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026). Overview This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting. Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.texttext-generation100M<n<1B5 likes796 downloads6mo agoHugging Face20tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes772 downloads7mo agoHugging Face21Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 61 runs, 2,689,200 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes757 downloads21h agoHugging Face22bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes742 downloads2y agoHugging Face23Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes710 downloads1y agoHugging Face24Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_html llm-jp-corpus-v4 — ja_sip_comprehensive_html Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_html Files: 181 × jsonl.gz (23.4 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.texttext-generation1M<n<10M0 likes645 downloads2mo agoHugging Face25orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes615 downloads2y agoHugging Face26neurondb /postgresql-llm postgresql-llm A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis. Dataset Summary postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source. Metric Value Total rows 211,539 PostgreSQL-specific rows 11,998 (5.7%) Schema fill rate 82.2% Explanation fill rate 17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.tabulartext-generation100K<n<1M4 likes602 downloads7mo agoHugging Face27lapa-llm /hermes3-uk Dataset Card for Hermes 3 Ukrainian Fixed Conversations Dataset Description Dataset Summary hermes3-uk-fixed is a Ukrainian translation of the [NousResearch/Hermes-3-Dataset]. The translation was produced with Gemma 3 27B (instruction-tuned). During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets. Languages Ukrainian (uk) Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-uk.texttext-generation100K<n<1M0 likes586 downloads11mo agoHugging Face28SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face29Psychotherapy-LLM /CBT-Bench CBT-Bench Dataset Overview CBT-Bench is a benchmark dataset designed to evaluate the proficiency of Large Language Models (LLMs) in assisting cognitive behavior therapy (CBT). The dataset is organized into three levels, each focusing on different key aspects of CBT, including basic knowledge recitation, cognitive model understanding, and therapeutic response generation. The goal is to assess how well LLMs can support various stages of professional mental health care… See the full description on the dataset page: https://huggingface.co/datasets/Psychotherapy-LLM/CBT-Bench.textquestion-answering1K<n<10K43 likes513 downloads1y agoHugging Face30Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes495 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.