datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epoch_ai_swebench_verified
Epoch AI SWE-bench Verified Traces
Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations.
Contents
34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run).
data/: loadable Parquet data, one exact trace per row.
original/: the byte-identical .eval archives published by Epoch AI.
run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.japanese-conversion
Awesome Japanese IME Training Data
Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の
ランキング学習例を作成したデータセットです。中間の読み付きデータセットは
作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した
例だけを採用できます。
context: 変換対象より前の本文
input: 変換対象のひらがな読み
correct: 元コーパスにある正解表記
incorrect: predict.py で全体または一部分を再変換した誤候補の配列
n_words: 抽出した連続形態素数
source_text と target_start / target_end により、元文章中の抽出位置を
復元できます。元データの利用条件は from と from_license を参照して
ください。
local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.quran
Dataset Card for the Quran
Summary
The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages.
TODO
Add Tafsirs
Add topics/ontology
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/quran", split="train")
ds
Output:
Dataset({
features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran.ClawBench
ClawBench — A Benchmark for AI Web Agents
Can AI Agents Complete Everyday Online Tasks?
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.PersonalLLM
Dataset Card for PersonalLLM
The PersonalLLM dataset is a collection of prompts, responses, and rewards designed for personalized language model methodology development and evaluation. This dataset is presented in the paper PersonalLLM: Tailoring LLMs to Individual Preferences.
Dataset Details
Dataset Description
Curated by: Andrew Siah*, Tom Zollo*, Naimeng Ye, Ang Li, Namkoong Hongseok
Funded by: Digital Future Initiative at Columbia Business School… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.XQ-MEval
Dataset Card for XQ-MEval
XQ-MEval is a quality-parallel benchmark dataset for automatic evaluation metrics on cross-lingual scoring bias.
Dataset Details
Dataset Description
XQ-MEval is a benchmark released under CC BY-S 4.0 for evaluating automatic metrics with respect to cross-lingual scoring bias. This dataset is constructed by injecting varying numbers of Multidimensional Quality Metric (MQM)-defined errors into high-quality translations… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/XQ-MEval.ManipuriGPT-Corpus-v1.0
ManipuriGPT Corpus v1.0
ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models.
Quick Summary
Total Sequences: 147,956
Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075
Total Characters: 16,019,401
Pipeline Version: 5.6
Release Version: v1.0.0
Build Timestamp: 2026-07-25T09:17:21.960438Z
Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.naijaweb
Naijaweb Dataset 🇳🇬
Naijaweb is a dataset that contains over 270,000+ documents, totaling approximately 230 million GPT-2 tokens. The data was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string
language_probability
float64… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb.CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.Nari-C4-ko-500MT
Lumia101/Nari-C4-ko-500MT
This dataset is a modified version of C4 dataset(multilingual, ko subset) made more useful for LLM training by applying additional filtering.
Since the number of tokens is only about 500M, it is recommended to mix it with other high-quality datasets.
Additional filtering methods used
Phase 1: Text normalization
Phase 2: Remove HTML-filled junk documents
Phase 3: Remove documents containing a lot of broken characters
Phase 4: Remove documents… See the full description on the dataset page: https://huggingface.co/datasets/Lumia101/Nari-C4-ko-500MT.isotope-bench
Isotope Bench
An indirect-prompt-injection benchmark for tool-calling agents, plus the
complete audit trail of one recorded run: 438 influence certificates, one for
every action an agent attempted across five defence conditions.
Built for Isotope, which tracks
untrusted influence inside the forward pass. The corpus is independent of that
method and usable with any defence.
💻 Code: https://github.com/NagaYu/isotope
🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope
🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.tiny-code-textbooks
Code Explanation Textbooks
A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook.
CodeGen-Deep-5K
CodeGen-Deep-5K: Deep Reasoning for Competitive Programming
Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset)
Dataset Description
CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches.
Key Statistics
Total samples: 5,000
Unique… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.entity-native-agent-sessions
Entity-Native vs File-Native Agent Sessions on SWE-bench Verified
Full session logs from a controlled A/B experiment measuring how a coding agent's
retrieval substrate changes its behaviour, cost, and success rate on real
software-engineering tasks.
Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the
same repository state. The only difference is how the agent is allowed to find code.
Arm
Label
Tools available
A
file-native
Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.canadian-hansard-44-1
Canadian Hansard — House of Commons Debates, 44th Parliament 1st Session
Speech-level dataset of the Debates of the Canadian House of Commons
(Hansard), 44th Parliament, 1st session (November 2021 – January 2025),
in English and French — both official languages, same 391 sitting days.
Source & terms
Fetched from the official per-sitting XML published by the House of
Commons of Canada: ourcommons.ca/Content/House/…/HAN{n}-{E|F}.XML
The text is Crown copyright… See the full description on the dataset page: https://huggingface.co/datasets/NathanielArfin/canadian-hansard-44-1.tamil_data_kalki_Fiction
Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி
Description
This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining.
Contents
2290 text chunks
Author: கல்கி ரா. கிருஷ்ணமூர்த்தி
Genre: பொன்னியின் செல்வன்
Total chunks: 2290
Usage
from datasets import load_dataset
dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")```
quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.atcoder_abc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests.nafsy
Dataset Card for nafsy
This arabic dataset is a set of mental health articles. The original dataset was scrapped from Nafsy.net.
Dataset Details
Language(s) (NLP): Arabic
Uses
Direct Use
Unsupervised Fine-tuning
RAG
Dataset Structure
Dataset Fields:
content: the articles
text_size: length of article
topic: top 10 words that describe the topics of the article
prob: topic prediction accuracy
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/MuhammadHelmy/nafsy.GlotCC-V1_malUSACO-Judge
USACO-Judge
A benchmark for judging competitive-programming solutions: given a problem and a candidate
solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code.
Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test
the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests
whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1… See the full description on the dataset page: https://huggingface.co/datasets/nahsa/in-the-wild-jailbreak-prompts.cbd-100pair-hf-natural-rewrite-provenance
cbd-100pair-hf-natural-rewrite-provenance
Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for
each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the
historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite
mapping cannot be reconstructed.
Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model
and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.TVPL
TVPL (thuvienphapluat.vn)
structured_data_doc.parquet: preprocessed version from only needed doc from tvpl . Please read by Datasets library
parent_nodes.parquet: parent nodes from [1] by chunking with SentenceSplitter, chunk_overlap=0, chunk_size=800, tokenizer="Viet-Mistral/Vistral-7B-Chat"
child_nodes.parquet: child nodes from [2] by chunking with SentenceSplitter, chunk_overlap=30, chunk_size=190 and using Word Segmentation, vietnamese-bi-encoder
Dedup
WARNING:… See the full description on the dataset page: https://huggingface.co/datasets/NaverHustQA/TVPL.
