CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01naderalfares /epoch_ai_swebench_verified Epoch AI SWE-bench Verified Traces Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations. Contents 34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run). data/: loadable Parquet data, one exact trace per row. original/: the byte-identical .eval archives published by Epoch AI. run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.tabulartext-generation10K<n<100K1 likes4.2k downloads1mo agoHugging Face02ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes1k downloads3mo agoHugging Face03naveenreddie-18 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.tabulartext-generation10M<n<100M0 likes874 downloads6mo agoHugging Face04nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes698 downloads2y agoHugging Face05nakasyou /japanese-conversion Awesome Japanese IME Training Data Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の ランキング学習例を作成したデータセットです。中間の読み付きデータセットは 作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した 例だけを採用できます。 context: 変換対象より前の本文 input: 変換対象のひらがな読み correct: 元コーパスにある正解表記 incorrect: predict.py で全体または一部分を再変換した誤候補の配列 n_words: 抽出した連続形態素数 source_text と target_start / target_end により、元文章中の抽出位置を 復元できます。元データの利用条件は from と from_license を参照して ください。 tabulartext-generation10M<n<100M1 likes506 downloads1mo agoHugging Face06nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes447 downloads8d agoHugging Face07nazimali /quran Dataset Card for the Quran Summary The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages. TODO Add Tafsirs Add topics/ontology Usage from datasets import load_dataset ds = load_dataset("nazimali/quran", split="train") ds Output: Dataset({ features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran.tabulartext-classification1K<n<10K20 likes408 downloads2y agoHugging Face08NAIL-Group /ClawBench ClawBench — A Benchmark for AI Web Agents Can AI Agents Complete Everyday Online Tasks? |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.tabulartext-generationn<1K2 likes363 downloads4mo agoHugging Face09namkoong-lab /PersonalLLM Dataset Card for PersonalLLM The PersonalLLM dataset is a collection of prompts, responses, and rewards designed for personalized language model methodology development and evaluation. This dataset is presented in the paper PersonalLLM: Tailoring LLMs to Individual Preferences. Dataset Details Dataset Description Curated by: Andrew Siah*, Tom Zollo*, Naimeng Ye, Ang Li, Namkoong Hongseok Funded by: Digital Future Initiative at Columbia Business School… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM.tabulartext-generation10K<n<100K18 likes195 downloads2y agoHugging Face10natong19 /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.tabularquestion-answering1K<n<10K0 likes182 downloads9mo agoHugging Face11naist-nlp /XQ-MEval Dataset Card for XQ-MEval XQ-MEval is a quality-parallel benchmark dataset for automatic evaluation metrics on cross-lingual scoring bias. Dataset Details Dataset Description XQ-MEval is a benchmark released under CC BY-S 4.0 for evaluating automatic metrics with respect to cross-lingual scoring bias. This dataset is constructed by injecting varying numbers of Multidimensional Quality Metric (MQM)-defined errors into high-quality translations… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/XQ-MEval.tabulartext-generation10K<n<100K2 likes163 downloads3mo agoHugging Face12nanskong /ManipuriGPT-Corpus-v1.0 ManipuriGPT Corpus v1.0 ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models. Quick Summary Total Sequences: 147,956 Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075 Total Characters: 16,019,401 Pipeline Version: 5.6 Release Version: v1.0.0 Build Timestamp: 2026-07-25T09:17:21.960438Z Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.tabulartext-generation100K<n<1M0 likes139 downloads2mo agoHugging Face13saheedniyi /naijaweb Naijaweb Dataset 🇳🇬 Naijaweb is a dataset that contains over 270,000+ documents, totaling approximately 230 million GPT-2 tokens. The data was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts. Dataset Summary Features Data Types text string link string token_count int64 section string int_score int64 language string language_probability float64… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb.tabulartext-generation100K<n<1M39 likes121 downloads2y agoHugging Face14Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes109 downloads10mo agoHugging Face15Lumia101 /Nari-C4-ko-500MT Lumia101/Nari-C4-ko-500MT This dataset is a modified version of C4 dataset(multilingual, ko subset) made more useful for LLM training by applying additional filtering. Since the number of tokens is only about 500M, it is recommended to mix it with other high-quality datasets. Additional filtering methods used Phase 1: Text normalization Phase 2: Remove HTML-filled junk documents Phase 3: Remove documents containing a lot of broken characters Phase 4: Remove documents… See the full description on the dataset page: https://huggingface.co/datasets/Lumia101/Nari-C4-ko-500MT.tabulartext-generation100K<n<1M0 likes108 downloads6mo agoHugging Face16NagaYu /isotope-bench Isotope Bench An indirect-prompt-injection benchmark for tool-calling agents, plus the complete audit trail of one recorded run: 438 influence certificates, one for every action an agent attempted across five defence conditions. Built for Isotope, which tracks untrusted influence inside the forward pass. The corpus is independent of that method and usable with any defence. 💻 Code: https://github.com/NagaYu/isotope 🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope 🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.tabulartext-generationn<1K1 likes106 downloads18d agoHugging Face17nampdn-ai /tiny-code-textbooksgated Code Explanation Textbooks A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook. tabulartext-generation100K<n<1M13 likes104 downloads3y agoHugging Face18Naholav /CodeGen-Deep-5K CodeGen-Deep-5K: Deep Reasoning for Competitive Programming Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset) Dataset Description CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches. Key Statistics Total samples: 5,000 Unique… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.tabulartext-generation1K<n<10K0 likes102 downloads10mo agoHugging Face19rs545837 /entity-native-agent-sessions Entity-Native vs File-Native Agent Sessions on SWE-bench Verified Full session logs from a controlled A/B experiment measuring how a coding agent's retrieval substrate changes its behaviour, cost, and success rate on real software-engineering tasks. Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the same repository state. The only difference is how the agent is allowed to find code. Arm Label Tools available A file-native Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.tabulartext-generationn<1K0 likes102 downloads22d agoHugging Face20nassimjp /quran-tafsir QuranLab — Multilingual Quran Tafsir Dataset A ready-to-use collection of Quran commentaries and annotated translations, aligned to the canonical 6,236 ayahs. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours. The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.tabulartext-generation100K<n<1M0 likes102 downloads5d agoHugging Face21NathanielArfin /canadian-hansard-44-1 Canadian Hansard — House of Commons Debates, 44th Parliament 1st Session Speech-level dataset of the Debates of the Canadian House of Commons (Hansard), 44th Parliament, 1st session (November 2021 – January 2025), in English and French — both official languages, same 391 sitting days. Source & terms Fetched from the official per-sitting XML published by the House of Commons of Canada: ourcommons.ca/Content/House/…/HAN{n}-{E|F}.XML The text is Crown copyright… See the full description on the dataset page: https://huggingface.co/datasets/NathanielArfin/canadian-hansard-44-1.tabulartext-generation100K<n<1M1 likes97 downloads16d agoHugging Face22Naveen934 /tamil_data_kalki_Fiction Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி Description This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining. Contents 2290 text chunks Author: கல்கி ரா. கிருஷ்ணமூர்த்தி Genre: பொன்னியின் செல்வன் Total chunks: 2290 Usage from datasets import load_dataset dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")``` tabulartext-generation1K<n<10K0 likes89 downloads1y agoHugging Face23nassimjp /quran QuranLab — Verse-Aligned Multilingual Quran Corpus A unified, verse-aligned multilingual Quran corpus spanning 79 languages and 185 translations. Every recension and translation is a separate config (subset), all row-aligned on the canonical 6,236-ayah verse_key (Hafs ʿan ʿAsim reading, 114 surahs). The corpus also contains 111 tafsir configs: verse-grain classical and openly licensed Arabic works, plus the native-passage and verse-expanded views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.tabulartext-generation1M<n<10M0 likes83 downloads5d agoHugging Face24Nan-Do /atcoder_abc_contestsgated Notification Atcoder is selling this data now. If you are interested in accessing it please contact them. Dataset Summary This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques. It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests.tabulartext-classification10M<n<100M2 likes77 downloads2mo agoHugging Face25MuhammadHelmy /nafsy Dataset Card for nafsy This arabic dataset is a set of mental health articles. The original dataset was scrapped from Nafsy.net. Dataset Details Language(s) (NLP): Arabic Uses Direct Use Unsupervised Fine-tuning RAG Dataset Structure Dataset Fields: content: the articles text_size: length of article topic: top 10 words that describe the topics of the article prob: topic prediction accuracy Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/MuhammadHelmy/nafsy.tabulartext-generation1K<n<10K0 likes69 downloads3y agoHugging Face26Navneeth017 /GlotCC-V1_maltabulartext-generation100K<n<1M0 likes68 downloads8mo agoHugging Face27n-alignment /USACO-Judge USACO-Judge A benchmark for judging competitive-programming solutions: given a problem and a candidate solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code. Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.tabulartext-generationn<1K0 likes65 downloads1mo agoHugging Face28nahsa /in-the-wild-jailbreak-prompts In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1… See the full description on the dataset page: https://huggingface.co/datasets/nahsa/in-the-wild-jailbreak-prompts.tabulartext-generation10K<n<100K0 likes60 downloads8d agoHugging Face29thoughtworks /cbd-100pair-hf-natural-rewrite-provenance cbd-100pair-hf-natural-rewrite-provenance Provenance-complete regeneration of 500 hf_natural-style conjunctive-pair rewrites: five rows for each of the 100 frozen AND-trigger pairs. This is a new deterministic sample, not a recovery of the historical 14,961 rows. The old published data removed source_id, so its exact OpenOrca-to-rewrite mapping cannot be reconstructed. Each source is an indexed row from Open-Orca/OpenOrca. Rewrites were generated with the same model and prompt… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-hf-natural-rewrite-provenance.tabulartext-generationn<1K0 likes54 downloads6d agoHugging Face30NaverHustQA /TVPL TVPL (thuvienphapluat.vn) structured_data_doc.parquet: preprocessed version from only needed doc from tvpl . Please read by Datasets library parent_nodes.parquet: parent nodes from [1] by chunking with SentenceSplitter, chunk_overlap=0, chunk_size=800, tokenizer="Viet-Mistral/Vistral-7B-Chat" child_nodes.parquet: child nodes from [2] by chunking with SentenceSplitter, chunk_overlap=30, chunk_size=190 and using Word Segmentation, vietnamese-bi-encoder Dedup WARNING:… See the full description on the dataset page: https://huggingface.co/datasets/NaverHustQA/TVPL.tabulartext-generation100K<n<1M0 likes51 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.