CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Nan-Do /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.texttext-generation100K<n<1M30 likes4.3k downloads3y agoHugging Face02naderalfares /epoch_ai_swebench_verified Epoch AI SWE-bench Verified Traces Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations. Contents 34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run). data/: loadable Parquet data, one exact trace per row. original/: the byte-identical .eval archives published by Epoch AI. run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.tabulartext-generation10K<n<100K1 likes4.2k downloads1mo agoHugging Face03Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.7k downloads1mo agoHugging Face04facebook /natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.texttext-generation1M<n<10M585 likes2.7k downloads2y agoHugging Face05SLPL /naabHuge corpora of textual data are always known to be a crucial need for training deep models such as transformer-based ones. This issue is emerging more in lower resource languages - like Farsi. We propose naab, the biggest cleaned and ready-to-use open-source textual corpus in Farsi. It contains about 130GB of data, 250 million paragraphs, and 15 billion words. The project name is derived from the Farsi word ناب which means pure and high-grade.textfill-mask10M<n<100M45 likes2.4k downloads4y agoHugging Face06AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.9k downloads2y agoHugging Face07nampdn-ai /tiny-codesgated Reasoning with Language and Code This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.texttext-generation1M<n<10M302 likes1.6k downloads3y agoHugging Face08NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.4k downloads8mo agoHugging Face09Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.3k downloads1mo agoHugging Face10Nan-Do /instructional_code-search-net-python Dataset Card for "instructional_code-search-net-python" Dataset Summary This is an instructional dataset for Python. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.texttext-generation100K<n<1M36 likes963 downloads3y agoHugging Face11naveenreddie-18 /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/naveenreddie-18/hacker-news.tabulartext-generation10M<n<100M0 likes857 downloads6mo agoHugging Face12Yujivus /nanochat-climbmix-170 nanochat ClimbMix: first 170 train shards Convenience mirror of the exact initial ClimbMix slice downloaded by python -m nanochat.dataset -n 170. Contents Training: shard_00000.parquet through shard_00169.parquet Validation: shard_06542.parquet manifest.json: pinned source revision, file list, and byte sizes The Parquet shards are copied without modifying their rows or text. Attribution and provenance nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.texttext-generation10M<n<100M0 likes850 downloads1mo agoHugging Face13nakasyou /awesome-japanese-corpus Awesome Japanese Corpus 4つの日本語データソースを text、from、from_license の3列に正規化した Parquet データセットです。 FineWeb以外の3ソースは、生成処理を停止した時点までに取得済みの公開データを 採用しています。FineWebは最大3シャードを並列先読みする1時間限定の処理で サンプリングしています。空文字は除外しています。Infini-News は取得済みの 年度について language_iso639_3 == "jpn" の行を採用しています。 Sources hotchpotch/fineweb-2-edu-japanese (odc-by) ruggsea/infini-news-corpus の language_iso639_3 == "jpn" (cc-by-4.0) turing-motors/MOMIJI (cc-by-4.0) AhmedSSabir/Japanese-wiki-dump-sentence-dataset… See the full description on the dataset page: https://huggingface.co/datasets/nakasyou/awesome-japanese-corpus.texttext-generation100M<n<1B1 likes830 downloads2mo agoHugging Face14ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes820 downloads3mo agoHugging Face15NationalLibraryOfScotland /encyclopaedia-britannica-lance Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/encyclopaedia-britannica-lance.imageimage-to-text100K<n<1M2 likes781 downloads8mo agoHugging Face16dvilasuero /natural-science-reasoning Natural Sciences Reasoning: the "smolest" reasoning dataset A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes: Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.) Knowledge sharing for domains other than Math and Code reasoning In this repo, you can find: The prompts and the pipeline (see the config file). The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.texttext-generationn<1K40 likes707 downloads2y agoHugging Face17Archangel-system /glaive-function-calling-v2-openai-native glaive-function-calling-v2-openai-native glaiveai/glaive-function-calling-v2 restructured into the native OpenAI / TRL format: tools is a typed column and tool_calls[].function.arguments is a real object — not JSON inside a string. The original is widely used (69k downloads/month) but inactive for ~3 years, and ships tool calls as <functioncall> text blobs with Python-quoted arguments. Existing repackagings either keep ShareGPT with tools as a string, or carry no license at all.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/glaive-function-calling-v2-openai-native.texttext-generation10K<n<100K1 likes683 downloads11d agoHugging Face18naveenmarthala /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/naveenmarthala/arxiv-latex.texttext-generation1M<n<10M0 likes653 downloads3mo agoHugging Face19nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes650 downloads2y agoHugging Face20Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes612 downloads1mo agoHugging Face21nakasyou /japanese-conversion Awesome Japanese IME Training Data Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の ランキング学習例を作成したデータセットです。中間の読み付きデータセットは 作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した 例だけを採用できます。 context: 変換対象より前の本文 input: 変換対象のひらがな読み correct: 元コーパスにある正解表記 incorrect: predict.py で全体または一部分を再変換した誤候補の配列 n_words: 抽出した連続形態素数 source_text と target_start / target_end により、元文章中の抽出位置を 復元できます。元データの利用条件は from と from_license を参照して ください。 tabulartext-generation10M<n<100M1 likes547 downloads1mo agoHugging Face22Namronaldo2004 /ViInfographicsVQA Introduction ViInfographicsVQA is a Vietnamese Visual Question Answering (VQA) dataset constructed from infographics sourced from 26 different news platforms. The dataset is designed to support research in multimodal learning by providing diverse questions and answers based on real-world visual data. The detailed distribution of sources is presented in the table below. Figure 1: The number of infographics per news source. Developed by: @Namronaldo2004, @Kiet2302… See the full description on the dataset page: https://huggingface.co/datasets/Namronaldo2004/ViInfographicsVQA.imagequestion-answering100K<n<1M3 likes499 downloads1y agoHugging Face23nassimjp /pashto-emoji-dataset Pashto Emoji Dataset This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text. The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label. Dataset Structure The dataset is provided in the following format: text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.texttext-to-image100K<n<1M0 likes494 downloads26d agoHugging Face24nanoswe /nanoswe-trajs-260812 nanoswe SWE-agent trajectories (v0) A consolidation of the SWE-bench-style coding-agent trajectory corpora used to train the nanoswe speedrun models. Each row is one multi-turn agent trajectory (issue → tool-using rollout → patch), stored untokenized. 1,582,701 trajectories, 34 parquet shards, content-deduplicated on traj_hash. Seed corpus: ricdomolm/mini-coder-trajs-400k; the rest are derived SWE-smith / openhands / swe-zero conversions. Schema column… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/nanoswe-trajs-260812.texttext-generation1M<n<10M0 likes470 downloads1mo agoHugging Face25nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes448 downloads8d agoHugging Face26xywang1 /NaturalConv NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation Introduction This dataset is described in the paper NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation. The entire dataset contains 5 data files. 1. dialog_release.json: It is a json file containing a list of dictionaries. After loading in python this way: import json import codecs dialog_list = json.loads(codecs.open("dialog_release.json"… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/NaturalConv.texttext-generation10K<n<100K22 likes447 downloads2y agoHugging Face27asaverren /native-sft native-sft A format-alignment remix, not new instruction data. Conversations come from AllenAI Dolci (ODC-By) and NVIDIA Nemotron-Post-Training-Dataset-v1 (CC BY 4.0). Each family config re-renders those chats through a real 2026 instruct template so SFT can keep native special tokens / think / tools markers. Trainers get prompt + completion, so they do not need {% generation %} in jinja. v1 2026-08-31: ~9609 canonical conversations; 57 unique-hash family configs; 539,326… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/native-sft.texttext-generation100K<n<1M2 likes440 downloads24d agoHugging Face28tohoku-nlp /nanochat-jp-pretrain nanochat-jp-pretrain nanochat の日本語フォーク nanochat-jp で使用する 事前学習用日本語コーパス です. LLM によるクリーニングを施した日本語ウェブテキストと,llm-jp の公開コーパスを混合したものを,nanochat のデータローダがそのまま読める parquet 形式で配布しています. 構成 以下の4つのソースを混合し,全体をシャッフルしています. ソース llm-jp-corpus-v4 の ja_fineweb-2 サブセット(後述の追加データクリーニングを適用) llm-jp-corpus-midtraining-v2 ja/llm-jp-IPT_v0.3.2/ja_general.jsonl.gz llm-jp-corpus-midtraining-v2 ja/llm-jp-IPT_v0.3.2/ja_reasoning.jsonl.gz llm-jp/scaling-data-constrained-llms… See the full description on the dataset page: https://huggingface.co/datasets/tohoku-nlp/nanochat-jp-pretrain.texttext-generation10M<n<100M0 likes431 downloads1mo agoHugging Face29nazimali /quran Dataset Card for the Quran Summary The Quran with metadata, translations, and multiple Arabic text (can use specific types for embeddings, search, classification, and display). There are 126+ columns containing 43+ languages. TODO Add Tafsirs Add topics/ontology Usage from datasets import load_dataset ds = load_dataset("nazimali/quran", split="train") ds Output: Dataset({ features: ['surah', 'ayah', 'surah-name', 'surah-total-ayas'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran.tabulartext-classification1K<n<10K20 likes419 downloads2y agoHugging Face30nanoswe /swesmith-qwen3.6-35b-a3b SWE-smith trajectories from Qwen3.6-35B-A3B Multi-turn coding-agent trajectories (issue → tool-using rollout → patch) produced by Qwen3.6-35B-A3B on SWE-smith tasks, stored untokenized. This is the exact SFT corpus used for the harbor arm of the nanoswe teacher-distillation experiments. 101,901 trajectories over 45,242 unique SWE-smith task instances (3 sampled rollouts per task, ~2.25 surviving filtering), 53 parquet shards, ~1.4 GB. ≈1.96B training tokens = exactly one epoch… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/swesmith-qwen3.6-35b-a3b.texttext-generation100K<n<1M0 likes398 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.