CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01krisbailey /cosmopedia-10B Cosmopedia 10B Dataset Description This is a 10.53 Billion token subset of the HuggingFaceTB/cosmopedia dataset. It was created by sampling approximately 45% of each subset (web_samples, stories, stanford, etc.) from the original dataset and deduplicating to ensure high utility. Motivation The original Cosmopedia dataset is massive (~25B+ tokens) and high quality. This 10B version serves as a "Goldilocks" dataset—large enough for meaningful pre-training… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-10B.texttext-generation10M<n<100M0 likes246 downloads8mo agoHugging Face02krisbailey /RedPajama-Data-V2-1B RedPajama-Data-V2 1B Dataset Description This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data. Motivation RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.texttext-generation100K<n<1M0 likes221 downloads8mo agoHugging Face03krisbailey /cosmopedia-1b Cosmopedia 1B Dataset Description This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia. It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source. Motivation While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.texttext-generation1M<n<10M0 likes194 downloads8mo agoHugging Face04krist67 /wikipedia Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/krist67/wikipedia.texttext-generation10M<n<100M0 likes136 downloads3mo agoHugging Face05hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes136 downloads4d agoHugging Face06krisbailey /RedPajama-10B-Weighted RedPajama-10B-Weighted A canonical 10 Billion token weighted subset of the RedPajama-Data-1T dataset. Dataset Description This dataset is a faithful reproduction of the original RedPajama-Data-1T distribution, scaled down to exactly 10 Billion tokens. It is designed to preserve the exact domain ratios of the original dataset (excluding the defunct 'Books' subset). This allows researchers and developers to prototype, debug, and test on a representative slice of the data… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-10B-Weighted.texttext-generation1M<n<10M0 likes99 downloads9mo agoHugging Face07krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes94 downloads6mo agoHugging Face08uralstech /kcc-krishi-rag-sft-advisory-corpus KCC-Krishi RAG/SFT Advisory Corpus The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem. It was created for: agricultural RAG research; supervised fine-tuning research; evidence-grounded response generation; safety-routing experiments; offline farmer-assistant prototyping; reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.texttext-generation100K<n<1M0 likes79 downloads2mo agoHugging Face09kristaller486 /wikisource_preferences_ru Wikisource Preferences [Russian] Датасет для оптимизации предпочтений. chosen тексты брались из kristaller486/wikisource-creative-ru, а rejected генерировались разнообразными LLM по сгенерированным промптам. Шаблон для DPO: axolotl chat_template.default Модели для генерации rejected семплов: google/gemma-3-27b-it gpt-4.1-mini gpt-4.1-nano gpt-4.1 gemini-2.0-flash Qwen/Qwen3-14B-FP8 (without reasoning) Moraliane/SAINEMO-reMIX (fp6-llm quantization) deepseek-v3-0324 (api)… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/wikisource_preferences_ru.texttext-generation10K<n<100K0 likes74 downloads1y agoHugging Face10michsethowusu /Code-170k-krio Dataset Description Code-170k-krio is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Krio, making coding education accessible to Krio speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Krio language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-krio.texttext-generation100K<n<1M0 likes63 downloads11mo agoHugging Face11hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes63 downloads9d agoHugging Face12kristaller486 /Nebo-T1-Russian Russian Description (English below) UPD: Dataset reuploaded, correct_format column added Nebo-T1-Russian (Вероятно) первый "longCoT" датасет для русского языка, созданный через Deeseek-R1 Подсказки взяты из датасета Sky-T1 и переведены через Llama3.3-70B Ответы и рассуждения сгенерированные Deeseek-R1 (685B) 16.4K сэмплов в целом, ≈12.4K только с русским языком (в остальных либо ответ, либо рассуждения на английском) Языки в ответе и рассуждениях размечены… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/Nebo-T1-Russian.texttext-generation10K<n<100K15 likes59 downloads2y agoHugging Face13krisbailey /fineweb-edu-1B FineWeb-Edu 1B Dataset Description FineWeb-Edu 1B is a high-quality, stratified subset of the HuggingFaceFW/fineweb-edu dataset. It contains approximately 1 billion tokens of educational web text, carefully sampled to preserve the original distribution of source data (CommonCrawl dumps). This dataset provides an accessible, lightweight alternative to the larger FineWeb-Edu subsets (like sample-10BT or sample-100BT) while maintaining the same data diversity and quality… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/fineweb-edu-1B.tabulartext-generation100K<n<1M0 likes59 downloads8mo agoHugging Face14krisdcosta /edge-llm-bench Edge LLM Bench — GGUF Quantization Benchmarks on Edge Devices Controlled inference benchmark dataset for 7 GGUF K-quant quantization variants (Q2_K through Q8_0) of Llama 3.2 3B Instruct and Qwen 2.5 1.5B Instruct across three hardware platforms: Device SoC / CPU RAM Backend Google Pixel 6a Google Tensor G1 (ARM Cortex-X1) 6 GB LPDDR5 llama.cpp CPU Apple M4 Mac Apple M4 (ARM, 10-core) 16 GB unified llama.cpp Metal HP Pavilion x86 Intel Core i5-1235U (12th gen) 16 GB… See the full description on the dataset page: https://huggingface.co/datasets/krisdcosta/edge-llm-bench.tabulartext-generation1K<n<10K0 likes59 downloads5mo agoHugging Face15hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes54 downloads9d agoHugging Face16krisbailey /RedPajama-Data-V2-100M RedPajama-Data-V2-100M Dataset Description This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.texttext-generation10K<n<100K0 likes44 downloads8mo agoHugging Face17kristaller486 /writingprompts-ru Переведенный датасет euclaise/writingprompts Модель переводчик - Gemma-3-27b-it-bf16 Переведены только подсказки (пока) Translated euclaise/writingprompts dataset Translator - Gemma-3-27b-it-bf16 Translated only prompts (yet) texttext-generation100K<n<1M0 likes43 downloads2y agoHugging Face18kristaller486 /hermes-3-dataset-ru-translated-prompts Переведенные промты из hermes-3-dataset Модель-переводчик Gemma-3-27b-it. Переведены все промты. Multi-turn промты переведены с учетом контекста англоязычного ответа. Будет полезно для создания крупных русскоязычных инструктивных датасетов или Online RL. Translated prompts from hermes-3-dataset Translator model: Gemma-3-27b-it. All prompts have been translated. Multi-turn prompts were translated considering the context of the English response. This will be useful… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/hermes-3-dataset-ru-translated-prompts.texttext-generation100K<n<1M2 likes43 downloads1y agoHugging Face19krishy-d /formatbench FormatBench: A Preference Dataset for Correcting LLM Formatting Bias Part of the Prosify project. Also available on Kaggle. The problem this dataset addresses Large language models trained with RLHF systematically over-format their outputs are defaulting to bullet points, bold headers, and templated structures even when flowing prose would serve the reader better. This shows up most visibly when people use LLMs for real-world tasks: a request to "polish this… See the full description on the dataset page: https://huggingface.co/datasets/krishy-d/formatbench.texttext-generationn<1K0 likes41 downloads4mo agoHugging Face20krisbailey /falcon-refinedweb-1B Falcon RefinedWeb 1B Dataset Description This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data. Motivation RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.texttext-generation1M<n<10M0 likes39 downloads8mo agoHugging Face21kristaller486 /wikisource-creative-ru wikisource-creative-ru Russian description Русскоязычная часть wikimedia/wikisource, отфильтрованная по по критерию креативных текстов (через регулярные выражения) и выделен небольшой обособленный по смыслу фрагмент текста, как в dostoevsky или gutenberg-dpo. Модель, которая выделяла сегменты - Gemma-3-27b-it. English description The Russian-language portion of wikimedia/wikisource, filtered by the criterion of creative texts (using regular expressions)… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/wikisource-creative-ru.texttext-generation10K<n<100K0 likes34 downloads1y agoHugging Face22krisbailey /RedPajama-1B-Weighted RedPajama-1B-Weighted A canonical 1 Billion token weighted subset of the RedPajama-Data-1T dataset. Dataset Description This is a strict, downsampled version of the RedPajama-10B-Weighted dataset. It maintains the exact domain distributions of the full 1T dataset, resized to a lightweight 1 Billion token footprint. This dataset is ideal for: Rapid Prototyping: Train small models or debug pipelines in minutes rather than days. Reference Baselines: Use a standard… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-1B-Weighted.texttext-generation100K<n<1M0 likes34 downloads9mo agoHugging Face23krisbailey /cosmopedia-100M cosmopedia-100M Dataset Description This is a 100.0 Million token subset of krisbailey/cosmopedia-1B, which is a subset of HuggingFaceTB/cosmopedia. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details Total Tokens: 100,000,060… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-100M.texttext-generation100K<n<1M0 likes33 downloads8mo agoHugging Face24Krishnapadala55 /Brahmastra-DarkNetra Brahmastra DarkNetra The Dark Eye that sees every vulnerability in the shadows. Security Research Dataset Notice: This dataset contains cybersecurity training data including descriptions of vulnerability exploitation techniques, security testing payloads, and attack methodologies for educational and defensive purposes. Antivirus software may flag files due to pattern matching on security-related text. This is expected behavior for cybersecurity datasets and the files do NOT contain… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/Brahmastra-DarkNetra.texttext-generation100K<n<1M1 likes33 downloads6mo agoHugging Face25AROY76 /KrishiGyan KrishiGyan (কৃষিজ্ঞান) A Bengali question–answer dataset for Bangladeshi agriculture, with an explicit chain-of-thought reasoning trace on every entry. Bangladesh has decades of agricultural research sitting in printed pamphlets, extension leaflets and encyclopedia entries. Very little of it is machine-readable, and almost none of it exists in a form a language model can be trained on. KrishiGyan is an attempt to close part of that gap: 5,529 Bengali Q&A pairs grounded in… See the full description on the dataset page: https://huggingface.co/datasets/AROY76/KrishiGyan.textquestion-answering1K<n<10K0 likes32 downloads2mo agoHugging Face26Krishnapadala55 /brahmastra-benchmark BRAHMASTRA Security LLM Benchmark Suite A 6-suite, 280-prompt benchmark for evaluating Large Language Models on Web Application Security Testing (DAST) tasks. This benchmark accompanies the release of BRAHMASTRA v0.3 and provides a reproducible methodology for measuring DAST-relevant capabilities of security-fine-tuned LLMs. Why this benchmark? Existing security LLM benchmarks (CyberSecEval, SecQA, HackBench) focus on penetration-testing scenarios or general security… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/brahmastra-benchmark.texttext-classificationn<1K0 likes21 downloads5mo agoHugging Face27kristaller486 /Ideya-preview-8k Креативные тексты от Gemma-3-27b-it на русском языке на основе kristaller486/writingprompts-ru actial_prompt - промт для генерации work in progress Creative writing text using Gemma-3-27b-it in Russian based on kristaller486/writingprompts-ru actial_prompt - generation prompt work in progress texttext-generation1K<n<10K0 likes20 downloads2y agoHugging Face28krinal /fifa_2022 Dataset Card for Dataset Name Dataset Summary Text corpus dataset (fifa world cup 2022) Additional Information Citation Information @misc{ enwiki:1154298520, author = "{Wikipedia contributors}", title = "2022 FIFA World Cup --- {Wikipedia}{,} The Free Encyclopedia", year = "2023", url = "https://en.wikipedia.org/w/index.php?title=2022_FIFA_World_Cup&oldid=1154298520" } textsummarizationn<1K0 likes19 downloads3y agoHugging Face29KrishnaGarg /afterlight-agent-trace Afterlight Agent Trace This dataset publishes a representative successful agent trace from Afterlight: The Last Signal. The trace records the responsibilities, validation boundaries, selected models, fallback state, and final structured result for one generated sector. Architecture nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 plans a route using only supplied, curated astrophysical concept IDs. openbmb/MiniCPM5-1B writes names, mission language, and a fictional… See the full description on the dataset page: https://huggingface.co/datasets/KrishnaGarg/afterlight-agent-trace.texttext-generationn<1K0 likes18 downloads3mo agoHugging Face30krist67 /ultrachat_200k Dataset Card for UltraChat 200k Dataset Description This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model. The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic: Selection of a subset of data for faster supervised fine tuning. Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: https://huggingface.co/datasets/krist67/ultrachat_200k.texttext-generation100K<n<1M0 likes14 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.