datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosmopedia-10B
Cosmopedia 10B
Dataset Description
This is a 10.53 Billion token subset of the HuggingFaceTB/cosmopedia dataset. It was created by sampling approximately 45% of each subset (web_samples, stories, stanford, etc.) from the original dataset and deduplicating to ensure high utility.
Motivation
The original Cosmopedia dataset is massive (~25B+ tokens) and high quality. This 10B version serves as a "Goldilocks" dataset—large enough for meaningful pre-training… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-10B.RedPajama-Data-V2-1B
RedPajama-Data-V2 1B
Dataset Description
This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data.
Motivation
RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.cosmopedia-1b
Cosmopedia 1B
Dataset Description
This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia.
It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source.
Motivation
While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.wikipedia
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/krist67/wikipedia.enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,087 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a
self-correction loop — and the benchmark is designed so that number cannot be
inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.RedPajama-10B-Weighted
RedPajama-10B-Weighted
A canonical 10 Billion token weighted subset of the RedPajama-Data-1T dataset.
Dataset Description
This dataset is a faithful reproduction of the original RedPajama-Data-1T distribution, scaled down to exactly 10 Billion tokens. It is designed to preserve the exact domain ratios of the original dataset (excluding the defunct 'Books' subset). This allows researchers and developers to prototype, debug, and test on a representative slice of the data… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-10B-Weighted.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
Punjabi
Gurmukhi
~408K
Urdu
Nastaliq
~374K
Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.kcc-krishi-rag-sft-advisory-corpus
KCC-Krishi RAG/SFT Advisory Corpus
The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem.
It was created for:
agricultural RAG research;
supervised fine-tuning research;
evidence-grounded response generation;
safety-routing experiments;
offline farmer-assistant prototyping;
reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.wikisource_preferences_ru
Wikisource Preferences [Russian]
Датасет для оптимизации предпочтений. chosen тексты брались из kristaller486/wikisource-creative-ru, а rejected генерировались разнообразными LLM по сгенерированным промптам.
Шаблон для DPO: axolotl chat_template.default
Модели для генерации rejected семплов:
google/gemma-3-27b-it
gpt-4.1-mini
gpt-4.1-nano
gpt-4.1
gemini-2.0-flash
Qwen/Qwen3-14B-FP8 (without reasoning)
Moraliane/SAINEMO-reMIX (fp6-llm quantization)
deepseek-v3-0324 (api)… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/wikisource_preferences_ru.Code-170k-krio
Dataset Description
Code-170k-krio is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Krio, making coding education accessible to Krio speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Krio language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-krio.text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.Nebo-T1-Russian
Russian Description (English below)
UPD: Dataset reuploaded, correct_format column added
Nebo-T1-Russian
(Вероятно) первый "longCoT" датасет для русского языка, созданный через Deeseek-R1
Подсказки взяты из датасета Sky-T1 и переведены через Llama3.3-70B
Ответы и рассуждения сгенерированные Deeseek-R1 (685B)
16.4K сэмплов в целом, ≈12.4K только с русским языком (в остальных либо ответ, либо рассуждения на английском)
Языки в ответе и рассуждениях размечены… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/Nebo-T1-Russian.fineweb-edu-1B
FineWeb-Edu 1B
Dataset Description
FineWeb-Edu 1B is a high-quality, stratified subset of the HuggingFaceFW/fineweb-edu dataset. It contains approximately 1 billion tokens of educational web text, carefully sampled to preserve the original distribution of source data (CommonCrawl dumps).
This dataset provides an accessible, lightweight alternative to the larger FineWeb-Edu subsets (like sample-10BT or sample-100BT) while maintaining the same data diversity and quality… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/fineweb-edu-1B.edge-llm-bench
Edge LLM Bench — GGUF Quantization Benchmarks on Edge Devices
Controlled inference benchmark dataset for 7 GGUF K-quant quantization variants
(Q2_K through Q8_0) of Llama 3.2 3B Instruct and Qwen 2.5 1.5B Instruct across three hardware platforms:
Device
SoC / CPU
RAM
Backend
Google Pixel 6a
Google Tensor G1 (ARM Cortex-X1)
6 GB LPDDR5
llama.cpp CPU
Apple M4 Mac
Apple M4 (ARM, 10-core)
16 GB unified
llama.cpp Metal
HP Pavilion x86
Intel Core i5-1235U (12th gen)
16 GB… See the full description on the dataset page: https://huggingface.co/datasets/krisdcosta/edge-llm-bench.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.RedPajama-Data-V2-100M
RedPajama-Data-V2-100M
Dataset Description
This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.writingprompts-ru
Переведенный датасет euclaise/writingprompts
Модель переводчик - Gemma-3-27b-it-bf16
Переведены только подсказки (пока)
Translated euclaise/writingprompts dataset
Translator - Gemma-3-27b-it-bf16
Translated only prompts (yet)
hermes-3-dataset-ru-translated-prompts
Переведенные промты из hermes-3-dataset
Модель-переводчик Gemma-3-27b-it.
Переведены все промты.
Multi-turn промты переведены с учетом контекста англоязычного ответа.
Будет полезно для создания крупных русскоязычных инструктивных датасетов или Online RL.
Translated prompts from hermes-3-dataset
Translator model: Gemma-3-27b-it.
All prompts have been translated.
Multi-turn prompts were translated considering the context of the English response.
This will be useful… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/hermes-3-dataset-ru-translated-prompts.formatbench
FormatBench: A Preference Dataset for Correcting LLM Formatting Bias
Part of the Prosify project.
Also available on Kaggle.
The problem this dataset addresses
Large language models trained with RLHF systematically over-format their outputs are defaulting to bullet points, bold headers, and templated structures
even when flowing prose would serve the reader better. This shows up most visibly when people use LLMs for real-world tasks: a request to "polish this… See the full description on the dataset page: https://huggingface.co/datasets/krishy-d/formatbench.falcon-refinedweb-1B
Falcon RefinedWeb 1B
Dataset Description
This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data.
Motivation
RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.wikisource-creative-ru
wikisource-creative-ru
Russian description
Русскоязычная часть wikimedia/wikisource, отфильтрованная по по критерию креативных текстов (через регулярные выражения)
и выделен небольшой обособленный по смыслу фрагмент текста, как в dostoevsky или gutenberg-dpo. Модель, которая выделяла сегменты - Gemma-3-27b-it.
English description
The Russian-language portion of wikimedia/wikisource, filtered by the criterion of creative texts (using regular expressions)… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/wikisource-creative-ru.RedPajama-1B-Weighted
RedPajama-1B-Weighted
A canonical 1 Billion token weighted subset of the RedPajama-Data-1T dataset.
Dataset Description
This is a strict, downsampled version of the RedPajama-10B-Weighted dataset. It maintains the exact domain distributions of the full 1T dataset, resized to a lightweight 1 Billion token footprint.
This dataset is ideal for:
Rapid Prototyping: Train small models or debug pipelines in minutes rather than days.
Reference Baselines: Use a standard… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-1B-Weighted.cosmopedia-100M
cosmopedia-100M
Dataset Description
This is a 100.0 Million token subset of krisbailey/cosmopedia-1B, which is a subset of HuggingFaceTB/cosmopedia.
Motivation
100M tokens is a standard size for:
CI/CD Pipelines: Fast enough to download and train for unit tests.
Debugging: Verifying training loops without waiting for hours.
Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).
Dataset Details
Total Tokens: 100,000,060… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-100M.Brahmastra-DarkNetra
Brahmastra DarkNetra
The Dark Eye that sees every vulnerability in the shadows.
Security Research Dataset Notice: This dataset contains cybersecurity training data
including descriptions of vulnerability exploitation techniques, security testing payloads,
and attack methodologies for educational and defensive purposes. Antivirus software may
flag files due to pattern matching on security-related text. This is expected behavior
for cybersecurity datasets and the files do NOT contain… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/Brahmastra-DarkNetra.KrishiGyan
KrishiGyan (কৃষিজ্ঞান)
A Bengali question–answer dataset for Bangladeshi agriculture, with an explicit chain-of-thought reasoning trace on every entry.
Bangladesh has decades of agricultural research sitting in printed pamphlets, extension leaflets and encyclopedia entries. Very little of it is machine-readable, and almost none of it exists in a form a language model can be trained on. KrishiGyan is an attempt to close part of that gap: 5,529 Bengali Q&A pairs grounded in… See the full description on the dataset page: https://huggingface.co/datasets/AROY76/KrishiGyan.brahmastra-benchmark
BRAHMASTRA Security LLM Benchmark Suite
A 6-suite, 280-prompt benchmark for evaluating Large Language Models on Web Application Security Testing (DAST) tasks.
This benchmark accompanies the release of BRAHMASTRA v0.3 and provides a reproducible methodology for measuring DAST-relevant capabilities of security-fine-tuned LLMs.
Why this benchmark?
Existing security LLM benchmarks (CyberSecEval, SecQA, HackBench) focus on penetration-testing scenarios or general security… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/brahmastra-benchmark.Ideya-preview-8k
Креативные тексты от Gemma-3-27b-it на русском языке
на основе kristaller486/writingprompts-ru
actial_prompt - промт для генерации
work in progress
Creative writing text using Gemma-3-27b-it in Russian
based on kristaller486/writingprompts-ru
actial_prompt - generation prompt
work in progress
fifa_2022
Dataset Card for Dataset Name
Dataset Summary
Text corpus dataset (fifa world cup 2022)
Additional Information
Citation Information
@misc{ enwiki:1154298520,
author = "{Wikipedia contributors}",
title = "2022 FIFA World Cup --- {Wikipedia}{,} The Free Encyclopedia",
year = "2023",
url = "https://en.wikipedia.org/w/index.php?title=2022_FIFA_World_Cup&oldid=1154298520"
}
afterlight-agent-trace
Afterlight Agent Trace
This dataset publishes a representative successful agent trace from
Afterlight: The Last Signal.
The trace records the responsibilities, validation boundaries, selected models,
fallback state, and final structured result for one generated sector.
Architecture
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 plans a route using only supplied,
curated astrophysical concept IDs.
openbmb/MiniCPM5-1B writes names, mission language, and a fictional… See the full description on the dataset page: https://huggingface.co/datasets/KrishnaGarg/afterlight-agent-trace.ultrachat_200k
Dataset Card for UltraChat 200k
Dataset Description
This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model.
The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic:
Selection of a subset of data for faster supervised fine tuning.
Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: https://huggingface.co/datasets/krist67/ultrachat_200k.
