datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.Rainbow-Pony-100m-Flutter-direct-eval
Rainbow-Pony-100M Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct,
a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files.
In direct mode, the model is given an existing file and an edit instruction and
generates the complete modified file in a single forward pass (as opposed to the
steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.Qwen2.5-Coder-0.5B-Flutter-direct-eval
Qwen2.5-Coder-0.5B Flutter — Direct Mode — Validation Results
Dataset Summary
Held-out evaluation results for bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct,
a fine-tune of Qwen2.5-Coder-0.5B for editing Flutter/Dart source files. In direct
mode, the model is given an existing file and an edit instruction and generates the
complete modified file in a single forward pass (as opposed to the steps /
iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Qwen2.5-Coder-0.5B-Flutter-direct-eval.DOD-Directive-type-Memorandum-22-001-Records-Management-Standards
DoD Records Management Standards for IT Systems and Services
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive-type Memorandum 22-001, “DoD Standards for Records Management Capabilities in Programs Including Information Technology,” dated March 3, 2022, and incorporating Change 2 effective February 22, 2024.
The source establishes… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-type-Memorandum-22-001-Records-Management-Standards.DOD-Directive-8000-01-Management-Of-Defense-Information
DoD Directive 8000.01 Management of the DoD Information Enterprise Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on Department of Defense Directive 8000.01, “Management of the Department of Defense Information Enterprise,” dated March 17, 2016, and incorporating Change 1 effective July 27, 2017.
The directive establishes Department-wide… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Directive-8000-01-Management-Of-Defense-Information.open-library-10k
Dir Bear Open Library — 10,000 distilled web documents
Ten thousand complete, cleaned, English web documents — every one at least 250
words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored
by the quality of the site it came from. This is the free, open slice of the
Dir Bear corpus: the same records, the same schema and the
same pipeline as the datasets we sell, at a size you can read through in an afternoon.
Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.gsm8k-direct
📦 GSM8K-Direct
GSM8K-Direct is a streamlined version of the original GSM8K dataset, containing only math word problems paired directly with numeric answers.
Each entry follows a simple format:
{
"question": "Tom has 3 apples and buys 2 more. How many apples does he have now?",
"answer": 5
}
This dataset simplifies AI training by removing complex explanatory text, making model training faster, lighter, and easier—ideal for situations where direct numeric responses are sufficient.wikipedia-paragraphs-direct-paraphrases
Wikipedia Paragraphs Direct Paraphrases
Paraphrases of Wikipedia paragraphs using AI large language models.
Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits
Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt:
Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.
