datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.10k_recipes
CookBookAI EDA
This Exploratory Data Analysis (EDA) is related to the following Hugging Face Space:
CookBookAI Space
Data Overview:
The dataset consists of 10,000 synthetically generated recipes.
Full Analysis:
To view the full EDA process, you can visit the notebook directly:
EDA_AppLegacy.ipynb
1. Data Validation & Structure
We began by performing rigorous validation, checking for row duplicates, empty columns, and title repetitions.
Duplicate Analysis: The… See the full description on the dataset page: https://huggingface.co/datasets/Liori25/10k_recipes.think-10k
think-10k
A dataset with extract rows the dataset in this collection.
List of categories:
general_qa
code
science_qa
math
creative_writing
brainstorming
summarization
information_extraction
classification
Dataset structure
main/train.csv -- the full 10k training datasft/train.csv -- 2k rows for SFT warmup before RLrl/train.csv -- 8k forws for RL
odia_context_10K_llama2_set
Dataset Card for odia_context_10k_llama2_set
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.1_pattern_10Kplus_myanmar_sentences
🧠 1_pattern_10Kplus_myanmar_sentences
A structured dataset of 11,452 Myanmar sentences generated from a single, powerful grammar pattern:
📌 Pattern:
Verb လည်း Verb တယ်။
A natural way to express repetition, emphasis, or causal connection in Myanmar.
💡 About the Dataset
This dataset demonstrates how applying just one syntactic pattern to a curated verb list — combined with syllable-aware rules — can produce a high-quality corpus of over 10,000 valid… See the full description on the dataset page: https://huggingface.co/datasets/freococo/1_pattern_10Kplus_myanmar_sentences.10K_MELD_Plus_v1.0
Synthetic MELD-Plus (10K Patients)
Watch a demo
This dataset contains 10,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days.
All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/10K_MELD_Plus_v1.0.
