datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
SimpleStories
📘📕 SimpleStories 📙📗
SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository.
If you'd like to commission other languages or story formats, feel free to send mail.
When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.ml-ai-engineer-sft
DuoNeural ML/AI Engineer SFT Dataset
A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments.
Why this dataset exists
Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/TinyStories.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.gsm8k-promise
GSM8K-Promise
A transform of openai/gsm8k (the
main subset) into a "promise" format: the original chain-of-thought is kept
in answer, and an added promise_answer rewrites the solution so the model
emits only operations and operands — every computed value is a var that an
external tool resolves. The model itself never computes, compares, or rounds, so
the reasoning is machine-translation-safe (numbers don't get mangled in
translation) and tool-verifiable.
train: 7,473 rows ·… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/gsm8k-promise.WebApp1K-Duo-React-Generations
WebApp1K-Duo-React-Generations
A comprehensive evaluation dataset containing React component generations from 32 state-of-the-art AI models on 1,000 paired web application scenarios.
Dataset Description
This dataset extends the original WebApp1K-Duo-React benchmark by including actual code generations from major AI models. Each row contains a paired web application scenario (combining two functionalities) along with generated React components from 32 different models and… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/WebApp1K-Duo-React-Generations.LPM-24-extend
LPM-24 dataset
This dataset was used in paper Mol2Lang-VLM: Vision- and Text-Guided Generative Pre-trained Language Models for Advancing Molecule Captioning through Multimodal Fusion
DOI: https://doi.org/10.18653/v1/2024.langmol-1.12
GitHub: https://github.com/nhattruongpham/mol-lang-bridge/tree/mol2lang
This dataset contains:
SELFIES strings (converted by selfies package)
SMILES strings
Molecular images (converted by RDKit)
Molecular captions.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/duongttr/LPM-24-extend.chebi-20-newThis dataset was used in paper Mol2Lang-VLM: Vision- and Text-Guided Generative Pre-trained Language Models for Advancing Molecule Captioning through Multimodal Fusion
DOI: https://doi.org/10.18653/v1/2024.langmol-1.12
GitHub: https://github.com/nhattruongpham/mol-lang-bridge/tree/mol2lang
This dataset contains:
SELFIES strings (converted by selfies package)
SMILES strings
Molecular images (converted by RDKit)
Molecular captions.
Citation
If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/duongttr/chebi-20-new.
