datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pile-uncopyrighted
Pile Uncopyrighted
In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA.
MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.pi-mono
Coding agent session traces for Pi
This dataset contains redacted coding agent session traces collected while working on the Pi OSS project.
Canonical source repository: git@github.com:earendil-works/pi.git
The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-mono.pi-mono-sessions
Coding agent session traces for thomasmustier/pi-mono-sessions
This dataset contains redacted coding agent session traces collected while working on earendil-works/pi. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-mono-sessions.pile-test-valcombust-labs_pi-mono-dockersouth-african-monolingual-corpora-jsonl
South African Languages Pretraining Dataset
This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections.
The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity
Languages Included
Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.c5_eng_nfp_nemo_4dumpsbadlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well.
All traces present are training safe and teich compatible
c5-nemotron-filtertest-2022-05This is the intersection of the 2022-05 snapshot of Nemotron-CC with monology/c5_2022-05_nofalsepositives, based on URL.Testing to see if this is a good filtering step for C5-style data. A good experiment would be expanding this to cover multiple dumps (to account for local vs global deduplication) and then training very small models over this and c5-nfp (fineweb subset) for ~100B tokens each to see which has greater potential for long token horizon training over multiple dumps.
minuszero-indian-autonomous-driving-monocam
Minus Zero Indian Urban Autonomous Driving Dataset - Single Camera
Overview
This dataset provides original single-camera autonomous driving recordings in MCAP format. It is designed for research on camera perception, H.265 video pipelines, localization, GNSS/pose integration, and robotics data tooling.
Depending on the recording, supporting channels include recorded or live GNSS/pose.
The dataset is public for personal, educational, and research use under CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-monocam.monoids-100
Monoids
Sequences of elements from various monoids along with their products. Contains data from monoids over symmetric groups (S2, S3, S4, S5, and S6); alternating groups (A3, A4, A5, and A6); and several cyclic groups chosen to match the order of the symmetric and alternating groups (Z6, Z12, Z24, Z60, Z120, Z360, and Z720).
Each group has its own eponymous split; each split contains 100,000 records of 5 features:
sequence: list[int] A sequence of elements in in the range… See the full description on the dataset page: https://huggingface.co/datasets/jowenpetty/monoids-100.bagel-v0.3Just a backup of jondurbin/bagel-v0.3 in .jsonl.zst format.
openmathinstruct-correct-traintelugu-monolingual-datasethwtcm-deepseek-r1-distill-data
简介
DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。
7B模型微调效果
模型表现出了推理能力,准确性有待继续验证。
我们的其他产品
中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。
。。。还有很多
Citation
If you find this project useful in your research, please consider cite:
@misc{hwtcm2024,
title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.hwtcm
Description
This dataset can be used to evaluate the capabilities of large language models in traditional Chinese medicine and contains multiple-choice, multiple-answer, and true/false questions.
Changelog
2024-08-28: Added 7226 questions.
2024-08-09: The benchmark code is available at https://github.com/huangxinping/HWTCMBench.
2024-08-02: System prompts are removed to ensure the purity of the evaluation results.
2024-07-20: Debut.
Examples
multiple-answers… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm.phictnlUsed in the reproduction of https://arxiv.org/abs/2309.08632. Not recommended for LLMs.
no-robots-vs-robotsBased on a subset of No Robots. Rejected responses generated with gpt-oss-20b.
hwtcm-sft-v1
A dataset of Tradictional Chinese Medicine (TCM) for SFT
一个用于微调LLM的传统中医数据集
Introduction
This repository contains a dataset of Traditional Chinese Medicine (TCM) for fine-tuning large language models.
Dataset Description
The dataset contains 7,096 Chinese sentences related to TCM. The sentences are collected from various sources on the Internet, including medical websites, TCM forums, and TCM books. The dataset is generated or judged by various LLMs, including… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-sft-v1.c5-en-filteredThis is the 2022-05 snapshot of BramVanroy/CommonCrawl-CreativeCommons, filtered by:
Extracting the URLs from the dataset
Getting documents that match those URLs from the corresponding snapshot of togethercomputer/RedPajama-Data-V2
Keeping only the head and middle partitions of ccnet
Keeping documents with at least 50 words and a mean word length between 3 and 10 inclusive
In total we keep 4,553,263 of the 15,239,155 total documents.
monomerMonomer molecule dataset with properties
This is dataset was curated in house and used to fine-tune chemistry language models.
african-transcribed-speech-monolingual
African Transcribed Speech — Monolingual
Sentence-level monolingual text for 15 African languages, derived from translated religious speech transcriptions. Intended as reference text for evaluating speech machine translation (e.g. BLEU scoring), and as a monolingual corpus for language modeling / tokenizer training.
Each language is a separate subset — load with e.g. load_dataset("<repo>", "fat").
Coverage
Language
Code
Sentences
Malagasy
mlg
108,333… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/african-transcribed-speech-monolingual.crh_monocorpus
Crimean Tatar corpus
Overview
In an effort to democratize research on low-resource languages, we release CrimeanTatarMonocorpus dataset, a books corpus consisting of materials from 100+ unique sources in the Crimean Tatar Language.
All annotation was done for Crimean Tatar corpus project. The text is provided as is, so some pre-processing is needed. For example, some subtitles contain timestamps information.
Both Cyrillic and Latin alphabets are used in text, so to switch… See the full description on the dataset page: https://huggingface.co/datasets/QIRIM/crh_monocorpus.70k_monomer_properties70K monomer compound formulas and properties
This is dataset was curated in house and used to fine-tune chemistry language models.
nuer_greetings_monolingual_pairs_eng_nuer_dinkathai-land-tax-full-triplets
Thai Land & Buildings Tax — Full Triplet Dataset
Dataset Name: monoboard/thai-land-tax-full-triplets
Language: Thai (th)
Tasks: Legal Retrieval • RAG • Contrastive Learning • Triplet Loss • Embedding Training
Overview
This dataset provides a legally verified retrieval corpus for training Thai-language retrieval models under the Land and Buildings Tax Act (B.E. 2562) and related regulations.
Each example includes:
A legal question (query)
One or more oracle passages (pos)… See the full description on the dataset page: https://huggingface.co/datasets/monoboard/thai-land-tax-full-triplets.medical_meadow_alpacarosettacode-sharegptno-robots-subsetearnings_call_mono
Motley Fool Earnings Call Mono (Private)
This private dataset contains cleaned monolingual earnings-call text chunks prepared from:
Kaggle dataset: tpotterer/motley-fool-scraped-earnings-call-transcripts
Split
train: 135306 rows
Columns
id: chunk identifier
text: cleaned source text chunk
ticker: ticker symbol
exchange: exchange string from source metadata
date: call date string from source metadata
section: prepared or qa
Processing Summary… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/earnings_call_mono.
