CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01monology /pile-uncopyrighted Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA. MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.text100M<n<1B175 likes81k downloads3y agoHugging Face02aaaaliou /pi-mono Coding agent session traces for Pi This dataset contains redacted coding agent session traces collected while working on the Pi OSS project. Canonical source repository: git@github.com:earendil-works/pi.git The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-mono.tabulartext-generationn<1K1 likes1.3k downloads4mo agoHugging Face03thomasmustier /pi-mono-sessions Coding agent session traces for thomasmustier/pi-mono-sessions This dataset contains redacted coding agent session traces collected while working on earendil-works/pi. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-mono-sessions.tabulartext-generationn<1K0 likes1.1k downloads3mo agoHugging Face04monology /pile-test-valtext100K<n<1M1 likes708 downloads3y agoHugging Face05rgruchalski /combust-labs_pi-mono-dockertabularn<1K0 likes282 downloads5mo agoHugging Face06SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes209 downloads1y agoHugging Face07monology /c5_eng_nfp_nemo_4dumpstext1M<n<10M0 likes203 downloads1y agoHugging Face08armand0e /badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well. All traces present are training safe and teich compatible tabularn<1K2 likes197 downloads4mo agoHugging Face09monology /c5-nemotron-filtertest-2022-05This is the intersection of the 2022-05 snapshot of Nemotron-CC with monology/c5_2022-05_nofalsepositives, based on URL.Testing to see if this is a good filtering step for C5-style data. A good experiment would be expanding this to cover multiple dumps (to account for local vs global deduplication) and then training very small models over this and c5-nfp (fineweb subset) for ~100B tokens each to see which has greater potential for long token horizon training over multiple dumps. text100K<n<1M0 likes193 downloads1y agoHugging Face10gagandeepreehal /minuszero-indian-autonomous-driving-monocamgated Minus Zero Indian Urban Autonomous Driving Dataset - Single Camera Overview This dataset provides original single-camera autonomous driving recordings in MCAP format. It is designed for research on camera perception, H.265 video pipelines, localization, GNSS/pose integration, and robotics data tooling. Depending on the recording, supporting channels include recorded or live GNSS/pose. The dataset is public for personal, educational, and research use under CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-monocam.textroboticsn<1K1 likes159 downloads27d agoHugging Face11jowenpetty /monoids-100 Monoids Sequences of elements from various monoids along with their products. Contains data from monoids over symmetric groups (S2, S3, S4, S5, and S6); alternating groups (A3, A4, A5, and A6); and several cyclic groups chosen to match the order of the symmetric and alternating groups (Z6, Z12, Z24, Z60, Z120, Z360, and Z720). Each group has its own eponymous split; each split contains 100,000 records of 5 features: sequence: list[int] A sequence of elements in in the range… See the full description on the dataset page: https://huggingface.co/datasets/jowenpetty/monoids-100.tabular1M<n<10M0 likes146 downloads1y agoHugging Face12monology /bagel-v0.3Just a backup of jondurbin/bagel-v0.3 in .jsonl.zst format. text1M<n<10M0 likes104 downloads3y agoHugging Face13monology /openmathinstruct-correct-traintext1M<n<10M0 likes91 downloads3y agoHugging Face14Nikhil8076 /telugu-monolingual-datasettext1M<n<10M0 likes87 downloads1mo agoHugging Face15Monor /hwtcm-deepseek-r1-distill-data 简介 DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。 7B模型微调效果 模型表现出了推理能力,准确性有待继续验证。 我们的其他产品 中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。 。。。还有很多 Citation If you find this project useful in your research, please consider cite: @misc{hwtcm2024, title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.textquestion-answering10K<n<100K3 likes76 downloads2y agoHugging Face16Monor /hwtcm Description This dataset can be used to evaluate the capabilities of large language models in traditional Chinese medicine and contains multiple-choice, multiple-answer, and true/false questions. Changelog 2024-08-28: Added 7226 questions. 2024-08-09: The benchmark code is available at https://github.com/huangxinping/HWTCMBench. 2024-08-02: System prompts are removed to ensure the purity of the evaluation results. 2024-07-20: Debut. Examples multiple-answers… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm.textquestion-answering10K<n<100K2 likes75 downloads2y agoHugging Face17monology /phictnlUsed in the reproduction of https://arxiv.org/abs/2309.08632. Not recommended for LLMs. text10K<n<100K1 likes74 downloads2y agoHugging Face18monology /no-robots-vs-robotsBased on a subset of No Robots. Rejected responses generated with gpt-oss-20b. text1K<n<10K0 likes66 downloads28d agoHugging Face19Monor /hwtcm-sft-v1 A dataset of Tradictional Chinese Medicine (TCM) for SFT 一个用于微调LLM的传统中医数据集 Introduction This repository contains a dataset of Traditional Chinese Medicine (TCM) for fine-tuning large language models. Dataset Description The dataset contains 7,096 Chinese sentences related to TCM. The sentences are collected from various sources on the Internet, including medical websites, TCM forums, and TCM books. The dataset is generated or judged by various LLMs, including… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-sft-v1.textquestion-answering1K<n<10K4 likes53 downloads2y agoHugging Face20monology /c5-en-filteredThis is the 2022-05 snapshot of BramVanroy/CommonCrawl-CreativeCommons, filtered by: Extracting the URLs from the dataset Getting documents that match those URLs from the corresponding snapshot of togethercomputer/RedPajama-Data-V2 Keeping only the head and middle partitions of ccnet Keeping documents with at least 50 words and a mean word length between 3 and 10 inclusive In total we keep 4,553,263 of the 15,239,155 total documents. tabular1M<n<10M0 likes49 downloads1y agoHugging Face21arcadianlee /monomerMonomer molecule dataset with properties This is dataset was curated in house and used to fine-tune chemistry language models. texttable-question-answering10K<n<100K0 likes39 downloads2mo agoHugging Face22AfriSpeech /african-transcribed-speech-monolingual African Transcribed Speech — Monolingual Sentence-level monolingual text for 15 African languages, derived from translated religious speech transcriptions. Intended as reference text for evaluating speech machine translation (e.g. BLEU scoring), and as a monolingual corpus for language modeling / tokenizer training. Each language is a separate subset — load with e.g. load_dataset("<repo>", "fat"). Coverage Language Code Sentences Malagasy mlg 108,333… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/african-transcribed-speech-monolingual.texttext-generation100K<n<1M0 likes36 downloads1mo agoHugging Face23QIRIM /crh_monocorpus Crimean Tatar corpus Overview In an effort to democratize research on low-resource languages, we release CrimeanTatarMonocorpus dataset, a books corpus consisting of materials from 100+ unique sources in the Crimean Tatar Language. All annotation was done for Crimean Tatar corpus project. The text is provided as is, so some pre-processing is needed. For example, some subtitles contain timestamps information. Both Cyrillic and Latin alphabets are used in text, so to switch… See the full description on the dataset page: https://huggingface.co/datasets/QIRIM/crh_monocorpus.texttext-generationn<1K0 likes35 downloads2y agoHugging Face24arcadianlee /70k_monomer_properties70K monomer compound formulas and properties This is dataset was curated in house and used to fine-tune chemistry language models. texttext-generation10K<n<100K0 likes35 downloads2mo agoHugging Face25NaathNLP /nuer_greetings_monolingual_pairs_eng_nuer_dinkatext10K<n<100K0 likes35 downloads1mo agoHugging Face26monoboard /thai-land-tax-full-triplets Thai Land & Buildings Tax — Full Triplet Dataset Dataset Name: monoboard/thai-land-tax-full-triplets Language: Thai (th) Tasks: Legal Retrieval • RAG • Contrastive Learning • Triplet Loss • Embedding Training Overview This dataset provides a legally verified retrieval corpus for training Thai-language retrieval models under the Land and Buildings Tax Act (B.E. 2562) and related regulations. Each example includes: A legal question (query) One or more oracle passages (pos)… See the full description on the dataset page: https://huggingface.co/datasets/monoboard/thai-land-tax-full-triplets.text1K<n<10K0 likes34 downloads10mo agoHugging Face27monology /medical_meadow_alpacatext1K<n<10K3 likes33 downloads3y agoHugging Face28monology /rosettacode-sharegpttext10K<n<100K1 likes33 downloads1y agoHugging Face29monology /no-robots-subsettext1K<n<10K0 likes33 downloads1mo agoHugging Face30alwaysgood /earnings_call_mono Motley Fool Earnings Call Mono (Private) This private dataset contains cleaned monolingual earnings-call text chunks prepared from: Kaggle dataset: tpotterer/motley-fool-scraped-earnings-call-transcripts Split train: 135306 rows Columns id: chunk identifier text: cleaned source text chunk ticker: ticker symbol exchange: exchange string from source metadata date: call date string from source metadata section: prepared or qa Processing Summary… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/earnings_call_mono.text100K<n<1M0 likes32 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.