CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mabidan /ganjoor Dataset Card for Dataset Name This is the csv format of the Ganjoor Database that is published in their github Dataset Details Curated by: Navid Abbaspoor Language(s) (NLP): Persian (Farsi) License: Creative Commons Attribution 4.0 International (cc-by-4.0) Dataset Description This dataset contains almost all of poems by Iran's great poets through many many past years till now. The original database was tabular, that I convert it to a csv format that… See the full description on the dataset page: https://huggingface.co/datasets/mabidan/ganjoor.texttext-generation100K<n<1M3 likes736 downloads2y agoHugging Face02dipta007 /Ganit Ganit: A Difficulty-Aware Bengali Mathematical Reasoning Dataset Dataset Description Ganit (গণিত, Bengali for "mathematics") is a rigorously-processed, difficulty-aware Bengali mathematical reasoning dataset designed for training and evaluating LLMs on Bengali math problems. It is the first Bengali math dataset with: Difficulty stratification based on LLM pass@k scores Decontamination against standard benchmarks (MGSM, MSVAMP) Verifiable numerical… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/Ganit.tabulartext-generation10K<n<100K0 likes128 downloads6mo agoHugging Face03ganeshjcs /hindi-article-summarization Summary hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.texttext-generation10K<n<100K0 likes57 downloads3y agoHugging Face04Ganasekhar /pii-masking-400k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.texttext-classification100K<n<1M0 likes55 downloads7mo agoHugging Face05gang100 /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/gang100/TinyStories.texttext-generation1M<n<10M0 likes42 downloads2mo agoHugging Face06Gandalf1 /indian-finance-synthetic-phase2-cleaned Indian Finance Synthetic Dataset (Phase 2 - Final Clean) Dataset Description 14,763 high-quality synthetic conversations about Indian personal finance, optimized for fine-tuning. Recent Updates ✅ v3 (Final): Removed 14 samples with empty content messages ✅ v2: Removed 58 incomplete conversations ✅ v1: Tools optimization (82.5% size reduction) All conversations are now complete and properly formatted for training. Key Features Clean… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2-cleaned.tabulartext-generation10K<n<100K1 likes37 downloads4mo agoHugging Face07ganeshjcs /hindi-headline-article-generation Summary hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.texttext-generation100K<n<1M1 likes27 downloads3y agoHugging Face08Gandalf1 /indian-finance-synthetic-phase2 Indian Finance Synthetic Dataset - Phase 2 A high-quality synthetic dataset of 14,835 Indian personal finance conversations for fine-tuning language models. Dataset Description This dataset contains synthetic conversations between users seeking personal finance advice and a financial assistant (FinEdge). All conversations are tailored to the Indian context, covering tax planning, investments, insurance, goal planning, and more, based on FY 2024-25 regulations. Key… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2.texttext-generation10K<n<100K0 likes20 downloads4mo agoHugging Face09Naveen934 /tamil_data_na_thozhar_gandhi Tamil தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல் Dataset by ஆர். பட்டாபிராமன் Description This dataset contains Tamil தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல் texts by ஆர். பட்டாபிராமன், processed for language model pretraining. Contents 154 text chunks Author: ஆர். பட்டாபிராமன் Genre: தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல் Total chunks: 154 Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Naveen934/tamil_data_na_thozhar_gandhi.texttext-generationn<1K0 likes19 downloads1y agoHugging Face10Ganz00 /Cleaned_ELI5_with_one_responseRedit response from the subredit explainlikeimfive. each question's had multiple response, here it's explode so you will have multiple time the same question with different answer textquestion-answering100K<n<1M0 likes18 downloads2y agoHugging Face11jungter /potao-gang-gc-032026_new is cleaned using a varity of methods like duplicate message removal, then fed to an LLM to go through A and B 0 - n classifying if they are related pairs or not. The result model is here _v3 is cleaned with the same algo, however the LLM is fed 7 messages plus 20 context padding before and after and asked to pair prompt and responses, with 3 messages overlap between chunks. The result model is here With the same training params, _v3 yielded a more talkitive and coherent model, however is… See the full description on the dataset page: https://huggingface.co/datasets/jungter/potao-gang-gc-032026.texttext-generation100K<n<1M0 likes16 downloads7mo agoHugging Face12nafisehNik /ganjoor-ipa-scansiongated Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration A corpus of 124,404 classical Persian poems collected via the Ganjoor API, enriched with two things every poem now has: Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem, including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet). Phonemic transliteration — Latin and IPA for every hemistich, produced by the Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.tabulartext-generation100K<n<1M0 likes12 downloads2mo agoHugging Face13GANBASS /GANBASS-Knowledge GANBASS Car Detailing Knowledge (GANBASS洗車知識データセット) 概要 (Overview) 洗車専門店・カーディテイリングブランド「GANBASS」が提供する、プロフェッショナルな洗車・メンテナンス知識のデータセットです。 AIに「塗装を傷つけない正しい洗車方法」や「適切なケミカルの使用順序」を学習させることを目的としています。 データ詳細 instruction: ユーザーからの質問(洗車、メンテナンス、製品選びなど) output: GANBASS流の回答(塗装保護を最優先とした論理的なアドバイス) 情報源 洗車専門店GANBASS公式の知識(マニュアル、SNS、ブログ等)に基づいています。 推奨用途 カーケア特化型AIチャットボットのトレーニング 洗車アドバイザーAIの開発 LLM(大規模言語モデル)への専門知識の注入 License MIT License texttext-generationn<1K0 likes11 downloads9mo agoHugging Face14Gan1108 /agripartstexttext-generation10K<n<100K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.