CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ganlinyang /Vlasertabular10K<n<100K0 likes930 downloads6mo agoHugging Face02farbodbij /Ganjoor-CorpusEnglish | فارسی Ganjoor-Corpus The Ganjoor poetry corpus as four tables covering poems, books of poetry, and poet information. This corpus can be used for training models, statistical work, and building other datasets. Configs poets: 234 poets. field description poet_id Ganjoor poet id name, nickname full name and pen name url Ganjoor path birth_year, death_year lunar Hijri; birth_year_valid / death_year_valid say whether Ganjoor marks the date as… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/Ganjoor-Corpus.tabular1M<n<10M0 likes395 downloads1mo agoHugging Face03ganchengguang /MMM-datasets-TestsetMultilingual Mutual Reinforcement Effect Mix Datasets This is a Training set of OIELLM. This Train set already formatted by OIELLM's format. The test set is in the another page in huggingface. The MMM support 3 languages (English, Chinese and Japanese). And you must use task instruct words to define kind of task. Mutual Reinforcement Effect. OIELLM's input and output MMM Dataset The following is input and output format: { "input": "In 1953, filming of "On the Waterfront" starring… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/MMM-datasets-Testset.text100K<n<1M1 likes247 downloads2y agoHugging Face04farbodbij /Ganjoor-Rhythm-BenchEnglish | فارسی Ganjoor-Rhythm-Bench A comprehensive dataset of Persian poems along with ther metrs (وزن عروضی) which is intended to be used for benchmarking LLMs, text classification models or any other model tasked with detecting the Rhythm of a Persian poem. Sourced from the Ganjoor dataset. Configs verses: 2,240,985 mesras, 107 metres. Use this for training or other work. field description verse one hemistich rhythm arkān string poem_id Ganjoor… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/Ganjoor-Rhythm-Bench.tabulartext-classification1M<n<10M1 likes115 downloads1mo agoHugging Face05Ganasekhar /pii-masking-400k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.texttext-classification100K<n<1M0 likes55 downloads7mo agoHugging Face06ganhteam321 /pov-authtextn<1K0 likes51 downloads20d agoHugging Face07gyanai /ganitaIf you use this dataset, please cite our paper, @misc{niyogi2024paramanuganitalanguagemodelmathematical, title={PARAMANU-GANITA: Language Model with Mathematical Capabilities}, author={Mitodru Niyogi and Arnab Bhattacharya}, year={2024}, eprint={2404.14395}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2404.14395}, } text100K<n<1M2 likes38 downloads2y agoHugging Face08pocasrocas /recipe-gantt Summary A very small dataset of input recipes and output recipe gantt charts in TSV format where each column represents a method step and each row represents a single ingredient. Cells of the output TSV are populated with X if that ingredient is used in that step. It was used to fine-tune pocasrocas/recipe-gantt-v0.1. Format It follows the alpaca instruction/input/response format, shared here in .jsonl format for easy use with libraries such as axolotl.… See the full description on the dataset page: https://huggingface.co/datasets/pocasrocas/recipe-gantt.textn<1K1 likes33 downloads2y agoHugging Face09VenkataRamanaKurumallajaddangi /Lord-ganeshtextn<1K0 likes30 downloads2mo agoHugging Face10gandhiraketla277 /finance-dpo-dataset Personal Finance DPO Dataset A comprehensive Direct Preference Optimization (DPO) dataset containing 5,000 high-quality examples focused on personal finance, investment advice, and tax planning topics. Dataset Description This dataset was created to train language models to provide better financial advice through preference-based learning. Each example contains a financial question paired with a "chosen" (high-quality) response and a "rejected" (lower-quality) response… See the full description on the dataset page: https://huggingface.co/datasets/gandhiraketla277/finance-dpo-dataset.text1K<n<10K0 likes19 downloads1y agoHugging Face11gangli71 /testdatatextn<1K0 likes17 downloads9mo agoHugging Face12gangiswag /python_ablationtext100K<n<1M0 likes15 downloads2y agoHugging Face13gangiswag /javascript_ablationtext100K<n<1M0 likes14 downloads2y agoHugging Face14GaniduA /OL-science-mcq_essays_json_DStext10K<n<100K0 likes14 downloads1y agoHugging Face15gannbayar /20251222cube 20251222cube This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularroboticsn<1K0 likes13 downloads9mo agoHugging Face16gangiswag /java_ablationtext100K<n<1M0 likes12 downloads2y agoHugging Face17gangiswag /php_ablationtext100K<n<1M0 likes12 downloads2y agoHugging Face18gandensang /masterchess-datasettext10M<n<100M0 likes12 downloads9mo agoHugging Face19GANBASS /GANBASS-Knowledge GANBASS Car Detailing Knowledge (GANBASS洗車知識データセット) 概要 (Overview) 洗車専門店・カーディテイリングブランド「GANBASS」が提供する、プロフェッショナルな洗車・メンテナンス知識のデータセットです。 AIに「塗装を傷つけない正しい洗車方法」や「適切なケミカルの使用順序」を学習させることを目的としています。 データ詳細 instruction: ユーザーからの質問(洗車、メンテナンス、製品選びなど) output: GANBASS流の回答(塗装保護を最優先とした論理的なアドバイス) 情報源 洗車専門店GANBASS公式の知識(マニュアル、SNS、ブログ等)に基づいています。 推奨用途 カーケア特化型AIチャットボットのトレーニング 洗車アドバイザーAIの開発 LLM(大規模言語モデル)への専門知識の注入 License MIT License texttext-generationn<1K0 likes11 downloads9mo agoHugging Face20gangiswag /go_ablationtext100K<n<1M0 likes10 downloads2y agoHugging Face21urjinchimed /gandan5textn<1K0 likes9 downloads1y agoHugging Face22Gandalf1 /personal-finance-sft-181ktext100K<n<1M0 likes9 downloads5mo agoHugging Face23gangiswag /ruby_ablationtext10K<n<100K0 likes8 downloads2y agoHugging Face24gannbayar /orangecube20250809 orangecube20250809 This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularroboticsn<1K0 likes7 downloads1y agoHugging Face25gandhiraketla277 /financial-qa-chat-formattedtext1K<n<10K0 likes5 downloads1y agoHugging Face26brockhouston /gandalf_therapistThis is a test textn<1K0 likes3 downloads2y agoHugging Face27ganeshnaiknavare3656 /legal-cases-1k-casestextn<1K0 likes3 downloads1y agoHugging Face28rayzox57 /Youtube_GaneshaGroup994tabularn<1K0 likes2 downloads1y agoHugging Face29urjinchimed /gandantextn<1K0 likes2 downloads1y agoHugging Face30Gantumur /mn_knowledgetextn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.