datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VlaserGanjoor-CorpusEnglish | فارسی
Ganjoor-Corpus
The Ganjoor poetry corpus as four tables covering poems, books of poetry, and poet information. This corpus can be used for training models, statistical work, and building other datasets.
Configs
poets: 234 poets.
field
description
poet_id
Ganjoor poet id
name, nickname
full name and pen name
url
Ganjoor path
birth_year, death_year
lunar Hijri; birth_year_valid / death_year_valid say whether Ganjoor marks the date as… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/Ganjoor-Corpus.MMM-datasets-TestsetMultilingual Mutual Reinforcement Effect Mix Datasets
This is a Training set of OIELLM.
This Train set already formatted by OIELLM's format. The test set is in the another page in huggingface.
The MMM support 3 languages (English, Chinese and Japanese). And you must use task instruct words to define kind of task.
Mutual Reinforcement Effect.
OIELLM's input and output
MMM Dataset
The following is input and output format:
{
"input": "In 1953, filming of "On the Waterfront" starring… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/MMM-datasets-Testset.Ganjoor-Rhythm-BenchEnglish | فارسی
Ganjoor-Rhythm-Bench
A comprehensive dataset of Persian poems along with ther metrs (وزن عروضی) which is intended to be used for benchmarking LLMs, text classification models or any other model tasked with detecting the Rhythm of a Persian poem. Sourced from the Ganjoor dataset.
Configs
verses: 2,240,985 mesras, 107 metres. Use this for training or other work.
field
description
verse
one hemistich
rhythm
arkān string
poem_id
Ganjoor… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/Ganjoor-Rhythm-Bench.pii-masking-400k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
AI4Privacy Dataset Analytics 📊
Dataset Overview
Total entries: 406,896
Total tokens: 20,564,179
Total PII tokens: 2,357,029
Number of PII classes in public dataset: 17
Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.pov-authganitaIf you use this dataset, please cite our paper,
@misc{niyogi2024paramanuganitalanguagemodelmathematical,
title={PARAMANU-GANITA: Language Model with Mathematical Capabilities},
author={Mitodru Niyogi and Arnab Bhattacharya},
year={2024},
eprint={2404.14395},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2404.14395},
}
recipe-gantt
Summary
A very small dataset of input recipes and output recipe gantt charts in TSV format where each column represents a method step and each row represents a single ingredient. Cells of the output TSV are populated with X if that ingredient is used in that step.
It was used to fine-tune pocasrocas/recipe-gantt-v0.1.
Format
It follows the alpaca instruction/input/response format, shared here in .jsonl format for easy use with libraries such as axolotl.… See the full description on the dataset page: https://huggingface.co/datasets/pocasrocas/recipe-gantt.Lord-ganeshfinance-dpo-dataset
Personal Finance DPO Dataset
A comprehensive Direct Preference Optimization (DPO) dataset containing 5,000 high-quality examples focused on personal finance, investment advice, and tax planning topics.
Dataset Description
This dataset was created to train language models to provide better financial advice through preference-based learning. Each example contains a financial question paired with a "chosen" (high-quality) response and a "rejected" (lower-quality) response… See the full description on the dataset page: https://huggingface.co/datasets/gandhiraketla277/finance-dpo-dataset.gang-of-four-training-data
Gang of Four Training Data
Training data for the Gang of Four neural AI, generated from ExpertStrategy self-play.
Dataset Details
Format: JSONL (one JSON object per line)
Size: ~1M game states
Source: ExpertStrategy vs ExpertStrategy games
Schema
Each line contains:
{
"state": [328 floats],
"action_mask": [40 floats],
"action_idx": int,
"declared_last_card": bool
}
Usage
from training.dataset import GangOfFourDataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/quintana42/gang-of-four-training-data.testdatapython_ablationjavascript_ablationOL-science-mcq_essays_json_DS20251222cube
20251222cube
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
java_ablationphp_ablationmasterchess-datasetGANBASS-Knowledge
GANBASS Car Detailing Knowledge (GANBASS洗車知識データセット)
概要 (Overview)
洗車専門店・カーディテイリングブランド「GANBASS」が提供する、プロフェッショナルな洗車・メンテナンス知識のデータセットです。
AIに「塗装を傷つけない正しい洗車方法」や「適切なケミカルの使用順序」を学習させることを目的としています。
データ詳細
instruction: ユーザーからの質問(洗車、メンテナンス、製品選びなど)
output: GANBASS流の回答(塗装保護を最優先とした論理的なアドバイス)
情報源
洗車専門店GANBASS公式の知識(マニュアル、SNS、ブログ等)に基づいています。
推奨用途
カーケア特化型AIチャットボットのトレーニング
洗車アドバイザーAIの開発
LLM(大規模言語モデル)への専門知識の注入
License
MIT License
go_ablationgandan5personal-finance-sft-181kruby_ablationorangecube20250809
orangecube20250809
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
financial-qa-chat-formattedgandalf_therapistThis is a test
legal-cases-1k-casesYoutube_GaneshaGroup994gandan
