datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Language_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.fixed-tokenizer-morphscore-segmentsAviationQAAviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering
https://aclanthology.org/2022.icon-main.26/
The paper is accepted in the main conference of ICON 2022.
We create a synthetic dataset, AviationQA, a set of 1 million factoid QA pairs from 12,000 National Transportation Safety Board (NTSB) reports using templates. These QA pairs contain questions such that answers to them are… See the full description on the dataset page: https://huggingface.co/datasets/sakharamg/AviationQA.agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
Saka-Alpaca-v1https://chatgpt.com
IndianLegal-QA
IndianLegal-QA
A question-and-answer dataset derived from Indian legal and government documents,
covering the Constitution of India, the Indian Penal Code, criminal and civil
procedural law, customs and tariff classifications, and numerous central and
state acts. The dataset is suitable for building, fine-tuning, and evaluating
retrieval and question-answering systems over Indian legal text.
This dataset is also hosted on GitHub at Sakib-Dalal/IndianLegal-QA.… See the full description on the dataset page: https://huggingface.co/datasets/Sakib-Dalal/IndianLegal-QA.AeroQArag-eval-ja-repro
RAG Eval JA Repro
Current version / 現行版: v1.1
2026-07-12 更新(v1.1): rag_evaluation_master.csv、採用PDF manifest、PDF checksumを更新し、6月30日公開時のローカル精度検証を同じ4条件で再実行しました。旧版の記述は取り消し線で残し、v1.1の値を併記します。
TL;DR (EN): A derived reproducibility dataset for allganize/RAG-Evaluation-Dataset-JA.
It adds (1) derived *_new answer/question columns (with per-item rationale), and (2) a Wayback-pinned + SHA-256 corpus manifest so anyone can fetch byte-identical source PDFs.
The original CSV is not modified;… See the full description on the dataset page: https://huggingface.co/datasets/SakataConsul/rag-eval-ja-repro.twitter_racism_datasetSakaEval-V1https://chatgpt.com
reddit_twitterDPO_datasetai-models-database
Convly AI Models Database
A continuously updated, hand-verified dataset of 30+ AI language models — specs, licenses, API pricing (USD per 1M tokens), and local-hardware (VRAM) requirements.
Maintained by Convly.ai · Live interactive version: https://convly.ai/models/
Fields
name, slug, convly_url, developer, model_type, modality, parameters, context_window, max_output, license, open_weights, release_date, input_price (USD/1M tokens), output_price (USD/1M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/sakd99/ai-models-database.Zefiromistral_dataFindSUMhachiwari#Origin
The name comes from "hachiwari/はちわれ" (chiikawa/ちいかわ).
RolePlay-v1https://chatgpt.com
vestasv52-scada-windturbine-granadaDesigned and generated by https://simulatexp.dev
Vestas V52 Wind Turbine SCADA Synthetic Dataset - Granada Peri-Urban Installation
This synthetic dataset contains comprehensive SCADA (Supervisory Control and Data Acquisition) data simulating a Vestas V52 wind turbine operating in a peri-urban environment in Granada, Spain. The dataset captures 40,000 one-minute aggregated sensor readings across 24 parameters, simulating realistic operational conditions and fault scenarios for… See the full description on the dataset page: https://huggingface.co/datasets/Sakura81537/vestasv52-scada-windturbine-granada.transcript-formatter-curriculum
Transcript Formatter Curriculum (L0–L5 + RL)
Training data behind
Akash-Sakala/gpt-oss-120b-transcript-formatter-lora:
a layered curriculum that turns raw speech-to-text transcripts into clean,
formatted transcripts. Every row is input (raw transcript) → output
(formatted), across 21 categories spanning punctuation, casing, fillers,
disfluencies, homophones, ITN, proper nouns, URLs/emails, and layout.
Subsets (use the Data Viewer dropdown)
Each curriculum level is… See the full description on the dataset page: https://huggingface.co/datasets/Akash-Sakala/transcript-formatter-curriculum.Multilingal-datasetgender-based-violence-ipv
Gender-Based Violence & Intimate Partner Violence Dataset
Abstract
This dataset provides 30,000 simulated GBV/IPV records (10,000 per scenario) of women in sub-Saharan Africa. Each record contains 45+ variables including violence type (physical, sexual, emotional, economic), risk factors, injuries, mental health consequences, help-seeking behaviour, barriers, and clinical response. Three settings: urban one-stop centre (26% help-seeking), district facility (20%), and… See the full description on the dataset page: https://huggingface.co/datasets/Saksham-443paudel/gender-based-violence-ipv.hachiwari-enhachiwari dataset english version
FindSUMTruncatedtesytinstructionshachiwari-en-csvair_quality_datasetexplorer-customer-objection-dataEmail_Assigning_2024-10-04Email Classification Dataset
Overview
This dataset contains 100 rows of simulated emails intended for training a text classification model. The model can classify emails into departments and determine the criticality of each email.
Columns:
Email: A string containing the email subject and body.
Department: An integer representing the department the email is associated with:
1: Catalogue Department
2: Integration Department
3: Sales Department
4: Onboarding Department
5: Support Department
6:… See the full description on the dataset page: https://huggingface.co/datasets/Sakhrani/Email_Assigning_2024-10-04.
