datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
mandarin-most-common-words-tr-en
Mandarin Most Common Words (TR-EN)
Overview
The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis.
This dataset was created by Stephanie Liu and Kamil Murat Yilmaz.
Dataset Content
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
gc_santegc_sante (Grande Cause Santé; Health Great Cause) is a French citizen consultation on how to act collectively for better health, prevention an well-being in France held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content and has a unique id proposal_id. the topic and subtopic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/gc_sante.ingerenceIngerence is a French citizen consultation on how to combat information manipulation due to foreign digital interference held in 2025-2026.
Data
This dataset contains three subsets:
proposals contains the written proposals in French. Each proposal has a text content written by an author with a unique author_id, and has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on proposals (defined by proposal_id).… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/ingerence.eurhopeEurhope is a European-Union wide multilingual citizens consultation on the future of the European Union. It was held in 2024.
Data
This dataset contains two subsets:
proposals contains the written proposals in French. Each proposal has a text content written by a author author_id and has a unique id proposal_id. Each proposal has one of 22 language given in its language column.
votes contains the votes of users on propositions. Each user has a unique id user_id and votes on… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/eurhope.steuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025.
Data
This dataset contains three subsets:
proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.common-voice-kinyarwanda-english-dataset
Kinyarwanda-English Commonvoice dataset
A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR
Note: The audio dataset shall be added in the future
commonvoice-mnhebrew-lexical-references
Hebrew Lexical Reference Indices
Four structured, Strong's-linked transcriptions of external Hebrew (and one Hebrew↔Greek) lexical
reference sources. These are not our own synonymy judgments — each config faithfully represents
what an established outside source, or an actual historical translation record, already asserts (an
etymological dictionary's own root groupings, a WordNet's own synset membership, five named scholars'
own verified structural analysis, the Septuagint's own… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/hebrew-lexical-references.mozilla-common-voice-23-bel-texts-exportCommonVoiceAzCommonVoices20_ro
Common Voices Corpus 20.0 (Romanian)
Common Voices is an open-source dataset of speech recordings created by
Mozilla to improve speech recognition technologies.
It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide.
Challenges: The raw dataset included numerous recordings with incorrect transcriptions
or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements
essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.commoncommonvoicebadini
Northern Kurdish (Arabic Script) ASR Dataset
Dataset Description
Northern Kurdish is the most widely spoken variant of the Kurdish language and is used across all parts of Kurdistan. Although it is mainly written today in the Latin script, it was historically written in the Arabic script. The Arabic script is still used for this dialect in Southern Kurdistan, particularly in the Duhok province of the Kurdistan Regional Government (KRG).Similarly, the primary writing… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/commonvoicebadini.common-variety-d04dad
common-variety-d04dad
Synthetic products test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.ovos-common-query-intentsDataset focusing on general knowledge questions, it distinguishs between queries aimed at a specific information source (wikipedia, wolfram alpha, duckduckgo..) and non specific queries
commoncatalog-cc-by-ext
CommonCatalog CC-BY Extention
このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。
以下の情報が追加されています。
Phi-3 VisionでDense Captioningした英語キャプション
英語キャプションをPhi-3 Mediumで日本語化した日本語キャプション
主キーはphotoidですので、CommonCatalog CC-BYと結合するなりして使ってください。
streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。
License
画像がCC BYなため、わかりやすくCC BYにしています。したがって、商用利用可能です。
Sample Code
import pandas
from datasets import load_dataset
df=pandas.read_csv("commoncatalog-cc-by-phi3-ja.csv")
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncatalog-cc-by-ext.philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.common-strategy-d0489c
common-strategy-d0489c
Synthetic sensors test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Ember-Wisp/common-strategy-d0489c.common_voice_13_0_zh_pseudo_labelledcommon-voicecommoncatalog-cc-by-recap
CommonCatalog CC-BY Recaptioning
このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。 以下の情報が追加されています。
Phi-3 VisionでDense Captioningした英語キャプション
主キーはphotoidですので、CommonCatalog CC-BYと結合するなりして使ってください。 streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。
Sample Code
import pandas
from datasets import load_dataset
from tqdm import tqdm
import json
df=pandas.read_csv("commoncatalog-cc-by-phi3.csv")
dataset = load_dataset("common-canvas/commoncatalog-cc-by",split="train",streaming=True)… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncatalog-cc-by-recap.common_voice_13_0_thai_small_pseudo_labelledcommon_voice_16_1_spanish_test_set
Dataset Card for Common Voice Corpus 16 Spanish Dataset
Acknowledgement
The dataset belongs to COMMON VOICE MOZILLA FOUNDATION.
I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main)
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Languages
Spanish
How to use
The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.dividend-common-question
Dividend Common Question Dataset
📖 개요
이 데이터셋은 배당주 투자 관련 자주 묻는 질문과 답변을 Alpaca 포맷(instruction, input, output)으로 구성한 학습용 자료입니다.총 10여 개의 샘플이 포함되어 있으며, 파인튜닝 실습이나 자연어 처리 모델 학습에 활용할 수 있습니다.
📂 데이터 구조
데이터는 CSV 파일(dividend-common-question.csv)로 제공되며, 다음과 같은 열을 포함합니다:
instruction: 모델에게 주는 지시문 (예: "배당성향이 높아야 좋은 건가요?")
input: 지시문을 수행하는 데 필요한 추가 입력 (없으면 빈칸)
output: 모델이 생성해야 하는 답변 (예: "꼭 그렇다고 할 수는 없습니다...")
예시
instruction,input,output
"시가배당수익률이 높으면 좋은… See the full description on the dataset page: https://huggingface.co/datasets/ycryu/dividend-common-question.common_voice_13_0_mn_pseudo_test_smallcommoncanvas-cc-by-recap-2
CommonCatalog CC-BY Recaptioning 2
このリポジトリはCommonCatalog CC-BYを拡張して、追加の情報を入れたものです。 以下の情報が追加されています。
Florence-2-large-ftでDense Captioning (More detailed caption) した英語キャプション
streaming=Trueで読み込むと同じ順に読み込まれますのでそれを利用するのが一番楽です。
Sample Code
import pandas
from datasets import load_dataset
from tqdm import tqdm
import json
df=pandas.read_csv("commoncatalog-cc-by-phi3.csv")
dataset = load_dataset("common-canvas/commoncatalog-cc-by",split="train",streaming=True)
data_info=[]
for… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/commoncanvas-cc-by-recap-2.
