datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
hausa
Hausa Dataset
The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Hausa words.
Key Features
Core Vocabulary Categories:
Pronouns with gender distinctions (kai/ke for masculine/feminine 'you')
Verbs covering daily activities and essential actions
Nouns spanning family, nature, time, and cultural concepts
Adjectives with proper Hausa formations
Numbers from basic counting to large values
Time… See the full description on the dataset page: https://huggingface.co/datasets/0xnu/hausa.ukr-tg-satire
Ukrainian Telegram satire & troll posts
A corpus of 92k posts from 17 public Telegram channels in the
satire/irony/parody register, harvested July 2026 via the public Telegram
API (history reaching back to 2018 for some channels). The channels are
Ukrainian-audience but the posts are a Ukrainian/Russian mix (the
main truha and sria_news feeds write mostly in Russian, the regional
truexa* branches mostly in Ukrainian), so the corpus carries
language: [ru, uk].
Two flavors are… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire.ukr-tg-media
Ukrainian Telegram media posts
A large corpus of Ukrainian news posts harvested from 23 public
Telegram channels (July 2026 snapshot; per-channel history reaching back
to 2016 for some). Channels include national and regional media
(BBC Ukrainian, Zaxidnet, Suspilne News, Nexta, Censor.net, Ukrainska
Pravda, Interfax-Ukraine, Texty, Hromadske, etc.), regional outlets
(Huyovy Kharkiv, Kyiv Real, Oko, Lacheny T, KSPZSU…), and official
accounts (Ministry of Defence of Ukraine, V.… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-media.dvach_chat
Двач чат — group chat (Telegram)
The group chat that runs alongside the public «Двач» channel
(hausmer/dvach). While the channel is a single-author broadcast, this is a
multi-participant group: every message carries the name of the person who
posted it. This corpus is a text-only snapshot of the whole group history
as exported by a member, with media stripped.
This corpus contains 568516 text messages from 2022-09-02T15:35:10+02:00 to
2026-09-15T21:18:52+02:00, posted by 4851… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/dvach_chat.hausa-stopwords-corpus
Hausa Stopword Candidates and Frequency Scores
A reproducible Hausa lexical resource containing frequency-scored stopword candidates. This repository is organized for inspection, preprocessing experiments, and future Hausa-speaker review. It does not publish a final stopword list or a final human-reviewed stopword count.
Quick navigation
Need
Go to
Browse candidates in the Dataset Viewer
data/hausa_stopword_candidates.jsonl
Efficient analysis… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/hausa-stopwords-corpus.truexa-comments
Truha audience comments
45,918 real audience comments from the public satirical news channel
«Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and
cleaned (ads, links, mentions, ultra-short fragments, duplicates and
phone numbers dropped — 50,000 raw → 45,918 clean).
The audience writes in both Ukrainian and Russian (a mix, not a
clean split — a share of comments code-switch between the two), so the
corpus carries language: [ru, uk].
These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.uxfromhell-chat
Adivy Chat (Адовый чатик)
Message history of the private-style group chat «Адовый чатик» — 5,578 messages
spanning 2024-11-15 to 2026-07-18, exported by the chat owner (84
participants). This is a raw Telegram chat export: the fields below mirror the
export 1:1 after scrubbing (see disclosure below).
Contents
Column
Type
Description
id
int64
Telegram message ID
date
string
ISO-8601 UTC timestamp of the message
text
string
Message body
sender_id… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/uxfromhell-chat.hausdorff-dimension-spectrum
Hausdorff Dimension Spectrum: All Subsets of {1,...,20}
First complete computation of the Hausdorff dimension of E_A for all 2^20 - 1 = 1,048,575 non-empty subsets A of {1,...,20}. Computed via transfer operator + Chebyshev collocation (N=40) on NVIDIA RTX 5090.
This dataset does not exist anywhere in the published literature.
Files
spectrum_n5.csv — All 31 subsets of {1,...,5}
spectrum_n10.csv — All 1,023 subsets of {1,...,10}
spectrum_n20.csv — All 1,048,575 subsets of… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/hausdorff-dimension-spectrum.hausa-ajami-blindspot-evalChEMBL36-SELFIES
ChEMBL 36 SELFIES
Pre-training dataset for ModernMolBERT. Contains ~2.4M drug-like small molecules from ChEMBL 36 represented as SELFIES strings.
Dataset details
field
value
source
lukaskim/ChEMBL-36
representation
SELFIES
train rows
2,390,314
validation rows
24,228
total rows
2,414,542
min heavy atoms
3
max heavy atoms
100
max MW
1000.0
deduplicated by
InChIKey
split method
deterministic hash on InChIKey
valid fraction
0.01… See the full description on the dataset page: https://huggingface.co/datasets/HauserGroup/ChEMBL36-SELFIES.lnl-hausa
Dataset Card for "lnl-hausa"
More Information needed
naijavoices-hausa-filteredHausaHate
Evaluation Benchmark for Hausa Hate Speech Detection
We introduce the first expert annotated corpus of Facebook comments for Hausa hate speech detection.
The corpus titled HausaHate comprises 2,000 comments extracted from Western African Facebook pages and
manually annotated by three Hausa native speakers, who are also NLP experts.
The corpus was annotated using two different layers. We first labeled each comment according to a
binary classification: offensive versus… See the full description on the dataset page: https://huggingface.co/datasets/franciellevargas/HausaHate.reklamation24_haus-reinigung-intent
Dataset Card for "reklamation24_haus-reinigung-intent"
More Information needed
africa-burkina-faso-enquete-longitudinale-a-haute-frequence-sur-limpact-de-la-ee7f0fa9
Enquete Longitudinale a Haute Frequence Sur Limpact De La | Africa (Institut national de la statistique et de la démographie)
52 rows - 1 Africa country/area - 2019-2020 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 52 rows from Institut national de la statistique et de la démographie, covering Enquete Longitudinale a Haute Frequence Sur Limpact De La. It is published as ML-ready Parquet with consistent Hugging Face metadata… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-burkina-faso-enquete-longitudinale-a-haute-frequence-sur-limpact-de-la-ee7f0fa9.reklamation24_haus-reinigung
Dataset Card for "reklamation24_haus-reinigung"
More Information needed
abfall-im-haushalt-simuliert
Abfall im Haushalt (simuliert)
High-quality synthetic dataset for machine learning and data analysis.
📊 Dataset Overview
Rows: 500,000 (sample: 1,000)
Columns: 11
Quality Score: 99%
Format: CSV
Type: 100% Synthetic
🎯 Features
datum
kosten
abfall_menge
abfall_art
kosten_pro_kg
abfall_quartal
kosten_quartal
abfall_art_quartal
kosten_abfall_art_pro_kg
abfall_quartal_jahr
... and 1 more
💡 Use Cases
🤖 Machine learning model training
📊 Data… See the full description on the dataset page: https://huggingface.co/datasets/MarvHins/abfall-im-haushalt-simuliert.Hausa_sentiment_analysisLerobot-datasethaulroll-new-motor-carriers
Haulroll: new US trucking companies (FMCSA new carrier registrations)
Need it fresh, filtered or via API? This free file is a snapshot (new FMCSA motor carriers at the last refresh), last updated 2026-09-25.
Get an email when this dataset updates: free, double opt-in, unsubscribe any time.
Every motor carrier, broker and freight forwarder that received a new USDOT number in the last 90 days, from the FMCSA Company Census File, one row per company, company-level fields… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/haulroll-new-motor-carriers.
