datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_pro_leaderboard_submissionarabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tigre-hubert-databarranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.mirandese_g2pArquivoDialetalCLUP_ipadataset info: https://cl.up.pt/arquivo/
portuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.EAT
EAT: Expected Answer Type Dataset
A high-quality dataset for Question Classification based on the TREC Question Taxonomy, enhanced with modern categories and strict Expected Answer Type (EAT) validation.
Dataset Summary
The EAT (Expected Answer Type) dataset is designed to train and evaluate NLP models in the task of classifying questions not by their surface keywords, but by the semantic category of their expected answer.
Unlike original TREC datasets, this version… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/EAT.bolt-tightening-quality
螺栓拧紧质量智能检测数据集(Bolt Tightening Quality Detection Dataset)
自建数据集:面向重型机械装配车间螺栓拧紧工序,基于螺栓拧紧「力矩-转角」力学模型模拟生成,用于拧紧质量智能检测与工艺追溯系统的算法开发与验证。
数据集概述
用途:螺栓拧紧过程质量状态自动判定(5 类多分类任务)
样本规模:3000 条完整拧紧周期(基准集 1200 条 + 扩展集 1800 条)
采样频率:100 Hz
采集量:力矩(N·m)、转角(deg)双通道时间序列
螺栓规格:M12 / M16 / M20 / M24
质量标签(5 类)
编码
标签
说明
0
合格
力矩升至目标扭矩后保持
1
欠拧
力矩提前停止,终值低于目标扭矩
2
过拧
力矩越过目标扭矩后停止
3
滑牙
力矩升至峰值后急剧跌落(螺纹滑丝)
4
虚拧
力矩全程低位(错扣/未正确啮合)
目录结构… See the full description on the dataset page: https://huggingface.co/datasets/zerozero01/bolt-tightening-quality.dicionario_barranquenho
Dicionário de Barranquenho
Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references.
Language
Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.sentence-types-multilingual
Little Questions: Multilingual Sentence Types Dataset
A multilingual dataset of 69,300 labeled sentences (9,900 per language) across 6 sentence type categories and 7 languages. Designed for training and evaluating sentence-type classifiers in multilingual contexts.
Dataset Details
Total entries: 69,300 (9,900 × 7 languages)
Languages: English (EN), Spanish (ES), French (FR), German (DE), Italian (IT), Portuguese (PT), Dutch (NL)
Class distribution: 13,200 entries per… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/sentence-types-multilingual.yes-no-multilingual
Yes/No Multilingual Answers Dataset
A dataset of 8,600 conversational utterances for classifying yes/no/ambiguous responses across 43 languages.
Dataset Description
Each sample is a natural language utterance a person might say in response to a yes/no question. The dataset covers three classes:
Label
Description
yes
Affirmation, agreement, or confirmation
no
Negation, refusal, or disagreement
None
Genuinely ambiguous — cannot be resolved without context… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/yes-no-multilingual.portuguese_phonetic_lexicon
📚 Portuguese Phonetic Lexicon Dataset
This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects.
🌍 Regional Coverage
The dataset includes words as spoken in ten regional variants:
🇵🇹 Lisbon (Standard and Non-Standard)
🇦🇴 Luanda
🇧🇷 Rio de Janeiro (Standard and Non-Standard)
🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.tigrinya-twi_sentence-pairs
Tigrinya-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Twi_Sentence-Pairs
Number of Rows: 142353
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-twi_sentence-pairs.clinical-5node-exp-buf-lag-tight-tox-casc-v0.1
What this repo does
This dataset tests whether a model can predict when exposure accumulation and governance delay push a clinical situation past the failure horizon into a toxicity cascade, using a four variable coupling pattern.
Core quad
expbuflagtight
Prediction target
label_horizon_breach
Row structure
One row represents a short case vignette with numeric signals for exposure pressure, remaining buffer, response lag, and coupling tightness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-5node-exp-buf-lag-tight-tox-casc-v0.1.clinical-5node-infl-buf-lag-tight-decomp-v0.1
What this repo does
This dataset tests whether a model can detect when rising inflammatory load, weakening buffer, governance lag, and coupling tightness cross the five-node cascade threshold into inflammatory decompensation.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node represents the nonlinear transition from recoverable drift to systemic cascade.
Core quad
inflbuflagtight
Prediction… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-5node-infl-buf-lag-tight-decomp-v0.1.chronologia-benchmark
Chronologia temporal-extraction benchmark
1,049 natural-language temporal expressions with hand-derived gold dates, across
24 languages, scored against three parsers. Every gold value was derived by a
human or by independent date arithmetic — never by any engine under test.
Each row is one utterance a person might actually say ("3 semanas dimpués de o
15 de chinero de 2019", "the week of july 20 2026", "עד יום שישי") with:
column
meaning
lang
BCP-47 primary language… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/chronologia-benchmark.tigrinya-wolof_sentence-pairs
Tigrinya-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Wolof_Sentence-Pairs
Number of Rows: 35192
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-wolof_sentence-pairs.tigrinya-tumbuka_sentence-pairs
Tigrinya-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tumbuka_Sentence-Pairs
Number of Rows: 152916… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tumbuka_sentence-pairs.somali-tigrinya_sentence-pairs
Somali-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tigrinya_Sentence-Pairs
Number of Rows: 169620
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tigrinya_sentence-pairs.packages_python_filtered
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
packages_python_filtered
This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis.
Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.kinyarwanda-tigrinya_sentence-pairs
Kinyarwanda-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Tigrinya_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-tigrinya_sentence-pairs.arabic-mantoq-synthetic-g2ptigrinya-tswana_sentence-pairs
Tigrinya-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tswana_Sentence-Pairs
Number of Rows: 154981
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tswana_sentence-pairs.afrikaans-tigrinya_sentence-pairs
Afrikaans-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Tigrinya_Sentence-Pairs
Number of Rows: 454333… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tigrinya_sentence-pairs.atc-role-classification-synthetic-2procedurally generated synthetic data, script is included in files
sentence-types
COMMAND: This label refers to statements that give instructions or orders. These can be further divided into sub-labels ACTION and DENIAL.
ACTION: This sub-label refers to commands that instruct the listener to perform an action.
DENIAL: This sub-label refers to commands that instruct the listener to refrain from performing an action.
QUESTION: This label refers to statements that ask for information. These can be further divided into sub-labels QUERY, YESNO and REQUEST.
QUERY: This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/sentence-types.Trading
Financial Sentiment Dataset
This dataset contains financial news articles and social media posts labeled with sentiment scores. It is designed to train and evaluate models for sentiment analysis in the context of financial markets.
Dataset Details
Data Sources: Financial news websites, social media platforms (e.g., Twitter, Reddit)
Labels: Positive, Negative, Neutral
Number of Samples: 50,000
Languages: English
Dataset Structure
The dataset is structured as… See the full description on the dataset page: https://huggingface.co/datasets/TigerTrading/Trading.tigrinya-tsonga_sentence-pairs
Tigrinya-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tsonga_Sentence-Pairs
Number of Rows: 157713
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tsonga_sentence-pairs.lingala-tigrinya_sentence-pairs
Lingala-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Tigrinya_Sentence-Pairs
Number of Rows: 97834… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-tigrinya_sentence-pairs.
