datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twngrams
Taiwanese Mandarin web n-grams
Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the
Taiwan slice of a large web crawl after variety filtering by
twfilter 0.1.0 with the published
twfilter-tables:
every sentence behind these counts passed the 教育部 character-inventory gate, the
simplified-character round-trip, the mainland-orthography, mainland-lexicon,
written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of
Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.korean-parallel-corporaWikipedia-Corpora-Report
Dataset Card for "Wikipedia-Corpora-Report"
This dataset is used as a metadata database for the online WIKIPEDIA CORPORA META REPORT dashboard that illustrates how humans and bots generate or edit Wikipedia editions and provides metrics for “pages” and “edits” for all Wikipedia editions (320 languages). The “pages” metric counts articles and non-articles, while the “edits” metric tallies edits on articles and non-articles, all categorized by contributor type: humans or bots. The… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Wikipedia-Corpora-Report.SA-Parallel-Corpora
SA-Parallel-Corpora
Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi
bitext, drawn from South African government publications.
Produced for the doctoral thesis Injecting Commonsense Knowledge into
Pretrained Language Models for Low Resource Languages (University of Cape Town,
2026). Code at https://github.com/sello-ralethe/SA-knowledge
Structure
One configuration per language pair, each with train, validation
and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.JIC-VQA
JIC-VQA
Dataset Description
Japanese Image Classification Visual Question Answering (JIC-VQA) is a benchmark for evaluating Japanese Vision-Language Models (VLMs). We built this benchmark based on the recruit-jp/japanese-image-classification-evaluation-dataset by adding questions to each sample. All questions are multiple-choice, each with four options. We select options that closely relate to their respective labels in order to increase the task's difficulty.
The… See the full description on the dataset page: https://huggingface.co/datasets/line-corporation/JIC-VQA.crh-parallel-corpora-document-level-noisykomi-russian-parallel-corpora
Source Datasets
1 - news from the website of the Komi administration (https://rkomi.ru/)
2 - Komi media library (http://videocorpora.ru/)
3 - Millet porridge by Ivan Toropov (adaptation)
Authors
Shilova Nadezhda
Chernousov Georgy
corporate-crypto-treasuries
Corporate Crypto Treasury Holdings
Every company, government and ETF known to hold Bitcoin, Ethereum or Solana on
its balance sheet, with the primary source document behind each figure.
359 positions across 297 entities in 39 countries. 358 of the 359 carry a
link to the filing or disclosure the number came from.
Maintained by CorpStacking.
What makes this different from a price feed
Most crypto datasets are market data. This one is balance-sheet data read out
of… See the full description on the dataset page: https://huggingface.co/datasets/corpstacking/corporate-crypto-treasuries.crh-parallel-corpora
Crimean Tatar-English/Russian/Ukrainian Parallel Corpora
Overview
This repository contains the Crimean Tatar-English/Russian/Ukrainian parallel corpora, a collection of sentences in Crimean Tatar and their corresponding translations in English/Russian/Ukrainian. The dataset is intended for use in natural language processing (NLP) tasks such as machine translation and cross-lingual analysis.
All sentences in Crimean Tatar are written in Cyrillic and/or Latin script. If any… See the full description on the dataset page: https://huggingface.co/datasets/QIRIM/crh-parallel-corpora.Georgian-Parallel-Corpora
English-Georgian Parallel Dataset 🇬🇧🇬🇪
📄 Dataset Overview
The English-Georgian Parallel Dataset is sourced from OPUS, a widely used open collection of parallel corpora. This dataset contains aligned sentence pairs in English and Georgian, extracted from Wikipedia translations.
Corpus Name: Wikimedia
Package: wikimedia.en-ka (Moses format)
Publisher: OPUS (Open Parallel Corpus)
Release: v20230407
Release Date: April 13, 2023
License: CC–BY-SA 4.0
🔗… See the full description on the dataset page: https://huggingface.co/datasets/Arseniy-Sandalov/Georgian-Parallel-Corpora.corporate_esg_risk_analyticsCorporate_Employee_Profilesen_corpora_parliament_processedagewerc_corporate-credit-rating
Corporate Credit Rating
Credit Ratings of Big US Firms and their Financials
Dataset Info
Source: Kaggle
Original Size: 0.29 MB
Kaggle Downloads: 4,884
Files: 1
Files
corporate_rating.csv
Mirrored from Kaggle
metra_western_european_drama_annotated_corporaCorporateMailCategorization
CorporateMailCategorization
tags: mails, classification, business
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CorporateMailCategorization' dataset is a collection of CEO office letters and business-related mails. Each entry in the dataset includes the original email text and a label that classifies the email's content into one of several categories relevant to corporate communication. This dataset can be used to train… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CorporateMailCategorization.CorporateCommMail
CorporateCommMail
tags: Email Classification, Priority, BusinessCommunication
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CorporateCommMail' dataset is a curated collection of business communication emails with a focus on understanding the priority levels assigned to emails based on employee responses. The dataset includes various features such as sender, subject, body, timestamp, receiver, CC (carbon copy recipients)… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CorporateCommMail.corporate-security-and-protection-protocols
