datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.shitto-mania-dic
English Summary
This dataset accompanies our NLP2026 study on language resource design in the RAG era.
Using a 321-episode Japanese dataset derived from the essay series Shitto Mania, we found that structured metadata substantially outperformed full text in the tested reference-retrieval setting.
Key finding: structured metadata alone achieves 11.1× better retrieval performance than full text in TF-IDF retrieval (59.0% vs 5.3% Recall@10, STRUCT queries, n=300). This advantage… See the full description on the dataset page: https://huggingface.co/datasets/samuraijun/shitto-mania-dic.64-que-kinh-dich
64 quẻ Kinh Dịch
The 64 hexagrams of the I Ching
1. Mô tả · Description
Đủ 64 quẻ theo thứ tự Chu Dịch, kèm tên Việt, tên Hán, quẻ thượng, quẻ hạ và tượng quẻ.
All 64 hexagrams in King Wen order, with Vietnamese and Han names, upper and lower trigrams, and the image.
Số dòng · Rows: 64
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/64-que-kinh-dich.sam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.english-khmer-dictionary
📖 English–Khmer Dictionary Dataset
A comprehensive bilingual English–Khmer (ភាសាខ្មែរ) dictionary dataset in CSV format containing 170,000+ entries. Each entry includes the original English word, its Khmer translation, part of speech, full definitions in both languages, and example sentences — making it one of the richer English–Khmer lexical resources available for NLP and language learning.
Dataset Description
This dataset provides structured dictionary entries pairing… See the full description on the dataset page: https://huggingface.co/datasets/mrrtmob/english-khmer-dictionary.eval-gliner2-ner-fin-dice-soft-20260709TCGA-SARC-dict-tumorlanguages_datasetThis dataset contains a set of 8612 languages from across the world as well as data such as Glottocode, ISO-639-3 codes, names, language families etc.
Original source: https://glottolog.org/glottolog/language
scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.Amazigh-English-DictionaryscoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.Corpus_DiCo
[!NOTE]
Dataset origin: http://redac.univ-tlse2.fr/lexiques/dico.html
Description
DiCo est un corpus qui contient de nombreuses informations sur les dictionnaires, telles que des listes d'articles nouveaux, d'articles sortis ou des statistiques. Ces informations ont été obtenues en comparant intégralement et manuellement les éditions successives d'un même dictionnaire, par exemple le Petit Larousse 2005 avec le Petit Larousse 2006, puis ce dernier avec le 2007, etc. La méthode de… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/Corpus_DiCo.balochi-dictionary
Zaanth Balochi Dictionary — v0.1 (Latin script)
The first release of the open Balochi dictionary by Zaanth, an open
platform for Balochi & Brahui language data.
30 entries of the most frequent Balochi words (Latin script / Syáhag),
each with an English meaning verified by a native Makrani Balochi speaker,
corpus frequency, and a real example sentence with translation.
Method
Words ranked by frequency across 18,930 Balochi sentences; candidate meanings
were… See the full description on the dataset page: https://huggingface.co/datasets/Zaanthai/balochi-dictionary.Dickens_Word_Frequency_Data
