datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tdtu_vietnamese_hsd_finaltdtu-student-regulations-qa
TDTU Vietnamese University Regulations QA Dataset
Dataset Description
Tập dữ liệu hỏi-đáp tiếng Việt về quy chế, quy định sinh viên của Trường Đại học Tôn Đức Thắng (TDTU), được xây dựng cho bài toán Retrieval-Augmented Generation (RAG) và fine-tuning LLM tư vấn sinh viên.
Ngôn ngữ: Tiếng Việt
Domain: Quy chế đại học, chính sách sinh viên
Mục đích: Huấn luyện chatbot tư vấn sinh viên TDTU
Dataset Details
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/hungminhss/tdtu-student-regulations-qa.tdtu_voice_dataset
Dataset Card for "tdtu_voice_dataset"
More Information needed
tdtu_vqa_dataset_herb
TDTU VQA Dataset — Vietnamese Medicinal Herbs 🌿
Dataset Description
TDTU VQA Dataset Herb is a Vietnamese Visual Question Answering (VQA) dataset focused on medicinal plants and herbs. It was developed for scientific research at Ton Duc Thang University (TDTU), with the goal of advancing AI models capable of recognizing and answering questions about Vietnamese medicinal herbs.
Homepage: Hugging Face Dataset
Repository: azan100an/tdtu_vqa_dataset_herb
Point of Contact:… See the full description on the dataset page: https://huggingface.co/datasets/azan100an/tdtu_vqa_dataset_herb.H-GEOThe problem texts of the dataset are included in FTtrain.json and FTtest.json.
H-GEO
H-GEO is a dataset of images and Chinese text featuring geometry problems, LLM errors, and their corrections. It supports fine-tuning to reduce hallucinations in geometric reasoning.
JSON structure
{
"messages": [
{
"content": "这是一个数学问题和它的一个错误答案。问题:如图,A、D是⊙O上的两个点,BC是直径,若∠OAC=55°,则∠D的度数是()… See the full description on the dataset page: https://huggingface.co/datasets/TDT000/H-GEO.CRAWL.TDT.mini.dai-hoc
TDT crawled website data (mini version)
This dataset is forked from (BroDeadlines/CRAWL.TDT.dai-hoc)
parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.QA.TDT.FQA_tu_van_hoc_duongTEST.PART_SUMMERIZE.UEH.raptor.edu_tdt_dataTEST.PART_SUMMERIZE.raptor.edu_tdt_data
Collection
TEST.basic_tdt_raptor
{
vec_idx: "vec-raptor-basic_index_tdt_clean"
text_idx: "text-raptor-basic_index_tdt_clean"
}
eda_tdtu_vietnamese_hsdtdtu_pdf_dataTEST.TDT.edu_tdt_data
Preprocess text
A dataset contains pre-processed text corpora that can be use as a knowledge source.
test-basic_test_tdt_dataset
{
"vector_index": "test-basic_test_tdt_dataset",
"method": "window_slide",
"step": 50,
"chunk_size": 1500
}
index.medium_index_tdt
{
"vector_index": "vec-index.medium_index_tdt",
"text_index": "text-index.medium_index_tdt",
"method": "window_slide",
"step": 50,
"chunk_size": 1500,
"time(min)": "5.35"
}
TEST.PART_CLUSTER.raptor.edu_tdt_dataTEST.TDT.large.tdt_copora_data
Notes
This dataset have not been run on proposition generation
TEST.TDT.mini.tdt_copora_dataCRAWL.TDT.admission.tdtu.edu.vn_dai-hoc
The crawled dataset from TDT website
This dataset is a filter of "admission.tdtu.edu.vn_dai-hoc" from a collection of web pages.
tdtu-vietnamese-hsdTEST.basic_test_tdt_datasetTEST.NEW.PART_SUMMERIZE.raptor.edu_tdt_data
Data
TEST.medium_tdt_raptor
{
text_idx: "text-raptor-medium_index_tdt",
vec_idx: "vec-raptor-medium_index_tdt",
size: 1935
}
TD_TVPLtdtu-newsTEST.TDT.mini.edu_tdt_proposition_dataTdtu_evaleval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
37.59%
20.64%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920.eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
11.34%
3.63%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926.tdtu-voz-unlabeledtdtu-hsd-augCRAWL.TDT.dai-hoc
TDT crawl website data
Crawled on 2023, this dataset has the following properties:
Filter out "news" information
Have information about university clubs, lecturers, other department websites,...
TEST.NEW.PART_CLUSTER.raptor.edu_tdt_data
