datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hi-en-noisy-vad-benchmark
Hindi-English Noisy VAD Benchmark
Version 0.1.0 is a deterministic, evaluation-only benchmark with 78
mono PCM16 WAV files at 16 kHz: six clean speech controls and 72 mixtures spanning
six speech sources, three real noise categories, and four SNRs (20, 10, 5, 0 dB).
Intended use
Use this dataset to compare voice-activity detectors under matched Hindi/English
noise conditions and to tune thresholds. It is too small and insufficiently diverse
for model training… See the full description on the dataset page: https://huggingface.co/datasets/Aakash22134/hi-en-noisy-vad-benchmark.BPCC-en-hi-300-cleanedcross-rag-enhiVietnam-History-200K-ENEnglish 200,000-sample Vietnamese history dataset in the same fine-tuning format (with ≈78% reasoning and ≈22% final-only).
Format & coverage
Language: English
Scope: 905–2025 (events, figures, dynasties, wars, reforms, culture, documents)
Structure: messages (ShareGPT/ChatML style)
With reasoning (≈78%): system → user → assistant (analysis) → assistant (final)
Final-only (≈22%): system → user → assistant (final)
Learn more on GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-200K-EN.Vietnam-History-500K-EnEnglish 500,000-sample Vietnamese history dataset, with ≈78% chain-of-thought (analysis) and ≈22% final-only answers.
Format & coverage
Scope: 905–2025 (events, figures, dynasties, wars, reforms, culture, documents)
Structure: ShareGPT/ChatML-style messages
With reasoning (≈78%): system → user → assistant (analysis) → assistant (final)
Final-only (≈22%): system → user → assistant (final)
Learn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
VietNam-History-100K_ENLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
Vietnam-History-1M-EnLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
Vietnamese History Q&A – English – 1,000,000 samples
Size: 1M conversations (JSONL, gzip)Language: EnglishDomain: Vietnamese history 905–2025 (events, figures, dynasties, wars, reforms, documents)Format: ShareGPT/ChatML-style messages with assistant channels analysis (reasoning) and final (answer).
Record
{
"messages": [
{"role":"system","content":"…"},
{"role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-1M-En.En_Hi_Sumqa_en_hiMedsiML_EN_HIN_Data
MedSiML
Dataset Description
This dataset contains simplified English and simplified Hindi sentence pairs derived from the MedSiML dataset. The data has been filtered using the Cynical Data Selection algorithm to retain examples that are most representative of the target distribution while reducing redundancy.
Source Dataset
This dataset is derived from the MedSiML dataset:
MedSiML: A Multilingual Approach for Simplifying Medical Texts… See the full description on the dataset page: https://huggingface.co/datasets/vishnu-vizz/MedsiML_EN_HIN_Data.challenge_enHindieval-results-Nayana-cognitivelab_exp-colpali-merged-hi-en-20k-vidore_arxivqa_test_subsampledscience_en_hillama_en_hieval-results-Nayana-cognitivelab_exp-colpali-merged-hi-en-10k-vidore_docvqa_test_subsampledCulinary_Terms_Translation_en-hieval-results-Nayana-cognitivelab_exp-colpali-merged-hi-en-10k-vidore_infovqa_test_subsampledeval-results-Nayana-cognitivelab_exp-colpali-merged-hi-en-20k-vidore_docvqa_test_subsampledeval-results-Nayana-cognitivelab_exp-colpali-merged-hi-en-20k-vidore_infovqa_test_subsampled
