datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nagi_no_Asukara_Videos_Captioned
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
molt-benchmark-results
Molt · elastic on-device inference measurements
Everything measured while building Molt, a
runtime that moves a running generation onto a smaller model between two
tokens, carrying the KV cache across, so an on-device LLM under memory pressure
is neither reclaimed by the OS nor restarted from the prompt.
Published so the claims can be checked rather than taken on trust. The figures in
the repo README and the results page are generated from these files; nothing is
transcribed by… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/molt-benchmark-results.nagamese-english-mt
Nagamese-English Machine Translation Corpus
Dataset Summary
This dataset contains 3,340 parallel sentence pairs in English and
Nagamese (Naga Pidgin), intended for training and evaluating machine
translation systems between the two languages. It has already been used to
fine-tune at least one NLLB-200-based translation model
(agnivamaiti/nllb-200-en-nagamese).
This is the first dataset card written for this dataset — no card
existed on the repository prior to this… See the full description on the dataset page: https://huggingface.co/datasets/agnivamaiti/nagamese-english-mt.nagri-language-dataset
Nagri Language Dataset - Description
Nagri is a writing system for the Sylheti language, which is spoken in Bangladesh and India.
The Nagri Language Dataset is a collection of text and image data specifically focused on Syloti Nagri script. This dataset is designed for OCR (Optical Character Recognition), handwriting recognition, and language modeling tasks related to the Syloti Nagri language.
Dataset Features:
✅ Text Samples: Contains a variety of words, phrases… See the full description on the dataset page: https://huggingface.co/datasets/mahdi-hasan-shuvo/nagri-language-dataset.nagisa_stopwords
Japanese Stopwords for nagisa
This dataset is the Japanese stopwords list built into nagisa (v0.2.12+). It is published here on Hugging Face for easy access and reproducibility.
Overview
Language
Japanese
Size
147 words
Source
CC-100, Wikipedia
License
MIT
Dataset Description
This dataset contains 147 frequently used Japanese words extracted from large-scale corpora. Each word is annotated with its part-of-speech (POS) tag according… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/nagisa_stopwords.osworld_tasks_filesprompt-variations
Prompt Variations and LLM Responses
Prompt variants and model responses used to evaluate the
Stability-Generalization Score (SGS) across eleven LLMs (eight
open-source + three closed-source) on six QA / instruction benchmarks
under six families of stylistic perturbations.
Splits
split
rows
source dataset
truthful_qa
99,888
TruthfulQA
natural_questions
41,040
Natural Questions
alpaca
13,872
Alpaca
simpleqa_verified
13,872
SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.naga-eng-fullgrasp_bboxWildFire-Ynaga-engnagri-sound-dataset
Sylheti Language Learning – Audio Interaction Dataset
Overview
This dataset contains letter-level reference pronunciation audio samples designed for a multilingual Augmented Reality (AR) language-learning system.
The system adapts its interface language dynamically based on user preference, while primarily aiming to teach and evaluate Sylheti pronunciation.The dataset is structured to support real-time pronunciation feedback, deterministic AR triggers, and multilingual… See the full description on the dataset page: https://huggingface.co/datasets/shivdi1999/nagri-sound-dataset.Agent_finetuning_dataNagpurs-Food-JointsEnglish-NagameseJLPT_Applicants_by_TestSiteI collected the “Applicants & Examinees by Test Site” data for the Japanese Language Proficiency Test (JLPT) from 2010 to 2024 directly from the Japan Foundation’s website.
Used the open-source tool Tabula. With this dataset, there are several interesting areas for exploration as a data analyst or researcher:
📊 Possible Analysis Directions:
Year-wise trends in applicants & examinees across test sites.
Geographical distribution of test-takers to spot high-participation vs underserved areas.… See the full description on the dataset page: https://huggingface.co/datasets/naga2hands/JLPT_Applicants_by_TestSite.lotSizenagamese2englishLanguage_preservation_datatesttest_uxHello world
TourismTourism_GLTourism_Cleanplay-store-revenue-analysis
