datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sekai2_Real_World
Sekai2 Real World
This repository releases the reproducible URL/timestamp metadata and paired
camera-pose/caption annotations for the perspective-video portion of
Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling.
Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper
The perspective MP4 clips are not redistributed here. Each row in
sekai2_clips.csv provides the source URL and the exact half-open frame range
[start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.kansas-layoffs-warn-act-notices-daily
Kansas WARN Act layoff notices — every filing we hold since 1998, one CSV, rebuilt daily
791 Kansas WARN notices — every one this dataset holds, back to 1998 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-05-01
· state source last checked 2026-09-20T12:28Z · official source: Kansas Department of Commerce — WARN notices.
Kansas employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/kansas-layoffs-warn-act-notices-daily.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/kankshith123/FinQA-hallucination-detection.kanari-wildfire-ignitions
kanari — worldwide wildfire ignitions archive
Continuously updated archive of significant wildfire events worldwide (136 countries), produced by
kanari, a free near-real-time map of wildfire ignitions.
Each row is one fire event: satellite hotspots from NASA FIRMS (VIIRS 375 m), NOAA GOES and
EUMETSAT Meteosat MTG are clustered into events (≈4 km cells); the first detection is the proxy for
ignition time. Public witness reports (Bluesky, press via GDELT, Telegram) are geoparsed… See the full description on the dataset page: https://huggingface.co/datasets/expansia/kanari-wildfire-ignitions.expenseskanurimt
KanuriMT: A Low-Resource English–Kanuri Parallel Corpus and Benchmark for Neural Machine Translation
Dataset Description
KanuriMT is the first publicly available cleaned, annotated English–Kanuri parallel corpus designed for Neural Machine Translation (NMT) research. Kanuri is a low-resource Nilo-Saharan language spoken primarily in northeastern Nigeria, Niger, Chad, and Cameroon, with an estimated 4–8 million speakers.
Language pair: English (en) → Kanuri (knc, Latin… See the full description on the dataset page: https://huggingface.co/datasets/IsahMBukar/kanurimt.middle-trust-c95b8b
middle-trust-c95b8b
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/kanakobayashi/middle-trust-c95b8b.english-kannada-cleaned
English–Kannada Cleaned
A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models.
Languages: English -> Kannada
License: Apache License 2.0
Dataset statistics
Train: 8,00,000 sentence pairs
Validation: 1,000 sentence pairs
Test: 1,000 sentence pairs
Total: 5,02,000 sentence pairs
These counts exclude per-file CSV headers.
Source and provenance
The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.KannadaPromptBench
KannadaPromptBench
A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language.
Dataset Summary
Language: Kannada (kn)
Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50)
Total: 225 culturally grounded samples
Inter-annotator agreement: Cohen's κ > 0.80
Dataset Structure
Each sample contains: id, task, input_text, label, difficulty, domain.
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/Anushhh/KannadaPromptBench.KandenAiHackathonAirflow
KandenAiHackathonAirflow — 室内気流CFDシミュレーションデータセット
English Summary
A tabular dataset of 54 CFD (Computational Fluid Dynamics) simulation cases for a 6m × 5m × 2.7m office room. Generated with OpenFOAM's buoyantSimpleFoam solver, the dataset systematically varies air conditioning speed, AC temperature, window state, and ventilation rate to capture indoor airflow, temperature, and CO2 distribution. Designed for training physics-informed neural network (PINN) surrogate models with… See the full description on the dataset page: https://huggingface.co/datasets/SeiyaCM/KandenAiHackathonAirflow.Sap_Kush_Med_Deepfake
Sap_Kush_Med_Deepfake Dataset
Paired medical-image forgery lineages across six modalities. Every lineage is
one source image, one mask, one seed: the arms differ only in what was done
inside the mask, so a comparison between arms isolates the manipulation rather
than an encoding artefact.
3956 lineages, 30385 files, 8.49 GiB.
v2 adds a removal arm grounded in human annotation for three more modalities
(endoscopy, ultrasound, MRI). v1 had removal for CT only.
What the… See the full description on the dataset page: https://huggingface.co/datasets/Kanhaiyya/Sap_Kush_Med_Deepfake.kannada-bedtime-tts
Kannada Bedtime Story TTS Dataset
Training data for fine-tuning IndicF5 on Kannada bedtime story narration.
Source
Base: SPRINGLab/IndicTTS_Kannada (800 clips, ~5h 53m)
Synthetic clips: 800 clips generated for bedtime story domain
Files
kannada_finetune/train.csv — Training metadata (text + audio paths)
kannada_finetune/val.csv — Validation split
kannada_manifest.jsonl — Full manifest with text, audio paths, durations
Usage
Used… See the full description on the dataset page: https://huggingface.co/datasets/sush0401/kannada-bedtime-tts.Kanun-Yonetmelik-Tuzuken-kannadaitalian-recipeskannada_news_classificationkan_academy_q_aWhile scraping 'textbooks' from Khan Academy, also scarped any community posted questions with the first (most voted) answer.
Karum
English-Karakalpak Parallel Corpus (en-kaa)
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/Kanzoet97/Karum.Kannada-Speech-Dataset
🎧 Kannada Speech Dataset
The Kannada Speech Dataset is a high-quality speech audio dataset designed to deliver structured and reliable audio data for AI and machine learning workflows. It includes 90 hours of audio data across 651 files, available in MP3 and WAV formats, with a total size of 220 MB. This well-organized audio dataset provides balanced and representative voice data, with 48% female and 52% male speakers, and an age range spanning from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Kannada-Speech-Dataset.xnli2.0_train_kannadadiffing-stats-qwen3_1_7B-kansas_abortion-L14-Crosscoder-s2-t100-k100-lr1e-04-x32GSM8k_Reward_Datakannada-cultural-dialogue-datasetxnli2.0_kannadahp_conversations_with_hermione_granger_movieKanParadiffing-stats-gemma3_1B-kansas_abortion-L19-k100-lr1e-03-x32-local-shuffling-Crosscodertop-30-active-restaurant-review-whales-in-kansas-us-177463
Top 30% Active Restaurant Review Whales in Kansas, US
Free sample dataset from BeamStation
--High-Volume, Active targets--
This dataset lists the Top 30% Active Restaurant Review Whales in Kansas, US, comprising 1,565 records of high‑traffic dining establishments. Each entry represents a restaurant that ranks among the most active in its local market, based on recent foot traffic and review volume, and is validated by an "Active Heartbeat" signal that confirms ongoing transactional… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/top-30-active-restaurant-review-whales-in-kansas-us-177463.kanpur-weather-datasetkanada_dataset
