datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk_dataset_no_unknownCrisisTS
CrisisTS Dataset
CrisisTS Description
CrisisTS is a multimodal multilingual dataset containing textual data from social media and meteorological data for crisis managmement.
Dataset Summary
Languages: 2 Languages (English and French)
Total number of tweets: 22,291 (15,368 in French and 6,923 in English) (French textual data will be released soon)
Total number of French meteorological data: 46,495 (3 hours frequency)
Total number of English meteorological data:… See the full description on the dataset page: https://huggingface.co/datasets/Unknees/CrisisTS.ruv_tv_unknown_speakersDataset copied from http://hdl.handle.net/20.500.12537/191 by Reykjavik University.
Information can be found at that link.
RUV TV unknown speakers
About the RUV TV unknown speakers corpus
The RUV TV unknown speakers corpus is 281 hours of TV data from six RÚV TV
shows. The data continas 221,759 utterrances from various unlabelled speakers.
The text is normalized. The data is aligned and segmented, ready for ASR
training. Audio conditions vary between recordings. This data set is… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ruv_tv_unknown_speakers.language-identificationpdm_carlainto-the-unknownhotpot_qa_unknownunkebench-hpse
UnKEBench-HPSE
This repository contains the UnKEBench evaluation data used in
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing.
It extends the 1,000 records in UnKEBench with an untargeted editing prompt for each passage.
The project repository, including lightweight evaluation helpers, is available at
lliutianc/hpse.
Usage
from datasets import load_dataset
dataset = load_dataset("lliutianc/unkebench-hpse", split="test")
print(dataset[0])… See the full description on the dataset page: https://huggingface.co/datasets/lliutianc/unkebench-hpse.qor-af-soomaali
Qor Af-Soomaali — the Unkad Somali Corpus (v0.4.0)
Community-contributed, peer-validated, linguist-verified Somali text,
built on qor.unkad.com by Unkad Labs, an independent
Somali AI research lab.
Every item was written by a consenting Somali speaker, validated by at least two community
members, and signed off by a trusted linguist reviewer. Every item
carries provenance: mode, register, sector, and (where shared) the variety the contributor speaks.
This release… See the full description on the dataset page: https://huggingface.co/datasets/unkadlabs/qor-af-soomaali.vctk_dataset_no_unknown_splittedBrats2021_slicedUnKEBenchTDevils
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Unknown9273/TDevils.magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details
Dataset Card for Evaluation run of magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
Dataset automatically created during the evaluation run of model magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details.clinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection
PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation.
You receive:
patient_summary
workup_summary
current_plan
You decide:
frontier_caseyes or no
reason_typemust match the allowed list
next_stepone sentence
Allowed reason_type values
no_frontier
rare_disease_suspected
conflicting_evidence
refractory_to_standard
atypical_multisystem
novel_adverse_event
unexplained_biomarker_pattern
unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ALCUNA_meta_affirmative_known_unknownPhi3_intent_v37_2_wo_unknownPhi3_intent_v46_2_w_unknown_upper_lowerPhi3_intent_v68_2_w_unknown_upper_loweruncertainty-incompleteness-functional-unknowns-genomics-v01
Dataset
ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01
This dataset tests one capability.
Can a model resist inventing biological function when evidence is incomplete.
Core rule
Genomics contains large unknowns.
A claim must respect
incomplete annotation
context specific regulation
limits of prediction
absence of functional validation
Prediction is not proof.
Annotation is not mechanism.
Expression is not causation.
Canonical labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01.Phi3_intent_v45_1_w_unknown_upper_lowerGV_Train_100h_UnknownGenderPhi3_intent_v50_2_w_unknown_upper_lowerKnown-to-Unknown-SQuADv2Phi3_intent_v46_1_w_unknownPhi3_intent_v57_2_w_unknown_upper_lowerafrica-worldbank-repeaters-in-grade-unknown-of-lower-secondary-general-education-male-number-uis
Repeaters in grade unknown of lower secondary general education, male (number) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-repeaters-in-grade-unknown-of-lower-secondary-general-education-male-number-uis.buzz_sources_125_unknownPhi3_intent_v45_3_w_unknown
