datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk_dataset_no_unknownCrisisTS
CrisisTS Dataset
CrisisTS Description
CrisisTS is a multimodal multilingual dataset containing textual data from social media and meteorological data for crisis managmement.
Dataset Summary
Languages: 2 Languages (English and French)
Total number of tweets: 22,291 (15,368 in French and 6,923 in English) (French textual data will be released soon)
Total number of French meteorological data: 46,495 (3 hours frequency)
Total number of English meteorological data:… See the full description on the dataset page: https://huggingface.co/datasets/Unknees/CrisisTS.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 22590,
"total_tasks":1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/UN-kk/so100_test.survivalanalysis_checkpointingruv_tv_unknown_speakersDataset copied from http://hdl.handle.net/20.500.12537/191 by Reykjavik University.
Information can be found at that link.
RUV TV unknown speakers
About the RUV TV unknown speakers corpus
The RUV TV unknown speakers corpus is 281 hours of TV data from six RÚV TV
shows. The data continas 221,759 utterrances from various unlabelled speakers.
The text is normalized. The data is aligned and segmented, ready for ASR
training. Audio conditions vary between recordings. This data set is… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ruv_tv_unknown_speakers.wikiart-artist-unknown-artistlanguage-identificationhotpot_qa_unknownpdm_carlainto-the-unknownsimply-datasetsSDXL_REGULARIZATION_IMAGESSDXL_REGULARIZATION_IMAGES
Dataset v1
Prompt: Beautiful girl
Negative Prompt: child
Resolution: (1024, 1024)
Base Model: sd_xl_base_1.0_0.9vae.safetensors, Refiner Model: sd_xl_refiner_1.0_0.9vae.safetensors
LoRA [sd_xl_offset_example-lora_1.0.safetensors] weight: 0.5
More Datasets will be added in future, Show your support by clicking like
unknown_oldA dataset of interesting contents.
unkebench-hpse
UnKEBench-HPSE
This repository contains the UnKEBench evaluation data used in
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing.
It extends the 1,000 records in UnKEBench with an untargeted editing prompt for each passage.
The project repository, including lightweight evaluation helpers, is available at
lliutianc/hpse.
Usage
from datasets import load_dataset
dataset = load_dataset("lliutianc/unkebench-hpse", split="test")
print(dataset[0])… See the full description on the dataset page: https://huggingface.co/datasets/lliutianc/unkebench-hpse.qor-af-soomaali
Qor Af-Soomaali — the Unkad Somali Corpus (v0.4.0)
Community-contributed, peer-validated, linguist-verified Somali text,
built on qor.unkad.com by Unkad Labs, an independent
Somali AI research lab.
Every item was written by a consenting Somali speaker, validated by at least two community
members, and signed off by a trusted linguist reviewer. Every item
carries provenance: mode, register, sector, and (where shared) the variety the contributor speaks.
This release… See the full description on the dataset page: https://huggingface.co/datasets/unkadlabs/qor-af-soomaali.vctk_dataset_no_unknown_splittedBrats2021_slicedUnKEBenchCHaMEL_v1.0
CHaMEL v1.0 Dataset
Controllable Harness for Model adaptability Evaluation under Latent rule-shifts
Dataset Description
CHaMEL is a benchmark harness for evaluating the adaptive reasoning capabilities of Large Language Models (LLMs) through dynamic rule-shift tasks. This dataset contains the pre-generated task stimuli and human evaluation baselines used in the CHaMEL paper.
CHaMEL implements a factorial experiment design where researchers can independently control the… See the full description on the dataset page: https://huggingface.co/datasets/unknown202612/CHaMEL_v1.0.combined-unknown-pneumonia-and-tuberculosisTDevils
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Unknown9273/TDevils.Douyin
Douyin Keyword Video Dataset
This dataset contains Douyin short videos collected from keyword search results.
The current upload is organized into batched .tar archives, with about 50 videos per batch.
The videos cover diverse keywords and scenarios, including:
vlog
cats
dogs
dance
football
basketball
cars
beauty
camera movement
daily life
food exploration
and more
Most videos are downloaded in 1080P resolution at 30fps. The dataset will continue to be updated
with more… See the full description on the dataset page: https://huggingface.co/datasets/Unknown0911xinyue/Douyin.industrialintelligenceb2d_cornerclinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection
PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation.
You receive:
patient_summary
workup_summary
current_plan
You decide:
frontier_caseyes or no
reason_typemust match the allowed list
next_stepone sentence
Allowed reason_type values
no_frontier
rare_disease_suspected
conflicting_evidence
refractory_to_standard
atypical_multisystem
novel_adverse_event
unexplained_biomarker_pattern
unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.oliy-maqsad-booksmagnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details
Dataset Card for Evaluation run of magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
Dataset automatically created during the evaluation run of model magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
qor-af-soomaali-frequency
Qor Somali Frequency List (from corpus v0.2.1)
Word and bigram frequencies computed from
unkadlabs/qor-af-soomaali
v0.2.1: 2,282 sentences of Somali, each written by a consenting, credited
Somali speaker, accepted by two independent community validators, and
signed off by a linguist reviewer on qor.unkad.com.
To our knowledge this is the first Somali frequency list in which the
origin of every counted occurrence is known. Frequency lists derived
from web crawls inherit whatever… See the full description on the dataset page: https://huggingface.co/datasets/unkadlabs/qor-af-soomaali-frequency.Phi3_intent_v37_2_wo_unknown
