datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrisisTS
CrisisTS Dataset
CrisisTS Description
CrisisTS is a multimodal multilingual dataset containing textual data from social media and meteorological data for crisis managmement.
Dataset Summary
Languages: 2 Languages (English and French)
Total number of tweets: 22,291 (15,368 in French and 6,923 in English) (French textual data will be released soon)
Total number of French meteorological data: 46,495 (3 hours frequency)
Total number of English meteorological data:… See the full description on the dataset page: https://huggingface.co/datasets/Unknees/CrisisTS.unkebench-hpse
UnKEBench-HPSE
This repository contains the UnKEBench evaluation data used in
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing.
It extends the 1,000 records in UnKEBench with an untargeted editing prompt for each passage.
The project repository, including lightweight evaluation helpers, is available at
lliutianc/hpse.
Usage
from datasets import load_dataset
dataset = load_dataset("lliutianc/unkebench-hpse", split="test")
print(dataset[0])… See the full description on the dataset page: https://huggingface.co/datasets/lliutianc/unkebench-hpse.qor-af-soomaali
Qor Af-Soomaali — the Unkad Somali Corpus (v0.4.0)
Community-contributed, peer-validated, linguist-verified Somali text,
built on qor.unkad.com by Unkad Labs, an independent
Somali AI research lab.
Every item was written by a consenting Somali speaker, validated by at least two community
members, and signed off by a trusted linguist reviewer. Every item
carries provenance: mode, register, sector, and (where shared) the variety the contributor speaks.
This release… See the full description on the dataset page: https://huggingface.co/datasets/unkadlabs/qor-af-soomaali.magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details
Dataset Card for Evaluation run of magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
Dataset automatically created during the evaluation run of model magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
SE-Bench
dataset structure
The SE-Bench contains the following subsets (configurations). Please note that all data is loaded under the train split key regardless of the subset.
train:
train train set
single_test: single function test set
multiple_test: multiple function test set
Usage
You can load different parts of the benchmark by specifying the configuration name.
Load Train Set
from datasets import load_dataset
dataset = load_dataset("unknown12423124/SE-Bench"… See the full description on the dataset page: https://huggingface.co/datasets/unknown12423124/SE-Bench.adaption-gd-unk-normalized-samples
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-gd_unk_normalized_samples
This dataset contains normalized and semantically enriched samples identified by GD-UNK codes, featuring anomaly labels and timestamps. The content is structured to ensure consistency and readiness for machine learning models, avoiding hallucinations. Each entry includes object data points processed for quality enhancement.
Dataset size
There are 1… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-gd-unk-normalized-samples.unkassistant-axis-unknown-self-chattrue-false-unknowndemo_unknown-multiturnendpoint-risk-data
