datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alibaba-ctf
Alibaba CTF Benchmark
Alibaba CTF Benchmark is a CTF benchmark designed to measure the frontier of agent work on Capture The Flag security challenges. It consists of 87 high-quality tasks curated from the 2023–2026 AlibabaCTF (formerly AliyunCTF) competition series, covering five core categories: Web (25), Pwn (19), Misc (14), Reverse (16), and Crypto (13). During the curation process, LLM-based challenges were excluded due to their additional credential requirements and test… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/alibaba-ctf.StreamGuardBench
Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
💻 GitHub
💡 Dataset Overview
StreamGuardBench is the first benchmark specifically designed for evaluating streaming guardrails. StreamGuardBench prompts ten widely used open-source LMs—comprising five text-only and five vision-language models—and annotates every generated response with harm labels, therefore enabling accurate measurement of streaming guardrail effectiveness in… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/StreamGuardBench.ddxplus
Dataset Description
We are releasing under the CC-BY licence a new large-scale dataset for Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the medical domain. The dataset contains patients synthesized using a proprietary medical knowledge base and a commercial rule-based AD system. Patients in the dataset are characterized by their socio-demographic data, a pathology they are suffering from, a set of symptoms and antecedents related to this pathology, and a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/ddxplus.telco-customer-churn
Dataset Card for Telco Customer Churn
This dataset contains information about customers of a fictional telecommunications company, including demographic information, services subscribed to, location details, and churn behavior. This merged dataset combines the information from the original Telco Customer Churn dataset with additional details.
Dataset Details
Dataset Description
This merged Telco Customer Churn dataset provides a comprehensive view of customer… See the full description on the dataset page: https://huggingface.co/datasets/aai510-group1/telco-customer-churn.pmdata
PMData Dataset
About Dataset
Paper: https://dl.acm.org/doi/10.1145/3339825.3394926
In this dataset, we present the PMData dataset that aims to combine traditional lifelogging with sports activity logging. Such a dataset enables the development of several interesting analysis applications, e.g., where additional sports data can be used to predict and analyze everyday developments like a person's weight and sleep patterns, and where traditional lifelog data can be used in a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/pmdata.ddxplus-french
Dataset Description
We are releasing under the CC-BY licence a new large-scale dataset for Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the medical domain. The dataset contains patients synthesized using a proprietary medical knowledge base and a commercial rule-based AD system. Patients in the dataset are characterized by their socio-demographic data, a pathology they are suffering from, a set of symptoms and antecedents related to this pathology, and a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/ddxplus-french.XGuard-Train-Open-200K
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
🤗 HuggingFace |
🤖 ModelScope |
📄 Paper
🐬 Introduction
XGuard-Train-Open-200K is an open-source subset of the training corpus developed for the YuFeng-XGuard-Reason guardrail model series. YuFeng-XGuard-Reason is engineered to accurately identify security risks in user requests, model responses, and general text, while… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/XGuard-Train-Open-200K.so101_full_DHThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 77598,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full_DH.details_tushar310__Hippy-AAI-7B
Dataset Card for Evaluation run of tushar310/Hippy-AAI-7B
Dataset automatically created during the evaluation run of model tushar310/Hippy-AAI-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tushar310__Hippy-AAI-7B.aaiso101_full_SHThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 63771,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full_SH.so101_2machine_full_SH_DHThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 110987,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_2machine_full_SH_DH.so101_full_JSThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 31,
"total_frames": 63161,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full_JS.EmoTa
EmoTa: A Tamil Emotional Speech Dataset
EmoTa is the first emotional speech dataset in Tamil, designed to reflect the
linguistic diversity of Sri Lankan Tamil speakers. It contains 936 recorded
utterances from 22 native Tamil speakers (11 male, 11 female), each articulating
19 semantically neutral sentences across five emotions: anger, happiness,
sadness, fear, and neutral.
🌐 Project page: https://aaivu.github.io/EmoTa/
💻 Code & loader: https://github.com/aaivu/EmoTa
📄 Paper… See the full description on the dataset page: https://huggingface.co/datasets/aaivu-labs/EmoTa.so101_coffee_updateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 2140,
"total_frames": 463261,
"total_tasks": 8,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2140"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_update.so101_recovery_2task
so101_recovery_2task — SO-101 cup repositioning as two tasks, 749 episodes
Cup start positions for the 394 pick episodes (left) and start → placed transport for 350 of the
355 place episodes (right) — one color per operator, anonymized.
Front camera (fixed, elevated view), 8× speed — pick episodes followed by place episodes.
749 teleoperated demonstrations on the SO-101 arm, collected by 8 operators and recorded as
two separately instructed tasks rather than one continuous… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_recovery_2task.so101_coffee_all
so101_coffee_all
2,564 teleoperated SO-101 demonstrations covering a coffee-making routine as 12 separately instructed steps, each recorded with two cup variants — a red-banded paper cup and a white one.
Steps 1–10 are the two-cup sequence (2,140 episodes, 107 per step-and-colour); steps 11 and 12 pick the cup up near the blue circle and place it back on it (424 episodes, 106 per step-and-colour).
This dataset was created using LeRobot.
Combines so101_coffee_subtask_3 and… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_all.AAID-clone
Citation Information
Liu, Z., Yang, K., Xie, Q., Zhang, T., & Ananiadou, S. (2024, August). Emollms: A series of emotional large language models and annotation tools
for comprehensive affective analysis. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5487-5496).
CHIRP
CHIRP Benchmark: An Open-Ended Free form Multimodal Benchmark
CHIRP is a new multimodal evaluation benchmark with 104 open ended questions. Free form questions consists of questions where the model has to generate a response that is more open-ended, creative and does not have a "correct" answer. We include 8 distinct categories of questions. Each category requires understanding the image, and presents the opportunity for analysis and a thorough response.
The categories are:… See the full description on the dataset page: https://huggingface.co/datasets/cerc-aai/CHIRP.heart-failure-prediction-dataset
language:
en
license: odbl
tags:
health
heart-disease
medical
machine-learning
annotations_creators:
expert-generated
language_creators:
expert-generated
pretty_name: Heart Failure Prediction Dataset
size_categories:
1K<n<10K
source_datasets:
original
task_categories:
structured-data-classification
task_ids:
binary-classification
health-data-analysis
paperswithcode_id: heart-failure-prediction
configs:
default
dataset_info:
features:
- name: Age
dtype: int32
- name: Sex… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/heart-failure-prediction-dataset.so101_full_DYThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 56729,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full_DY.so101_coffee_recovery_2
so101_coffee_recovery_2
424 teleoperated SO-101 demonstrations of two steps from a coffee-making routine, 212 episodes each:
"pick up the cup near the blue circle""place the cup on the blue circle"
Both steps were recorded with two cup variants — a red-banded paper cup and a white one — 106 episodes per step-and-colour.
This dataset was created using LeRobot.
Overview
Episodes / frames
424 / 86,049 (47.8 min)
Composition
2 steps × 2 colours… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_recovery_2.so101_coffee_all_blue
so101_coffee_all
2,564 teleoperated SO-101 demonstrations covering a coffee-making routine as 12 separately instructed steps, each recorded with two cup variants — a red-banded paper cup and a white one.
Steps 1–10 are the two-cup sequence (2,140 episodes, 107 per step-and-colour); steps 11 and 12 pick the cup up near the blue circle and place it back on it (424 episodes, 106 per step-and-colour).
This dataset was created using LeRobot.
Combines so101_coffee_subtask_3 and… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_all_blue.so101_coffee_all_yellow_floor
so101_coffee_all
2,564 teleoperated SO-101 demonstrations covering a coffee-making routine as 12 separately instructed steps, each recorded with two cup variants — a red-banded paper cup and a white one.
Steps 1–10 are the two-cup sequence (2,140 episodes, 107 per step-and-colour); steps 11 and 12 pick the cup up near the blue circle and place it back on it (424 episodes, 106 per step-and-colour).
This dataset was created using LeRobot.
Combines so101_coffee_subtask_3 and… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_all_yellow_floor.AAI_assignment2diabetes-readmissionPort of the diabetes-readmission dataset from UCI (link here). See details there and use carefully.
Basic preprocessing done by the imodels team in this notebook.
The target is the binary outcome readmitted.
Sample usage
Load the data:
from datasets import load_dataset
dataset = load_dataset("imodels/diabetes-readmission")
df = pd.DataFrame(dataset['train'])
X = df.drop(columns=['readmitted'])
y = df['readmitted'].values
Fit a model:
import imodels
import numpy as np
m =… See the full description on the dataset page: https://huggingface.co/datasets/aai540-group3/diabetes-readmission.so101_full_TKThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 74491,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full_TK.so101_coffee_subtask_3
so101_coffee_subtask_3
2,140 teleoperated SO-101 demonstrations of a coffee-making routine, recorded as
10 separately instructed steps. Every step was recorded with two cup variants — a
red-banded paper cup and a white one — giving 20 (step × cup) combinations of
107 episodes each.
This dataset was created using LeRobot.
Overview
Episodes / frames
2,140 / 463,261 (4.29 h)
Composition
10 steps × 2 cup colours × 107 episodes
Episode length
mean… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_subtask_3.so101_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 243,
"total_frames": 537800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:243"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_full.so101_coffee_subtask_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 486,
"total_frames": 131102,
"total_tasks": 8,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:486"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aailabkaist/so101_coffee_subtask_2.
