datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
online-retailfailbench-robocasa-v2
FailBench RoboCasa v2 — contact-prediction dataset
Labeled robot-failure trials built on RoboCasa kitchen demos (PandaMobile / Franka).
Each trial injects a hardware failure partway through a teleop/MimicGen demo, then records the
contacts the failure causes during a 1-second settle. The supervised target is a
240×320 force-weighted contact heatmap in the agentview camera — the model learns to predict
where a failure at a given pre-failure configuration will drive the… See the full description on the dataset page: https://huggingface.co/datasets/aaronngx/failbench-robocasa-v2.riddlesenseplusplusruleswapNon-compete fact patterns feature agreements that are a modified sample of the noncompete agreement from Our non-compete fact patterns feature non-compete agreements that are a modified version of a sample noncompete agreement from Chapter 23 of the book, "Massachusetts Employment Law” written by Jennifer A. Henricks & Alicia J. Samolis 2026. https://www.mcle.org/product/catalog/code/2000269B00
Indonesian-Health-Newsaart-ai-safety-datasetaaronschlegel_austin-animal-center-shelter-outcomes-and
Austin Animal Center Shelter Outcomes
30,000 shelter animals
Dataset Info
Source: Kaggle
Original Size: 3.24 MB
Kaggle Downloads: 9,138
Files: 2
Files
aac_shelter_cat_outcome_eng.csv
aac_shelter_outcomes.csv
Mirrored from Kaggle
aarne-1910-tale-types
Aarne 1910 Tale-Type Index (Verzeichnis der Märchentypen)
A complete structured digitization of Antti Aarne's Verzeichnis der Märchentypen (Folklore Fellows' Communications 3, Helsinki 1910) — the founding catalogue of the Aarne–Thompson–Uther tale-type system. Every type with its German title and description, part/division/subsection structure, group captions, Grundtvig and Grimm cross-references, page numbers, and English title glosses. Parsed from the proofread German… See the full description on the dataset page: https://huggingface.co/datasets/wheelofheaven/aarne-1910-tale-types.planetmath
Planet Math Data
This dataset contains (most of) the pages from the website Planet Math. The data are organized into the columns name, url, and content. This was compiled using a modified version of the gist aarjaneiro/planetmath_docset.py.
NCERT-c10-12DH201_meta-kaggle_workshopMORBenchOfficial multilingual dataset for "Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary".
NOTATION
Here train subset is our test set used in paper.
Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC
Indian OSCC Somatic Driver Mutation Dataset
This repository contains the processed datasets used in the study: "Identification of Somatic Driver Mutations in Indian Oral Squamous Cell Carcinoma Using XGBoost and Integrative Genomic Features".
Files
train_MAF.csv – Somatic mutation data used for training
gene_variants.csv – Curated cancer driver gene list
ROH.csv – Runs of homozygosity intervals
SBS_Signature_HNSC.csv – Gene-level SBS13 annotation
synthetic_mutations.csv… See the full description on the dataset page: https://huggingface.co/datasets/aarushidas/Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC.ULMAR-80mahaepicrecallSHDL_DatasetMMU - Siti Hasmah Digital Library Training Dataset for LLM-based Virtual Assistants
Overview
This dataset is specifically designed to fine-tune Large Language Models (LLMs) like GPT, Mistral, and OpenELM for tasks in the context of Multimedia University (MMU) and the Siti Hasmah Digital Library. It has been crafted to address user interactions related to MMU services, admissions, scholarships, and library operations.
The dataset's goal is to facilitate domain adaptation, allowing institutions… See the full description on the dataset page: https://huggingface.co/datasets/AaronLim/SHDL_Dataset.small-aml-data
Dataset Credits
This project uses the IBM Transactions for Anti Money Laundering (AML) synthetic dataset, published by Erik Altman and collaborators at IBM Research.
Original dataset source:
Kaggle: https://www.kaggle.com/datasets/ealtman2019/ibm-transactions-for-anti-money-laundering-aml
IBM Research publication: https://research.ibm.com/publications/realistic-synthetic-financial-transactions-for-anti-money-laundering-models
IBM AML-Data repository: https://github.com/IBM/AML-Data… See the full description on the dataset page: https://huggingface.co/datasets/aaronzeller/small-aml-data.aarkoodataset24gameidentifying-debiasing-online-media-with-chatgptIS_298C_DatasetVotes.csvtraining_data_mistral
