datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MARI-dataset
MARI Dataset
MARI dataset for music instruction-tuning.
Repository: https://github.com/Cactooz/MARI-dataset
Language: English
License: CC BY-NC-SA 4.0
Uses
Music Add Remove Instruction (MARI) dataset is a dataset for instruction-following music edits.
The dataset is used to train and evaluate text-to-music models for ADD and REMOVE editing operations.
Dataset Structure
The mari-dataset.parquet has the following structure.
Each row represents a… See the full description on the dataset page: https://huggingface.co/datasets/Cactooz/MARI-dataset.IDNet-2025
IDNet-2025 Dataset
IDNet-2025 is a novel dataset for identity document analysis and fraud detection. The dataset is entirely synthetically generated and does not contain any private information.
Dataset Structure
The dataset contains 1 models.tar.gz file, 10 LOC.tar.gz files, and 10 LOC_scanned.tar.gz files. Each pair of LOC.tar.gz file and LOC_scanned.tar.gz files belongs to a separate location in the world (European countries). Each LOC.tar.gz file includes a meta… See the full description on the dataset page: https://huggingface.co/datasets/cactuslab/IDNet-2025.hg38_cactus447waycactusIDSpace
IDSpace Dataset
Dataset Summary
IDSpace contains a large-scale synthetic dataset designed for the evaluation and benchmarking of digital identity verification and document fraud detection systems. The dataset was generated using the IDSpace framework, a model-guided synthetic document generation methodology that aligns generated documents with a target domain using only a small number of real samples.
Unlike existing synthetic identity document datasets that focus… See the full description on the dataset page: https://huggingface.co/datasets/cactuslab/IDSpace.counseling-cactus
Cactus — Counseling Dialogues (processed)
英文 CBT 咨询对话(Cactus, EMNLP'24 Findings);含 intake form / CBT plan / cognitive patterns / client attitude。
本仓库是 counselor_agent 项目中,经统一预处理器落地到 dataset/processed/ 的
Cactus 数据集。所有记录采用统一 schema(case_id / source / lang /
messages[] + 各数据集特有的可选标注 / profile)。
规模
cactus_train: 31,577 dialogues, 963,907 turns (avg 30.53)
cactus_eval: 450 dialogues, 900 turns (avg 2.0)
cactus_eval_zh: 450 dialogues, 900 turns (avg 2.0)… See the full description on the dataset page: https://huggingface.co/datasets/XuShihao6715/counseling-cactus.cactusFirehouse-Cactus-1.04-DatasetThis is the dataset Cactus 1.04 was trained on, it has the personality baked into it already.
needle-tokenizercactus-instruction-template
References
You can find the original dataset at cactus-camel/cactus and at
LangAGI-Lab/cactus. I only reformatted the the training prompts suggested in Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
Research Paper to be in instruction template format. I also removed some rows and did some small modification. You could also view the original collection at LangAGI-Lab
's Collections
Counter-Strike-2-rectangles-yoloCounselingEvalcactus-chat-template
Recognition and References
You can find the original dataset at cactus-camel/cactus and at
LangAGI-Lab/cactus. I only reformatted the the training prompts suggested in Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
Research Paper to be in chat template format. I also removed some rows and did some small modification.
edgeqa-resource-release
EdgeQA resource release (OLP + OSP)
This Hugging Face dataset repository contains a release-ready snapshot of the EdgeQA resource family:
EdgeQA (grounded QA exports at multiple budgets), EdgeCoverBench (robustness + abstention stress tests), and BEIR-style IR test collections.
Links
Code: https://github.com/cactusYuri/edgeqa
Author: Yuli Zhang (Beijing University of Posts and Telecommunications) — zhangyuli@bupt.edu.cn
Contents
corpora/: redistributable… See the full description on the dataset page: https://huggingface.co/datasets/cactusYuri/edgeqa-resource-release.cactus-dialogues-fr
Cactus (FR) : Dataset de dialogues traduits pour la Thérapie Cognitivo-Comportementale (TCC)
Description
Cactus (FR) est une adaptation en français d'une partie du dataset original "Cactus", un ensemble de dialogues multi-tours conçu pour simuler des interactions réalistes dans le cadre de séances de thérapie cognitivo-comportementale (CBT). Mon travail a consisté à traduire la colonne des dialogues afin de rendre ce dataset accessible à la communauté francophone, tout en… See the full description on the dataset page: https://huggingface.co/datasets/innermost47/cactus-dialogues-fr.gemma4-tflite-encoderscactus447way-zarr-zipcactoCactus-Mental-Health-datasetcactus_toygemma-4-e2b-it-cqgdelt-disaster-datasetsmp-extra-data
