datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ndlj_tosho_1
国会図書館に収蔵される著作権切れのデータです
claudette_tos
Dataset Card for "claudette_tos"
More Information needed
TOS_Dataset
TOS_Dataset
This dataset contains clauses from Terms of Service (ToS) documents with annotations indicating the fairness level of each clause. The dataset includes clauses labeled as clearly_fair, potentially_unfair, and clearly_unfair.
Dataset Summary
The dataset comprises clauses extracted from various ToS documents. Each clause is annotated with a fairness level, indicating whether it is clearly fair, potentially unfair, or clearly unfair.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/CodeHima/TOS_Dataset.tos_pp_dataset
A collection of Terms of Service or Privacy Policy datasets
Annotated datasets
CUAD
Specifically, the 28 service agreements from CUAD, which are licensed under CC BY 4.0 (subset: cuad).
Code
import datasets
from tos_datasets.proto import DocumentQA
ds = datasets.load_dataset("chenghao/tos_pp_dataset", "cuad")
print(DocumentQA.model_validate_json(ds["document"][0]))
100 ToS
From Annotated 100 ToS, CC BY 4.0 (subset: 100_tos).
Code
import… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/tos_pp_dataset.tosti-qwen-image-lora-datasetTOS_DatasetV3
TOS_DatasetV3
Dataset Description
TOS_DatasetV3 is a dataset designed for analyzing the unfairness of terms of service (ToS) clauses. It includes sentences from various terms of service agreements categorized into three unfairness levels: clearly_fair, potentially_unfair, and clearly_unfair. This dataset aims to aid in the development of models that can assess the fairness of legal documents.
Dataset Structure
The dataset consists of the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/CodeHima/TOS_DatasetV3.key-data
🔑 KEY: Neuroevolution Dataset
40,000+ logged events from real evolutionary runs — every mutation, crossover, selection, and fitness evaluation.
KEY evolves LoRA adapters on frozen base models (MiniLM-L6, DreamerV3) using NEAT-style neuroevolution. This dataset captures the complete evolutionary history.
🎮 Links
🌌 Live Demo
Watch evolution in action
🧠 Champion Model
The evolved DreamerV3 model
Loading the Dataset
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/tostido/key-data.unfair_tosmab_english
Dataset Card for [MAB]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/tosin/mab_english.ToS-SummariesLLaVA-CC3M-Pretrain-595K-JA
Dataset Card for "LLaVA-CC3M-Pretrain-595K-JA"
Dataset Details
Dataset Type:
Japanese LLaVA CC3M Pretrain 595K is a localized version of the original LLaVA Visual Instruct CC3M 595K dataset. This version is translated into Japanese using cyberagent/calm2-7b-chat and is aimed at serving similar purposes in the context of Japanese language.
Resources for More Information:
For information on the original dataset: liuhaotian/LLaVA-CC3M-Pretrain-595K
License:
Must comply with… See the full description on the dataset page: https://huggingface.co/datasets/toshi456/LLaVA-CC3M-Pretrain-595K-JA.c4-en-2k-tos-game-replay
Fixed English C4 replay subset
A subset of allenai/c4, English
configuration, training split. C4 is derived from Common Crawl; see the upstream
card for provenance and licensing. Sized against cfierro/tos_game_synthetic_docs, split
train, using raw text tokens without special tokens or truncation.
All 6,219 documents are in train, with 2,982,687 raw tokens.
Whole documents are kept until the target is reached; exact duplicate texts
are skipped. id is SHA-256 of the original… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/c4-en-2k-tos-game-replay.llava-bench-in-the-wild-jaThis dataset is the data that corrected the translation errors and untranslated data of the Japanese data in MBZUAI/multilingual-llava-bench-in-the-wild.
Original dataset is liuhaotian/llava-bench-in-the-wild.
kontext-tosti-lora-datasetTOSD
Dataset Card for Tamazight Open Speech Dataset
This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice recordings and corresponding text transcripts in Standard Moroccan Amazigh (ⵜⴰⵎⴰⵣⵉⵖⵜ ⵜⴰⵏⴰⵡⴰⵢⵜ ⵜⴰⵎⵓⵔⴰⴽⵓⵛⵜ) intended for training Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models.
This specific repository is published by a collaborator. You may visit the raw dataset repository which has additional dataset that hasn't… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/TOSD.tosti-qwen-loraRakuten-Alpaca-Data-32K
Dataset Card for "Rakuten-Alpaca-Data-32K"
Dataset Detail
Dataset Type: Rakuten-Alpaca-Data-32KはStanford Alpacaの手法を参考にRakuten/RakutenAI-7B-chatを使用して自動生成した日本語インストラクションデータです。
データ生成を行う際のSEEDデータには有志の方々が作成したseed_tasks_japanese.jsonlを利用させていただきました。
データの品質が低いため、何かしらの方法でフィルタリングして有益なデータのみ利用するのをおすすめします。
License: Apache license 2.0
Acknowledgement
Stanford Alpaca
Rakuten
seed_tasks_japanese.jsonl
tosti-v2-qwen-image-datasetunfair_tosacf-co24-tossupstoxic-dpo-v0.2-dutch
Toxic DPO v0.2 - Dutch Translation
Dataset Description
This is a direct machine-translated Dutch version of the original datasetunalignment/toxic-dpo-v0.2.
Translation method:English → Dutch using the Helsinki-NLP/opus-mt-en-nl model from the MarianMTModel translations.No manual edits, additions or filtering were applied besides the automated translation.
Data set is checked on NULL values and duplicates.
All fields (prompt, chosen, rejected) were translated… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic-dpo-v0.2-dutch.Tos_Datasetdocci_jaThis data was translated from the "DOCCI" into Japanese by DeepL
DOCCI: https://google.github.io/docci/
Lisence
CC-BY-4.0
structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.batterypass-12Kunfair_tos_fewshot_eval
Dataset Card for "unfair_tos_fewshot_eval"
More Information needed
acf-co24-tossupsTOS_DatasetV2TOS_sentence_embedded_all_minilm_l6_v2beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/tosiyama/beans.
