datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-animal-illustrations
SVG Pet Illustrations Dataset
A dataset of 1,416 prompt-SVG pairs for training text-to-SVG generation models, with a focus on cute animal illustrations.
Dataset Description
This dataset contains text prompts paired with their corresponding SVG code, designed for training models to generate vector graphics from natural language descriptions.
Dataset Statistics
Total examples: 1,416
Format: JSONL (JSON Lines)
Fields: prompt, svg
Content
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yoavf/svg-animal-illustrations.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full annotation across five dimensions.
Annotator: Mandy Hathaway — AI ethics specialist and technical writer with an MA in Ethical Technology & Artificial Intelligence. mandyhathaway.com
Dataset Summary
Most public preference datasets optimize for general helpfulness or… See the full description on the dataset page: https://huggingface.co/datasets/animasuri/Ai_ethics_dataset.Animal-nutritionsharechat-animal-welfare-coarse-filter
ShareChat Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: tucnguyen/ShareChat
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 129,584
matched conversations: 7,606
match rate: 5.8696%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms, _text_preview, and… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/sharechat-animal-welfare-coarse-filter.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
sharelm-animal-welfare-coarse-filter
ShareLM Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: shachardon/ShareLM
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 3,551,155
matched conversations: 228,192
match rate: 6.4259%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms, _text_preview, and… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/sharelm-animal-welfare-coarse-filter.anima-prompt-expander-v0.4
Anima Prompt Expander Dataset v0.4
作者:kitt3n
用于将中文视觉描述转换为适合 Anima-Aesthetic 的简洁英文 prompt。输出为英文标签与短关系短语,保留主体、动作、关系、构图和指定风格,只补充有视觉用途的细节。
数据组成
Split
样本
概念
train
900
450
validation
50
25
test
50
25
合计
1,000
500
按概念划分,同一概念的两种表述不会跨 split。保留来源表中的 3 条完全相同的重复输入与答案;校验未发现跨 split 的相同输入或概念泄漏,也未发现相同输入对应冲突答案。
原始 val.jsonl 文件的 split 字段仍为 val;HF 加载时将该文件映射到标准 validation split。
来源
来源于本仓库的 anima_pairs_1000_review_v0.4.md 审阅表:200 条已确认 pilot 加 800… See the full description on the dataset page: https://huggingface.co/datasets/kitt3n/anima-prompt-expander-v0.4.lmsys-chat-1m-animal-welfare-coarse-filter
LMSYS Chat 1M Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: lmsys/lmsys-chat-1m
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 1,000,000
matched conversations: 16,527
match rate: 1.6527%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/lmsys-chat-1m-animal-welfare-coarse-filter.anima-corpus-ko-fineweb2-broad
anima-corpus-ko-fineweb2-broad
🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining.
anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다.
Source
Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트).
Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards).
Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.animal-alignment-feedback
Open Paws Animal Alignment Feedback
🐾 Human feedback and preference data for aligning AI with animal advocacy values
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Feedback Data
Format: CSV (Comma-separated values)
Languages: Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/animal-alignment-feedback.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
Italian_latin_parallel_animals
descrizioni di animali e habitat - Synthetic Dataset
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: descrizioni di animali e habitat
Field 1: italiano
Field 2: latino antico(traduzione)
Rows: 280
Generated on: 2025-05-27T00:07:49.042Z
