datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-animal-illustrations
SVG Pet Illustrations Dataset
A dataset of 1,416 prompt-SVG pairs for training text-to-SVG generation models, with a focus on cute animal illustrations.
Dataset Description
This dataset contains text prompts paired with their corresponding SVG code, designed for training models to generate vector graphics from natural language descriptions.
Dataset Statistics
Total examples: 1,416
Format: JSONL (JSON Lines)
Fields: prompt, svg
Content
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yoavf/svg-animal-illustrations.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
Animal-nutritionsharechat-animal-welfare-coarse-filter
ShareChat Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: tucnguyen/ShareChat
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 129,584
matched conversations: 7,606
match rate: 5.8696%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms, _text_preview, and… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/sharechat-animal-welfare-coarse-filter.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
sharelm-animal-welfare-coarse-filter
ShareLM Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: shachardon/ShareLM
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 3,551,155
matched conversations: 228,192
match rate: 6.4259%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms, _text_preview, and… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/sharelm-animal-welfare-coarse-filter.lmsys-chat-1m-animal-welfare-coarse-filter
LMSYS Chat 1M Animal Welfare Coarse Filter
Public working dataset for Compassion in Machine Learning.
Source dataset: lmsys/lmsys-chat-1m
Filter package: flpc
Filter used: original coarse animal-welfare keyword list provided by the project team.
Counts:
total conversations scanned: 1,000,000
matched conversations: 16,527
match rate: 1.6527%
Files:
matches.parquet: one row per matched conversation, preserving all original source fields/columns plus _matched_terms… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/lmsys-chat-1m-animal-welfare-coarse-filter.animal-alignment-feedback
Open Paws Animal Alignment Feedback
🐾 Human feedback and preference data for aligning AI with animal advocacy values
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Feedback Data
Format: CSV (Comma-separated values)
Languages: Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/animal-alignment-feedback.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
Italian_latin_parallel_animals
descrizioni di animali e habitat - Synthetic Dataset
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: descrizioni di animali e habitat
Field 1: italiano
Field 2: latino antico(traduzione)
Rows: 280
Generated on: 2025-05-27T00:07:49.042Z
