datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bird-species-10k
Bird-10K SigLIP Training Dataset
Taxonomy-aware image-text dataset for fine-tuning SigLIP on 10,753 bird species.
Dataset Structure
taxonomy.json # eBird v2025 taxonomy (species/genus/family/order hierarchy)
hard_negatives.json # Same-genus & same-family negative sampling index
siglip_train.jsonl # 692,236 training image-text pairs
siglip_val.jsonl # 32,125 validation pairs
images_224_*.tar.gz # 224x224 bird images organized by species… See the full description on the dataset page: https://huggingface.co/datasets/Hakureirm/bird-species-10k.open-instruct-v1
Open Instruct V1 - A dataset for having LLMs follow instructions.
Open Instruct V1 is an amalgamation of different datasets which are cleaned and then collated into a singular format for training.
Dataset Breakdown
Dataset
Amount of Samples
Alpaca
51759
Self Instruct
82599
GPT-4 Instruct
18194
Code Alpaca
18019
Dolly
15015
Synthetic
33143
Roleplay
3146
asss
448
instruction-dataset
327
Total
222650
HakHukuk-mevzuat-bge-m3-s2
HakHukuk — Mevzuat Arama İndeksi (bge-m3, şema s2)
Bu, bge-m3 yoğun gömme + BM25 hibrit arama indeksinin gömme matrisidir. HakHukuk hukuk
asistanının retriever katmanı, sorguyu ve mevzuat maddelerini bu indeksle eşleştirir.
İçerik
Bu depoda yalnızca şu iki dosya vardır — başka bir şey yoktur:
dosya
bayt
sha256
gomme.npy
82.935.936
16f54cab972eaccb4b30143a70d728faca07253385ce75d5455b4b407a62bedc
KUNYE.json
2.354… See the full description on the dataset page: https://huggingface.co/datasets/Rfetha/HakHukuk-mevzuat-bge-m3-s2.screenplay_emotions
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset was created by scrapping the screenplays from the imsdb website and then splitting them into 100 segments.
Each segment has been fed into a emotion classification model and classified into the emotion it evokes and represented as a number from 1 to 6.
Each number represents one of six emotions:
1 - joy
2 - love
3 - surprise
4 - sadness
5 - anger
6 - fear
These numbers are then… See the full description on the dataset page: https://huggingface.co/datasets/hakkam10/screenplay_emotions.Hakimi_test
哈基米(Hakimi)自我认知数据集
数据集描述
这是一个关于虚拟角色"哈基米"的自我认知对话数据集。哈基米是一只独特的"耄耋"(猫),由"南北绿豆"养育长大。该数据集包含了哈基米与用户之间的对话交互,展现了其独特的个性和自我认知。
数据格式
数据集采用 JSONL 格式,每行包含一个 JSON 对象,具有以下字段:
system: 定义哈基米的角色设定
conversation: 包含对话历史的数组,每个元素包含 human(用户输入)和 assistant(哈基米回复)
数据内容
数据集包含 59 个对话样本,涵盖了以下主题:
自我介绍(姓名、身份、特点)
与用户的问候互动
对自身身份的澄清(强调自己是"猫"而非AI)
与"南北绿豆"的关系
"曼波"等特色表达方式
数据特征
语言: 中文为主,包含少量英文
角色特点: 哈基米坚持自己是一只猫,而非AI助手
核心表达: 频繁提及"曼波"、"哈基米"等特色词汇
情感倾向: 友好、亲切,致力于传播"爱与和平"… See the full description on the dataset page: https://huggingface.co/datasets/weisiren/Hakimi_test.FaultPremiseDreadPoor__hakuchido-8B-MODEL_STOCK-details
Dataset Card for Evaluation run of DreadPoor/hakuchido-8B-MODEL_STOCK
Dataset automatically created during the evaluation run of model DreadPoor/hakuchido-8B-MODEL_STOCK
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__hakuchido-8B-MODEL_STOCK-details.HaKhanhPhuong_LLama2_Dataset_v1HAKE_TrainFightingICELLMconv-ai-yks-lgs-sfthako_contentconv-ai-yks-lgs-dpotest_apple
