datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DORI-instruction-tuning-dataset
DORI Spatial Reasoning Instruction Dataset
Dataset Description
This dataset contains instruction tuning data for spatial reasoning tasks across multiple question types and visual datasets.
Dataset Structure
Dataset Splits
train: 26,626 samples
test: 6,672 samples
Total: 33,298 samples
Question Types
q1
q2
q3
q4
q5
q6
q7
Source Datasets
3d_future
cityscapes
coco
coco_space_sea
get_3d
jta
kitti
nocs_real
objectron… See the full description on the dataset page: https://huggingface.co/datasets/appledora/DORI-instruction-tuning-dataset.behavioral-fine-tuning-v1
Why This Dataset Exists
"A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks."
The Problem
Our Solution
Uncensored data → helpful but uncontrolled
Surgical 85% helpfulness + 13% safety + 2% eval mix
Safety-only data → lobotomized, over-refusing models
Calibrated ratio preserves full helpfulness
Raw data → PII, leaked secrets, duplicates
7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.cartoonization
Instruction-prompted cartoonization dataset
This dataset was created from 5000 images randomly sampled from the Imagenette dataset. For more
details on how the dataset was created, check out this directory.
Following figure depicts the data preparation workflow:
Known limitations and biases
The dataset was derived from Imagenette, which, in turn, was derived from ImageNet. So, naturally, this
dataset inherits the limitations and biases of ImageNet.… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/cartoonization.low-level-image-proc
Instruction-prompted low-level image processing dataset
To construct this dataset, we took different number of samples from the following datasets for each task and constructed
a single dataset with prompts added like so:
Task
Prompt
Dataset
Number of samples
Deblurring
“deblur the blurry image”
REDS (train_blur and train_sharp)
1200
Deraining
“derain the image”
Rain13k
686
Denoising
“denoise the noisy image”
SIDD
8
Low-light image enhancement
"enhance the… See the full description on the dataset page: https://huggingface.co/datasets/instruction-tuning-sd/low-level-image-proc.AutoDims-Visual-Tuning-V1.1
AutoDims Visual Tuning V1.0
Dataset pushed from Dataset Engine.
vlm_fine_tuningvisual_instruction_tuning_ID_referenceAPR-single-hunk-fine-tuningThis dataset is used to fine-tune LLMs for automated program repair at single-function level, especially for single-hunk bugs. We provide two versions with different outputs:
The input is a buggy function and the output is a fixed function.
The input is a buggy function and the output is a unified diff for the bug.
frontend-instruction-tuningcat-breed-blip-fine-tuningfine_tuning_diffusionArtificial-super-girlfriend-for-fine-tuningリアル系モデルに特有の肖像権の問題について比較的クリアなモデルを作ることが可能なように、私が私自身から作り出した人工超彼女(ver 2.1系、ver 2.6系)のデータセット(約2800枚)を作成しました。
全ての元画像(加工前)がbeauty score 87以上なのが特徴であり、特にbeauty score 90以上の女性画像のデータセットとして、1000枚以上揃えているのは有数の規模だと思います。
具体的には、以下のように構成されています(87はこの子/私の最大のライバルが到達した最高得点、90は今のところ実在人物では確認できていない得点ラインです)。
version \ beauty score
87~89
90~
2.1(可愛いと綺麗のバランスを追求)
kawaii (無加工362枚/加工後724枚)
exceptional (無加工140枚/加工後280枚)
2.6(綺麗さ・美しさに特化)
beautiful (無加工464枚/加工後928枚)
perfect (無加工416枚/加工後832枚)… See the full description on the dataset page: https://huggingface.co/datasets/ThePioneer/Artificial-super-girlfriend-for-fine-tuning.ecu_tuningfine_tuning05Feb-fine-tuning-batch-3win_fine_tuning_stab_diff_train
Dataset Card for "win_fine_tuning_stab_diff_train"
More Information needed
coco_fine_tuning_diffusers
Dataset Card for "coco_fine_tuning_diffusers"
More Information needed
FINE_TUNING_LLM_QWEN2win_fine_tuning_stab_diff_eval
Dataset Card for "win_fine_tuning_stab_diff_eval"
More Information needed
sdxl_fine_tuning_datasetAutoDims-Visual-Tuning-V1.0
AutoDims Visual Tuning V1.0
Dataset pushed from Dataset Engine.
AutoDims-Visual-Tuning-V1.2
AutoDims Visual Tuning V1.0
Dataset pushed from Dataset Engine.
VVM-Tuning_Train_Datavision_fine_tuning_withimagesFine-Tuning-Deck-Image-Generatorstable-diffusion-fine-tuning-svg-data
