datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dream_NLP_FineTuneCars196_dino3_finetune_b
FAISS Index and Results for Cars196_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the Cars196 dataset.
Dataset: Cars196
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_b.CUB_dino3_finetune_b
FAISS Index and Results for CUB_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the CUB dataset.
Dataset: CUB
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path, split, and… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/CUB_dino3_finetune_b.benchmark-finetune-dpo-v1
Odyn benchmark: DPO LoRA fine-tuning peak VRAM (V1)
Curated benchmark rows for validating GPU memory estimators during DPO + LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance.
This dataset is not preference-pair training JSONL (UltraFeedback-style). It is evaluation ground truth for placement / scheduler memory… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-dpo-v1.StanfordOnlineProducts_dino3_finetune_b
FAISS Index and Results for StanfordOnlineProducts_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the StanfordOnlineProducts dataset.
Dataset: StanfordOnlineProducts
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/StanfordOnlineProducts_dino3_finetune_b.benchmark-finetune-lora-v1
Odyn benchmark: LoRA fine-tuning peak VRAM (V1)
Curated benchmark rows for validating GPU memory estimators during LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance.
This dataset is not Alpaca-style training JSONL. It is evaluation ground truth for placement / scheduler memory models (Odyn Smart Digester math… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-lora-v1.Siamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc)
and text scraped from books, news articles, reviews etc
license: apache-2.0
Cars196_dino3_finetune_s
FAISS Index and Results for Cars196_dino3_finetune_s
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (s) model, evaluated on the Cars196 dataset.
Dataset: Cars196
DINO Version: 3
DINO Size: s
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_s.CUB_dino3_finetune_s
FAISS Index and Results for CUB_dino3_finetune_s
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (s) model, evaluated on the CUB dataset.
Dataset: CUB
DINO Version: 3
DINO Size: s
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path, split, and… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/CUB_dino3_finetune_s.benchmark-dataset-finetune
Fine-Tuning VRAM Benchmark Dataset
Benchmark dataset for evaluating the accuracy of the Odyn Smart Digester VRAM Math Engine for fine-tuning workloads.
Compares the V1 (initial) and V2 (updated) engine estimates against expected peak VRAM values sourced from published research papers and hardware measurements.
Dataset Details
10 workload rows — all with gradient checkpointing enabled
Methods covered — LoRA (bf16) and QLoRA (NF4)
Models — Llama 2 7B, Llama… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-finetune.StanfordOnlineProducts_dino3_finetune_s
FAISS Index and Results for StanfordOnlineProducts_dino3_finetune_s
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (s) model, evaluated on the StanfordOnlineProducts dataset.
Dataset: StanfordOnlineProducts
DINO Version: 3
DINO Size: s
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/StanfordOnlineProducts_dino3_finetune_s.isear_finetuned_embeddingsfinetunetrydebbie_fine_tunebatch2_uncleaned_finetunefineTuneDeepSeekdatasets_model_finetuned_frmy-fine-tuned-chatbot-eval
