datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sprintduplicatequestions-pairclassification
SprintDuplicateQuestions
An MTEB dataset
Massive Text Embedding Benchmark
Duplicate questions from the Sprint community.
Task category
t2t
Domains
Programming, Written
Reference
https://www.aclweb.org/anthology/D18-1131/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["SprintDuplicateQuestions"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/sprintduplicatequestions-pairclassification.sprintduplicatequestions-pairclassification-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["SprintDuplicateQuestions-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/sprintduplicatequestions-pairclassification-vn.Property_Inferencellm-watermark-detectionllm-eval-sprint_duplicate_questionsFL_Data_Reconstructionsam-controlnet-sprint-large-groundedDINO-mask
Dataset Card for "sam-controlnet-sprint-large-groundedDINO-mask"
More Information needed
watermark_localizationSprintDataset-0.1sam-controlnet-sprint-larg-v1
Dataset Card for "sam-controlnet-sprint-larg-v1"
More Information needed
DPhi_Sprint_25_Flowers
Dataset Card for "DPhi_Sprint_25_Flowers"
All images in this archive are licensed under the Creative Commons By-Attribution License, available at:
https://creativecommons.org/licenses/by/2.0/
The photographers are listed in LICENSE.txt, thanks to all of them for making their work available.
However, you will observe the image file names are different in this file than those we have provided. The file names were changed solely for the purpose of the data sprint.
base_model_sprint
Base Model Metadata Sprint
Description
Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models.
🤗 Strong contributions will win prizes!! 🤗
Why It Matters
Adding base_model metadata helps users:
Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.ALRAGE-Benchmark-sprint-resultssprint2-opc-sft-stage2-splitP4Ms-hackathon-vision-tasktest-chunked-sprint-D1
Generated Responses Dataset
This dataset contains generated responses for prompts from davanstrien/haiku_dpo.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Source Split: train
Input Column: question (plain text prompts)
Model: Qwen/Qwen2.5-3B-Instruct
Rows Processed: 5
Batches: 3 (chunk size: 2)
Generation Date: 2026-04-14T13:42:30.674651
Script: generate-responses-chunked.py (experimental streaming version)
Sampling Parameters
Temperature: 0.7… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-chunked-sprint-D1.sprint2
Sprint2
Made with ❤️ using 🦥 Unsloth Studio
Sprint2 was generated with Unsloth Recipe Studio. It contains 100 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("dtorette/sprint2", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 100
📋 Columns: 3
📋 Schema & Statistics
Column
Type
Column Type
Unique (%)
Null (%)
Details
llm_structured_1
dict… See the full description on the dataset page: https://huggingface.co/datasets/dtorette/sprint2.dissertation-sprint1-vuln-dataDUCIsam-controlnet-sprint-small-v1
Dataset Card for "sam-controlnet-sprint-small-v1"
More Information needed
llm_watermark_removalTabular_Attribute_InferenceSprintDataset-0.2SprintDataset-0.2.2sprintduplicatequestions-pairclassification-explodeddiffusers_sprint_cartoonify_yourself
Dataset Card for "diffusers_sprint_cartoonify_yourself"
More Information needed
vlm-guidance-sprint-3-votessprint_plannerlangley-sprint05-datasprintduplicatequestions-pairclassification-exploded-vn
