datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
databricks-qa-ja
License & Attribution
MTEB-format derivative of yulanfmy/databricks-qa-ja (Japanese Databricks/Dolly-style technical QA). Query = question; corpus = answer. Licensed under CC-BY-SA-3.0 (same as source).
databricks_dolly_15k
Databricks Dolly task samples
Standalone task subsets derived from
databricks/databricks-dolly-15k at
revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a:
general_qa (source category: general_qa)
open_qa (source category: open_qa)
closed_qa (source category: closed_qa)
brainstorm (source category: brainstorming)
classify (source category: classification)
extract_information (source category: information_extraction)
summarize (source category: summarization)
creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.databricks-dolly15k-semantic-complexity
Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation)
Overview
This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks.
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.databricks-dolly-15k_standardizeddatabricks-dolly-15k-modernbert-train-kmeans-dim768-20250723databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130databricks-dolly-15k_standardizeddatabricks-dolly-15k-modernbert-split-kmeans-dim768-20250917databricks-dolly-15k-tfidf-sweep-kmeans-dim10000-20250914databricks-dolly-15k-modernbert-train-kmeans-dim768-20250316databricks-dolly_ratedFirst, I merged instruction and context columns because it's weird to have instructions saying "summarize this" without the passage itself.
Then I used Senku-70B Q2 GGUF to rate each example out of 10 using a custom-made prompt based on clarity, completeness, correctness, relevance and formatting. Here are a few examples of below 5 pairs:
Observations & Thoughts:
There 1734 examples with <6.5 score and 562 examples with <5 score. Around 10% of the dataset looks low quality and/or confusing.… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/databricks-dolly_rated.databricks-dolly-15k-ja__resCount-rounddatabricks-dolly-15k-modernbert-split-kmeans-dim768-20250130databricks-dolly-15k-tfidf-train-kmeans-dim10000-20250914databricks-dolly-15k-ja__resCount-round-prompt_1-summarydatabricks-dolly-15k-modernbert-split-kmeans-dim768-20250917
