CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bowang0911 /databricks-qa-ja License & Attribution MTEB-format derivative of yulanfmy/databricks-qa-ja (Japanese Databricks/Dolly-style technical QA). Query = question; corpus = answer. Licensed under CC-BY-SA-3.0 (same as source). tabulartext-retrieval1K<n<10K0 likes329 downloads3mo agoHugging Face02Alberto1231 /databricks_dolly_15k Databricks Dolly task samples Standalone task subsets derived from databricks/databricks-dolly-15k at revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a: general_qa (source category: general_qa) open_qa (source category: open_qa) closed_qa (source category: closed_qa) brainstorm (source category: brainstorming) classify (source category: classification) extract_information (source category: information_extraction) summarize (source category: summarization) creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.tabulartext-generationn<1K0 likes107 downloads2mo agoHugging Face03GenAIDevTOProd /databricks-dolly15k-semantic-complexity Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation) Overview This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks. Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.tabular10K<n<100K1 likes30 downloads1y agoHugging Face04HydraLM /databricks-dolly-15k_standardizedtabular10K<n<100K0 likes20 downloads3y agoHugging Face05albertge /databricks-dolly-15k-modernbert-train-kmeans-dim768-20250723tabular10K<n<100K0 likes19 downloads1y agoHugging Face06albertge /databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130tabular10K<n<100K0 likes16 downloads2y agoHugging Face07kranthigv /databricks-dolly-15k_standardizedtabular10K<n<100K0 likes14 downloads3y agoHugging Face08albertge /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250917tabular10K<n<100K0 likes13 downloads1y agoHugging Face09albertge /databricks-dolly-15k-tfidf-sweep-kmeans-dim10000-20250914tabular10K<n<100K0 likes10 downloads1y agoHugging Face10NamburiSrinath /databricks-dolly-15k-modernbert-train-kmeans-dim768-20250316tabular10K<n<100K0 likes9 downloads2y agoHugging Face11Ba2han /databricks-dolly_ratedFirst, I merged instruction and context columns because it's weird to have instructions saying "summarize this" without the passage itself. Then I used Senku-70B Q2 GGUF to rate each example out of 10 using a custom-made prompt based on clarity, completeness, correctness, relevance and formatting. Here are a few examples of below 5 pairs: Observations & Thoughts: There 1734 examples with <6.5 score and 562 examples with <5 score. Around 10% of the dataset looks low quality and/or confusing.… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/databricks-dolly_rated.tabular10K<n<100K0 likes8 downloads3y agoHugging Face12shimejii /databricks-dolly-15k-ja__resCount-roundtabular10K<n<100K0 likes6 downloads2y agoHugging Face13rchu233 /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes6 downloads2y agoHugging Face14albertge /databricks-dolly-15k-tfidf-train-kmeans-dim10000-20250914tabular10K<n<100K0 likes6 downloads1y agoHugging Face15shimejii /databricks-dolly-15k-ja__resCount-round-prompt_1-summarytabular1K<n<10K0 likes4 downloads2y agoHugging Face16NamburiSrinath /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250917tabular10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.