CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alberto1231 /databricks_dolly_15k Databricks Dolly task samples Standalone task subsets derived from databricks/databricks-dolly-15k at revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a: general_qa (source category: general_qa) open_qa (source category: open_qa) closed_qa (source category: closed_qa) brainstorm (source category: brainstorming) classify (source category: classification) extract_information (source category: information_extraction) summarize (source category: summarization) creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.tabulartext-generationn<1K0 likes117 downloads2mo agoHugging Face02open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face03open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face04GenAIDevTOProd /databricks-dolly15k-semantic-complexity Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation) Overview This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks. Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.tabular10K<n<100K1 likes24 downloads1y agoHugging Face05open-llm-leaderboard /databricks__dolly-v1-6b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v1-6b Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face06open-llm-leaderboard /databricks__dolly-v2-3b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-3b Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face07Inversta /rationale-databricks-dolly-cqa Dataset Overview Filtered and annotated version of the closed-question answering part (~1.5k datapoints) of the Databricks Dolly Dataset intended for the task of rationale extraction. Citation @article{pirenne2024exploration, title={Exploration of Closed-Domain Question Answering Explainability Methods With a Sentence-Level Rationale Dataset}, author={Pirenne, Lize and Mokeddem, Samy and Ernst, Damien and Louppe, Gilles}, year={2024} }… See the full description on the dataset page: https://huggingface.co/datasets/Inversta/rationale-databricks-dolly-cqa.tabularquestion-answering1K<n<10K1 likes21 downloads2y agoHugging Face08albertge /databricks-dolly-15k-modernbert-train-kmeans-dim768-20250723tabular10K<n<100K0 likes21 downloads1y agoHugging Face09HydraLM /databricks-dolly-15k_standardizedtabular10K<n<100K0 likes20 downloads3y agoHugging Face10Cleanlab /databricks-dolly-15k-cleanset Summary databricks-dolly-15k-cleanset can be used to produced CLEANed up versions of the popular databricks-dolly-15k dataSET, which was used to fine-tune the Dolly 2.0. The original databricks-dolly-15k contains 15,000 human-annotated instruction-response pairs covering various categories. However, there are many low-quality responses, incomplete/vague prompts, and other problematic text lurking in the dataset (as with for all real-world instruction tuning datasets). We ran… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/databricks-dolly-15k-cleanset.tabular10K<n<100K2 likes20 downloads3y agoHugging Face11kranthigv /databricks-dolly-15k_standardizedtabular10K<n<100K0 likes19 downloads3y agoHugging Face12albertge /databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130tabular10K<n<100K0 likes15 downloads2y agoHugging Face13albertge /databricks-dolly-15k-tfidf-sweep-kmeans-dim10000-20250914tabular10K<n<100K0 likes14 downloads1y agoHugging Face14albertge /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250917tabular10K<n<100K0 likes13 downloads1y agoHugging Face15NamburiSrinath /databricks-dolly-15k-modernbert-train-kmeans-dim768-20250316tabular10K<n<100K0 likes10 downloads2y agoHugging Face16Ba2han /databricks-dolly_ratedFirst, I merged instruction and context columns because it's weird to have instructions saying "summarize this" without the passage itself. Then I used Senku-70B Q2 GGUF to rate each example out of 10 using a custom-made prompt based on clarity, completeness, correctness, relevance and formatting. Here are a few examples of below 5 pairs: Observations & Thoughts: There 1734 examples with <6.5 score and 562 examples with <5 score. Around 10% of the dataset looks low quality and/or confusing.… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/databricks-dolly_rated.tabular10K<n<100K0 likes8 downloads3y agoHugging Face17albertge /databricks-dolly-15k-tfidf-train-kmeans-dim10000-20250914tabular10K<n<100K0 likes7 downloads1y agoHugging Face18shimejii /databricks-dolly-15k-ja__resCount-roundtabular10K<n<100K0 likes5 downloads2y agoHugging Face19rchu233 /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes5 downloads2y agoHugging Face20shimejii /databricks-dolly-15k-ja__resCount-round-prompt_1-summarytabular1K<n<10K0 likes4 downloads2y agoHugging Face21NamburiSrinath /databricks-dolly-15k-modernbert-split-kmeans-dim768-20250917tabular10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.