CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Wanfq /Explore_Instruct_Rewriting_32k Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration | 📑 Paper | 🤗 Data | 🤗 Model | 🐱 Github Repo | Fanqi Wan†, Xinting Huang‡, Tao Yang†, Xiaojun Quan†, Wei Bi‡, Shuming Shi‡ † Sun Yat-sen University, ‡ Tencent AI Lab News Oct 16, 2023: 🔥 We're excited to announce that the Explore-Instruct datasets in brainstorming, rewriting, and math domains are now available on 🤗 Huggingface Datasets! Additionally, we've… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/Explore_Instruct_Rewriting_32k.text10K<n<100K7 likes109 downloads3y agoHugging Face02gabrielmbmb /rewriting-assistant Dataset Card for rewriting-assistant This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: pipeline.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/gabrielmbmb/rewriting-assistant/raw/main/pipeline.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/rewriting-assistant.textn<1K0 likes65 downloads2y agoHugging Face03argilla-warehouse /smollm-v2-rewriting Dataset Card for smollm-v2-rewriting This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/argilla-warehouse/smollm-v2-rewriting/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/smollm-v2-rewriting.text100K<n<1M0 likes59 downloads2y agoHugging Face04spacemanidol /query-rewriting-dense-retrieval2 likes55 downloads4y agoHugging Face05Wanfq /Explore_Instruct_Rewriting_10k Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration | 📑 Paper | 🤗 Data | 🤗 Model | 🐱 Github Repo | Fanqi Wan†, Xinting Huang‡, Tao Yang†, Xiaojun Quan†, Wei Bi‡, Shuming Shi‡ † Sun Yat-sen University, ‡ Tencent AI Lab News Oct 16, 2023: 🔥 We're excited to announce that the Explore-Instruct datasets in brainstorming, rewriting, and math domains are now available on 🤗 Huggingface Datasets! Additionally, we've… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/Explore_Instruct_Rewriting_10k.text10K<n<100K3 likes32 downloads3y agoHugging Face06mudasir13cs /ecommerce-query-rewriting #e-commerce-query-rewriting-dataset Hub: mudasir13cs/ecommerce-query-rewriting A dataset of 10,000 examples pairing ambiguous, context-dependent user queries with their fully resolved, context-aware rewrites for e-commerce product search. Built for fine-tuning LLMs to resolve pronouns, ellipsis, ordinals, and other conversational shortcuts using prior search context — the kind of resolution real shopping assistants need to handle turns like "show me that one" or "the cheaper… See the full description on the dataset page: https://huggingface.co/datasets/mudasir13cs/ecommerce-query-rewriting.texttext-generation10K<n<100K1 likes23 downloads2mo agoHugging Face07tcotter /squadv2-query-rewriting Dataset Card for squadv2_query_rewriting Synthetic data on top of SquadV2, focused on follow-up questions and then query rewriting optimized for retrieval. Dataset Details Dataset Sources Repository: SquadV2 texttext-generationn<1K1 likes21 downloads2y agoHugging Face08trieunh /Text-Rewritingtext1K<n<10K0 likes18 downloads2d agoHugging Face0911-47 /self_rewriting_meta_learning_god_seed_25kgated Self-Rewriting God Seed AI — Max Distill + Self Meta-Learning The ultimate dataset for creating truly autonomous, self-modifying, god-level recursive superintelligence. This 25,000-example dataset is specifically engineered to turn any LLM into a Self-Rewriting AI with Self Meta-Learning Thinking and God-Level Recursive Seed AI Mindset. What Makes This Dataset God-Level This is not regular instruction tuning. This is intelligence explosion engineering at the… See the full description on the dataset page: https://huggingface.co/datasets/11-47/self_rewriting_meta_learning_god_seed_25k.text10K<n<100K2 likes16 downloads5mo agoHugging Face10kilicai /turkish-sft-rewriting-10k kilicai/turkish-sft-rewriting-10k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-rewriting-10k') text10K<n<100K1 likes10 downloads4mo agoHugging Face11unhif /rewritingtext10M<n<100M0 likes8 downloads2mo agoHugging Face12mlfoundations-dev /r1_rewriting_math_with_gttextn<1K0 likes7 downloads2y agoHugging Face13mlfoundations-dev /multiple_samples_rewriting_baselinetabular10K<n<100K0 likes7 downloads2y agoHugging Face14mlfoundations-dev /r1_rewriting_math_without_gttextn<1K0 likes6 downloads2y agoHugging Face15shanaka95 /arxiv-abstract-rewritingtext100K<n<1M0 likes6 downloads5mo agoHugging Face16kienhoang123 /Query-Rewriting-datasettabular1K<n<10K0 likes5 downloads2y agoHugging Face17mlfoundations-dev /multiple_samples_rewritingtabular10K<n<100K0 likes5 downloads2y agoHugging Face18Ashfaq1988 /Urdu-NLP-Style-Rewritingtextn<1K0 likes4 downloads2y agoHugging Face19kilicai /turkish-sft-rewriting_20k kilicai/turkish-sft-rewriting_20k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-rewriting_20k') text10K<n<100K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.