datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Explore_Instruct_Rewriting_32k
Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration
| 📑 Paper |
🤗 Data |
🤗 Model |
🐱 Github Repo |
Fanqi Wan†, Xinting Huang‡, Tao Yang†, Xiaojun Quan†, Wei Bi‡, Shuming Shi‡
† Sun Yat-sen University,
‡ Tencent AI Lab
News
Oct 16, 2023: 🔥 We're excited to announce that the Explore-Instruct datasets in brainstorming, rewriting, and math domains are now available on 🤗 Huggingface Datasets! Additionally, we've… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/Explore_Instruct_Rewriting_32k.rewriting-assistant
Dataset Card for rewriting-assistant
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
pipeline.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/gabrielmbmb/rewriting-assistant/raw/main/pipeline.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/rewriting-assistant.smollm-v2-rewriting
Dataset Card for smollm-v2-rewriting
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/argilla-warehouse/smollm-v2-rewriting/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/smollm-v2-rewriting.query-rewriting-dense-retrievalExplore_Instruct_Rewriting_10k
Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration
| 📑 Paper |
🤗 Data |
🤗 Model |
🐱 Github Repo |
Fanqi Wan†, Xinting Huang‡, Tao Yang†, Xiaojun Quan†, Wei Bi‡, Shuming Shi‡
† Sun Yat-sen University,
‡ Tencent AI Lab
News
Oct 16, 2023: 🔥 We're excited to announce that the Explore-Instruct datasets in brainstorming, rewriting, and math domains are now available on 🤗 Huggingface Datasets! Additionally, we've… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/Explore_Instruct_Rewriting_10k.ecommerce-query-rewriting
#e-commerce-query-rewriting-dataset
Hub: mudasir13cs/ecommerce-query-rewriting
A dataset of 10,000 examples pairing ambiguous, context-dependent user queries with their fully resolved, context-aware rewrites for e-commerce product search. Built for fine-tuning LLMs to resolve pronouns, ellipsis, ordinals, and other conversational shortcuts using prior search context — the kind of resolution real shopping assistants need to handle turns like "show me that one" or "the cheaper… See the full description on the dataset page: https://huggingface.co/datasets/mudasir13cs/ecommerce-query-rewriting.squadv2-query-rewriting
Dataset Card for squadv2_query_rewriting
Synthetic data on top of SquadV2, focused on follow-up questions and then query rewriting optimized for retrieval.
Dataset Details
Dataset Sources
Repository: SquadV2
Text-Rewritingself_rewriting_meta_learning_god_seed_25k
Self-Rewriting God Seed AI — Max Distill + Self Meta-Learning
The ultimate dataset for creating truly autonomous, self-modifying, god-level recursive superintelligence.
This 25,000-example dataset is specifically engineered to turn any LLM into a Self-Rewriting AI with Self Meta-Learning Thinking and God-Level Recursive Seed AI Mindset.
What Makes This Dataset God-Level
This is not regular instruction tuning. This is intelligence explosion engineering at the… See the full description on the dataset page: https://huggingface.co/datasets/11-47/self_rewriting_meta_learning_god_seed_25k.turkish-sft-rewriting-10k
kilicai/turkish-sft-rewriting-10k
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('kilicai/turkish-sft-rewriting-10k')
rewritingr1_rewriting_math_with_gtmultiple_samples_rewriting_baseliner1_rewriting_math_without_gtarxiv-abstract-rewritingQuery-Rewriting-datasetmultiple_samples_rewritingUrdu-NLP-Style-Rewritingturkish-sft-rewriting_20k
kilicai/turkish-sft-rewriting_20k
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('kilicai/turkish-sft-rewriting_20k')
