datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Optimized_Reasoning
Optimized_Reasoning
SUPPORT ME ON PATREON
https://www.patreon.com/c/Rombodawg
Optimized_Reasoning was created because even modern LLM's are not very good at handling reasoning very well, and if they are, they still waste tons of tokens in the process. With this dataset I hope to accomplish 2 things:
Reduce token usage
Increase model strength in reasoning
So how does this dataset accomplish that? By Adding a "system_prompt" like reasoning tag to the beggining of every data… See the full description on the dataset page: https://huggingface.co/datasets/Rombo-Org/Optimized_Reasoning.optimized-sd-configmedical-o1-reasoning-SFT-jsonl-optimizedlean-expert-optimized-2000
lean-expert-optimized-2000
Dataset Description
Optimized 2000-example dataset for training Lean trading algorithm optimization agents with 94%+ success rate target.
Dataset Statistics
Total Examples: 2,000
Training Examples: 1800
Validation Examples: 200
Target Success Rate: 94%+
Expected Performance: 96% (94-98% range)
Category Distribution
JSON Parsing: 1,333 examples (CRITICAL - 0% → 95% impact)
Optimization Workflows: 182 examples (HIGH… See the full description on the dataset page: https://huggingface.co/datasets/Kronu/lean-expert-optimized-2000.optimized_prompts_llama_3llm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.tifr-data-merged-optimized
license: mit
task_categories:
- translation
language:
- en
optimized-solidity-datasetmarketing-leads-optimized
Lead Response Prediction - Fine-tuning with Unsloth
📋 项目概述
使用Unsloth微调Llama 3.1模型,预测潜在客户(Lead)的响应行为,包括:
响应类型: replied, email_opened, connection_accepted, meeting_completed等
响应推理: 为什么lead会这样反应
下一步建议: 应该采取什么行动
📊 数据集优化总结
原始数据 → 优化数据对比
指标
优化前
优化后
提升
总样本
1,689
2,588
+53%
replied样本
116 (6.9%)
200 (12.1%)
+72%
meeting样本
16 (0.9%)
200 (12.1%)
+1150%
多touchpoint占比
19.6%
~40%
+2倍
主要优化
类别平衡 - Smart策略
关键标签过采样到200… See the full description on the dataset page: https://huggingface.co/datasets/MotionG-ai/marketing-leads-optimized.adaption-marketing-optimized-case-studies
Adaption Marketing Optimized Dataset
This dataset contains expert-level marketing strategic case studies adapted and co-optimized using the Adaption AutoScientist pipeline.
Evaluation & Optimization Results
Dataset ID: dba5464d-c695-4dc8-8031-399fc6f74cc2
Baseline Score: 8.0
Optimized Score: 8.6
Improvement Percent: 7.5%
Pipeline Settings
Deduplication: Enabled
Prompt Rephrasing: Enabled
Reasoning Traces: Enabled (Chain-of-Thought reasoning… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-case-studies.finance_customer_service_10000_optimizednyxmed-icd-optimized-20241221nyxmed-v4-icd-optimized-20241222rsh345__mistral-ft-optimized-1218-NeuralHermes-2.5-Mistral-7B-details
Dataset Card for Evaluation run of rsh345/mistral-ft-optimized-1218-NeuralHermes-2.5-Mistral-7B
Dataset automatically created during the evaluation run of model rsh345/mistral-ft-optimized-1218-NeuralHermes-2.5-Mistral-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rsh345__mistral-ft-optimized-1218-NeuralHermes-2.5-Mistral-7B-details.optimized_qa_commucation将选择题目对调整为问答对,并使用claude和gemma进行问题优化
adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.seo_optimized_bulletpoints_training_datagdrive-sbpn-fresh-diarization-optimized-terminal-l4-20260814-benchmarkoptimized_vantage_data_mehmood-ramllama_optimizedhf_evaluator_optimized_v10Aleex-Dataset-Optimized
