datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.MMLU-ProX_EN_Cleaned
MMLU-ProX English Cleaned
Dataset Description
This is a cleaned version of the English subset from MMLU-ProX (arXiv:2503.10497),
a comprehensive multilingual benchmark for evaluating large language models. The original MMLU-ProX dataset
contains 11,829 questions across 29 languages, built on the English MMLU-Pro benchmark.
Why This Cleaned Version?
The original English subset of MMLU-ProX contained spacing issues where words were concatenated without
proper… See the full description on the dataset page: https://huggingface.co/datasets/ZQ-Dev/MMLU-ProX_EN_Cleaned.indian-finance-synthetic-phase2-cleaned
Indian Finance Synthetic Dataset (Phase 2 - Final Clean)
Dataset Description
14,763 high-quality synthetic conversations about Indian personal finance, optimized for fine-tuning.
Recent Updates
✅ v3 (Final): Removed 14 samples with empty content messages
✅ v2: Removed 58 incomplete conversations
✅ v1: Tools optimization (82.5% size reduction)
All conversations are now complete and properly formatted for training.
Key Features
Clean… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2-cleaned.kowiki-cleaned
NLP-07-ODQA/kowiki-cleaned
Dataset Description
이 데이터셋은 한국어 위키피디아 XML 덤프에서 추출하고 정제한 데이터셋입니다. 청킹 전 단계의 정제된 문서를 포함합니다.
데이터 소스
원본: 한국어 위키피디아 XML 덤프
네임스페이스: 0 (일반 문서)
처리된 페이지 수: 565,484개
전체 페이지 수: 2,157,147개
데이터 정제 과정
위키 마크업 제거: [[링크]], {템플릿}, ==제목== 등 제거
불필요한 섹션 제거: '같이 보기', '외부 링크', '참고 문헌', '각주' 등 제거
필터링: 리다이렉트, 빈 페이지, 스텁 페이지 제거
최소 길이: 200자 이상만 포함
통계
평균 텍스트 길이: 1871.8자
데이터 구조
각 데이터 포인트는 다음 필드를 포함합니다:
content: 정제된 텍스트 (전체… See the full description on the dataset page: https://huggingface.co/datasets/NLP-07-ODQA/kowiki-cleaned.cleaned_nvidia_OpenCodeReasoning元データ: https://huggingface.co/datasets/nvidia/OpenCodeReasoning
データ件数: 11,275
平均トークン数: 11251
最大トークン数: 19,802
合計トークン数: 126,859,041
ファイル形式: JSONL
ファイルサイズ: 707.4 MB
難易度スコアが15, カテゴリがcompetition、ライセンスがmitとcc-by-4.0をピックアップ
繰り返し除去
極端に少ない・多いなどを除去
詳しいコードはGithub
https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/opencodereasoning
