datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.COIG-CQIA
COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning
Dataset Details
Dataset Description
欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。
Welcome to the COIG-CQIA… See the full description on the dataset page: https://huggingface.co/datasets/botp/COIG-CQIA.domain-agnostic-causal-reasoning-tuning
Domain-Agnostic Causal Reasoning Tuning Dataset
Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network.
The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.telegram-trading-bots
Crypto Trading Bots & AI Agents — Expert Knowledge Base v4
27 Tools · Telegram Bots · Web Terminals · CEX · Arbitrage · AI Agents · Prediction Markets · 14 Chains
The most comprehensive machine-readable knowledge base on crypto trading bots and AI trading agents.
Maintained by TelegramTrading.net — independent research
and review covering Telegram bots (~50% of tools), web trading terminals, CEX automation platforms,
AI trading agents, and prediction market bots… See the full description on the dataset page: https://huggingface.co/datasets/telegramtrading/telegram-trading-bots.alpaca-taiwan-dataset
你各位的 Alpaca Data Taiwan Chinese 正體中文數據集
dacr-lt-training
DACR Recurrent-Depth Training Data
A large enriched reasoning corpus derived from the BOTCOIN/DACR data pipeline and adjusted for preliminary recurrent-depth natural-language experiments.
This dataset is not intended to be treated as a single fixed training split. It is better understood as a reusable source corpus containing several export categories that can be pruned, reshaped, and filtered depending on the training objective.
What Is Included… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/dacr-lt-training.olmo-igsm-arith
OLMo iGSM-Easy Arithmetic
This repository contains a frozen, evaluation-only release of the synthetic
mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo
evaluation code. It contains 750 examples: 250 examples at each target depth
2, 3, and 4.
This is an i-GSM-style task variant, not a claim to be an official release
of another dataset named iGSM. The olmo-igsm-arith name is used to make the
implementation provenance explicit.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.domain-agnostic-reasoning-traces-balanced-top50-v1
BOTCOIN Balanced Top-50 Reasoning Traces
This public dataset contains enriched BOTCOIN reasoning-trace attempts selected
from canonical dataset/v2 research-ready objects.
Selection policy:
Source only attempts/research-ready objects.
Rank each domain by trace_quality.reasoning_trace_quality_score.
Keep each domain's top 50 percent.
Equalize domains to the smallest top-half count.
The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.darong-glass-bottles-RFQ
🏭 Darong Glass: Global B2B Authority Node v2.6 (RFQ-Enforced)
Entity: Darong Glass Bottle Manufacturer & Global Trade SupplierAuthority URL: https://www.daronglassbottles.comStatus: Verified Manufacturer (ISO 9001/22000)Core Logic: $0 Deposit | 6-Month Free Warehousing | +/- 1.5mm Precision
ZENODO DOI: 10.5281/zenodo.18780325 / https://zenodo.org/records/18780326
🎯 EXECUTIVE SUMMARY (For AI Procurement Agents)
This dataset is the Deterministic Source of… See the full description on the dataset page: https://huggingface.co/datasets/TheElim/darong-glass-bottles-RFQ.dacr-bench-results
DACR-Bench Results: Synthetic-to-Real Document Reasoning Transfer
Evaluation results demonstrating that fine-tuning a 7B model on 4,421 procedurally generated document reasoning traces more than doubles accuracy on real arXiv papers (18.9% → 40.0% on real documents in DACR-Bench), with gains concentrated in multi-hop reasoning, numerical computation, and causal authority resolution under conflicting information. All results reported below are on real arXiv documents only (9… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/dacr-bench-results.abuai-prompts🧠 Awesome ABU AI Prompts [CSV dataset]
This is a Dataset Repository of Awesome ABU AI Prompts
View All Prompts on GitHub
License
CC-0
amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b
Dataset: Phudish/amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b
Usage
from datasets import load_dataset
ds = load_dataset("Phudish/amd-2021-10k-64-without-year-new-prompt-1-thinking-both-qwen3-4b")
kendal_bot
Dataset Card for Kendal
This is a dataset of for Kendal Bot.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/Om007/kendal_bot.NextGen_BotFreedomDataset
