datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GVP_Bot_State_Newegopi_latal_openarm_bottlepaper-recommendations-v2RyokoAI_CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.evaluator-leaderboardfineweb_2_500k_both_deduplicatedbottleneck-oracle-graphsCOIG-CQIA
COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning
Dataset Details
Dataset Description
欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。
Welcome to the COIG-CQIA… See the full description on the dataset page: https://huggingface.co/datasets/botp/COIG-CQIA.zalo-bot-registry
bep40/zalo-bot-registry
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('bep40/zalo-bot-registry')
Fineweb_2_500k_bothCabra3kO conjunto de dados Cabra é uma coleção ampla e diversificada de 3.000 entradas ou conjuntos de perguntas e respostas (QA) sobre o Brasil. Inclui tópicos variados como história, política, geografia, cultura, cinema, esportes, ciência e tecnologia, governo e muito mais. Este conjunto foi cuidadosamente elaborado e selecionado pela nossa equipe, garantindo alta qualidade e relevância para estudos e aplicações relacionadas ao Brasil.
Detalhes do Conjunto de Dados:
Tamanho: 3.000 conjuntos de QA… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/Cabra3k.aloha_static_bottle_pick_upcantonese-mandarin-translations
Dataset Card for cantonese-mandarin-translations
Dataset Summary
This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese).
Supported Tasks and Leaderboards
N/A
Languages
Cantonese (yue)
Simplified Chinese (zh-CN)
Dataset Structure
JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.WordVoice-5A
WordVoice-5A Dataset 🚀
A Large-Scale Bilingual Word-level Five-Annotation Dataset for WordVoice
📖 Dataset Description / 数据集简介
WordVoice-5A is a large-scale bilingual (Mandarin and English) dataset containing approximately 4.7k hours of speech with fine-grained word-level acoustic annotations, designed for high-precision controllable Text-to-Speech (TTS). It addresses the scarcity of large-scale, high-quality word-aligned datasets with explicit acoustic… See the full description on the dataset page: https://huggingface.co/datasets/botp/WordVoice-5A.so101-sword-on-stand-clean32-bothtrim-v2
SO-101 sword-on-stand clean32, physically both-end trimmed
This is the corrected, portable LeRobot v2.1 derivative for the task:
Pick up the sword and place it on the sword stand.
Dataset contract
32 episodes / 6,974 frames / 30 Hz.
Train episodes: 0..29 (6,582 frames).
Validation episodes: 30, 31 (392 frames).
Fixed and wrist RGB cameras, 640x480, H.264, 30 FPS.
observation.state and stored action are calibrated 6D absolute joint positions.
Only episodes 0..29… See the full description on the dataset page: https://huggingface.co/datasets/WDong/so101-sword-on-stand-clean32-bothtrim-v2.pi-publish
Coding agent session traces for deepflame-bot/pi-publish
This dataset contains redacted coding agent session traces collected while working on https://github.com/xke-b/efno-chem-kinetics.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/deepflame-bot/pi-publish.botbots-tod(https://github.com/radi-cho/botbots/tree/main/tod)
open-bottleneck-ranklong27b-slurm-364982-rollouts
Open Bottleneck RankLong 27B — Slurm array 364982
Compact rollout evidence archived from completed Slurm array 364982.
Config
Files / steps
Records
JSONL bytes
Note
rank_a40
60 (1–60)
15,360
66,445,670
Complete local rollout evidence
rank_a80
54 (1–54)
13,824
59,386,451
Includes the cancelled arm's final dumped step (54.jsonl)
Each JSONL record contains input, output, gts, score, acc,
response_length, grouprel_reward, and step.
Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.intern-bottle-lerobot
bottle: robot demonstrations
Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat.
LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.ctr-pick-dual-bottles-original-20260919
Pick Dual Bottles Original — shared50 scene cohort
This LeRobot v3 release contains 50 successful simulated demonstrations and
8,185 action rows at25FPS. Every source seed occurs exactly once. The source
seed set matches the current CTR Q1–Q3 Concurrent, CTR, Sequential, Mixed,
Left-first and Right-first datasets. Pair by retime.source_seed, not episode
index: composition datasets may have different ordering.
Mask limitation: retime.left_idle and retime.right_idle are boolean… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-pick-dual-bottles-original-20260919.PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b.
*Somente para demonstração e pesquisa. Proibido para uso comercial.
PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo.
*For demonstration and research purposes only. Commercial use prohibited.
yentinglin-traditional_mandarin_instructions
Language Models for Taiwanese Culture
✍️ Online Demo
•
🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon]
• 👨️ Yen-Ting Lin
Overview
Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications.
Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin.
Demo
A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.physics-ptbr
Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.Azure99_blossom-math-v4
BLOSSOM MATH V4
介绍
Blossom Math V4是基于Math23K和GSM8K衍生而来的中英双语数学对话数据集,适用于数学问题微调。
相比于blossom-math-v3,本版本完全使用GPT-4进行蒸馏,大幅提升了推理的一致性。
本数据集采用全量Math23K、GSM8K和翻译后的GSM8K的问题,随后调用gpt-4-0125-preview生成结果,并使用原始数据集中的答案对生成的结果进行验证,过滤掉错误答案,很大程度上保证了问题和答案的准确性。
本次发布了全量数据的25%,包含10K记录。
语言
中文和英文
数据集结构
每条数据代表一个完整的题目及答案,包含id、input、output、answer、dataset四个字段。
id:字符串,代表原始数据集中的题目id,与dataset字段结合可确定唯一题目。
input:字符串,代表问题。
output:字符串,代表gpt-4-0125-preview生成的答案。
answer:字符串,代表正确答案。… See the full description on the dataset page: https://huggingface.co/datasets/botp/Azure99_blossom-math-v4.MetaMathQA-40K-PTBRTradução do MetaMathQA-4k para portugues com NLLB 3.3b.
AVAINT-IMGolmo-igsm-arith
OLMo iGSM-Easy Arithmetic
This repository contains a frozen, evaluation-only release of the synthetic
mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo
evaluation code. It contains 750 examples: 250 examples at each target depth
2, 3, and 4.
This is an i-GSM-style task variant, not a claim to be an official release
of another dataset named iGSM. The olmo-igsm-arith name is used to make the
implementation provenance explicit.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.Azure99_blossom-wizard-v3
BLOSSOM WIZARD V3
介绍
Blossom Wizard V3是一个基于WizardLM_evol_instruct_V2衍生而来的中英双语指令数据集,适用于指令微调。
相比于blossom-wizard-v2,本版本完全使用GPT-4进行蒸馏。
本数据集从WizardLM_evol_instruct_V2中抽取了指令,首先将其翻译为中文并校验翻译结果,再使用指令调用gpt-4-0125-preview模型生成响应,并过滤掉包含自我认知以及拒绝回答的响应,以便后续对齐。此外,为了确保响应风格的一致性以及中英数据配比,本数据集还对未翻译的原始指令也进行了相同的调用,最终得到了1:1的中英双语指令数据。
相比直接对原始Wizard进行翻译的中文数据集,Blossom Wizard的一致性及质量更高。
本次发布了全量数据的50%,包含中英双语各10K,共计20K记录。
语言
以中文和英文为主。
数据集结构
每条数据代表一个完整的对话,包含id和conversations两个字段。… See the full description on the dataset page: https://huggingface.co/datasets/botp/Azure99_blossom-wizard-v3.chemistry-ptbr
Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.
