CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Askhat777 /GVP_Bot_State_Newtabularn<1K0 likes946 downloads37m agoHugging Face02junbrro /egopi_latal_openarm_bottletabularn<1K0 likes783 downloads2mo agoHugging Face03librarian-bots /paper-recommendations-v2text10K<n<100K16 likes748 downloads12h agoHugging Face04botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes500 downloads3y agoHugging Face05botay /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.documenttable-question-answering10K<n<100K0 likes495 downloads5mo agoHugging Face06giskard-bot /evaluator-leaderboardtabularn<1K0 likes452 downloads2y agoHugging Face07JupiterLLM /fineweb_2_500k_both_deduplicatedtabular1M<n<10M0 likes452 downloads1y agoHugging Face08preethamvj /bottleneck-oracle-graphstabularn<1K0 likes216 downloads5mo agoHugging Face09botp /COIG-CQIA COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning Dataset Details Dataset Description 欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。 Welcome to the COIG-CQIA… See the full description on the dataset page: https://huggingface.co/datasets/botp/COIG-CQIA.textquestion-answering10K<n<100K0 likes204 downloads2y agoHugging Face10bep40 /zalo-bot-registry bep40/zalo-bot-registry Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('bep40/zalo-bot-registry') textn<1K0 likes200 downloads8d agoHugging Face11JQL-AI /Fineweb_2_500k_bothtabular10M<n<100M0 likes188 downloads2y agoHugging Face12botbotrobotics /Cabra3kO conjunto de dados Cabra é uma coleção ampla e diversificada de 3.000 entradas ou conjuntos de perguntas e respostas (QA) sobre o Brasil. Inclui tópicos variados como história, política, geografia, cultura, cinema, esportes, ciência e tecnologia, governo e muito mais. Este conjunto foi cuidadosamente elaborado e selecionado pela nossa equipe, garantindo alta qualidade e relevância para estudos e aplicações relacionadas ao Brasil. Detalhes do Conjunto de Dados: Tamanho: 3.000 conjuntos de QA… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/Cabra3k.text1K<n<10K6 likes161 downloads2y agoHugging Face13Ai-robot-001 /aloha_static_bottle_pick_uptabularn<1K0 likes149 downloads2y agoHugging Face14botisan-ai /cantonese-mandarin-translations Dataset Card for cantonese-mandarin-translations Dataset Summary This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese). Supported Tasks and Leaderboards N/A Languages Cantonese (yue) Simplified Chinese (zh-CN) Dataset Structure JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.texttranslation10K<n<100K31 likes143 downloads3y agoHugging Face15botp /WordVoice-5A WordVoice-5A Dataset 🚀 A Large-Scale Bilingual Word-level Five-Annotation Dataset for WordVoice 📖 Dataset Description / 数据集简介 WordVoice-5A is a large-scale bilingual (Mandarin and English) dataset containing approximately 4.7k hours of speech with fine-grained word-level acoustic annotations, designed for high-precision controllable Text-to-Speech (TTS). It addresses the scarcity of large-scale, high-quality word-aligned datasets with explicit acoustic… See the full description on the dataset page: https://huggingface.co/datasets/botp/WordVoice-5A.text1M<n<10M0 likes140 downloads2mo agoHugging Face16WDong /so101-sword-on-stand-clean32-bothtrim-v2 SO-101 sword-on-stand clean32, physically both-end trimmed This is the corrected, portable LeRobot v2.1 derivative for the task: Pick up the sword and place it on the sword stand. Dataset contract 32 episodes / 6,974 frames / 30 Hz. Train episodes: 0..29 (6,582 frames). Validation episodes: 30, 31 (392 frames). Fixed and wrist RGB cameras, 640x480, H.264, 30 FPS. observation.state and stored action are calibrated 6D absolute joint positions. Only episodes 0..29… See the full description on the dataset page: https://huggingface.co/datasets/WDong/so101-sword-on-stand-clean32-bothtrim-v2.tabularn<1K0 likes110 downloads21d agoHugging Face17deepflame-bot /pi-publish Coding agent session traces for deepflame-bot/pi-publish This dataset contains redacted coding agent session traces collected while working on https://github.com/xke-b/efno-chem-kinetics.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/deepflame-bot/pi-publish.tabulartext-generationn<1K0 likes98 downloads5mo agoHugging Face18taskydata /botbots-tod(https://github.com/radi-cho/botbots/tree/main/tod) tabularn<1K0 likes97 downloads2y agoHugging Face19ryankim17920 /open-bottleneck-ranklong27b-slurm-364982-rollouts Open Bottleneck RankLong 27B — Slurm array 364982 Compact rollout evidence archived from completed Slurm array 364982. Config Files / steps Records JSONL bytes Note rank_a40 60 (1–60) 15,360 66,445,670 Complete local rollout evidence rank_a80 54 (1–54) 13,824 59,386,451 Includes the cancelled arm's final dumped step (54.jsonl) Each JSONL record contains input, output, gts, score, acc, response_length, grouprel_reward, and step. Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.tabular10K<n<100K0 likes88 downloads2mo agoHugging Face20cloudfan /intern-bottle-lerobot bottle: robot demonstrations Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat. LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec") sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.tabularn<1K0 likes74 downloads2d agoHugging Face21Shiki42 /ctr-pick-dual-bottles-original-20260919 Pick Dual Bottles Original — shared50 scene cohort This LeRobot v3 release contains 50 successful simulated demonstrations and 8,185 action rows at25FPS. Every source seed occurs exactly once. The source seed set matches the current CTR Q1–Q3 Concurrent, CTR, Sequential, Mixed, Left-first and Right-first datasets. Pair by retime.source_seed, not episode index: composition datasets may have different ordering. Mask limitation: retime.left_idle and retime.right_idle are boolean… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-pick-dual-bottles-original-20260919.tabularn<1K0 likes72 downloads3d agoHugging Face22botbotrobotics /PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b. *Somente para demonstração e pesquisa. Proibido para uso comercial. PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo. *For demonstration and research purposes only. Commercial use prohibited. text10K<n<100K7 likes69 downloads3y agoHugging Face23botp /yentinglin-traditional_mandarin_instructions Language Models for Taiwanese Culture ✍️ Online Demo • 🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon] • 👨️ Yen-Ting Lin Overview Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications. Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin. Demo A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.texttext-generation100K<n<1M0 likes57 downloads3y agoHugging Face24botbotrobotics /physics-ptbr Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.texttext-generation10K<n<100K2 likes54 downloads3y agoHugging Face25botp /Azure99_blossom-math-v4 BLOSSOM MATH V4 介绍 Blossom Math V4是基于Math23K和GSM8K衍生而来的中英双语数学对话数据集,适用于数学问题微调。 相比于blossom-math-v3,本版本完全使用GPT-4进行蒸馏,大幅提升了推理的一致性。 本数据集采用全量Math23K、GSM8K和翻译后的GSM8K的问题,随后调用gpt-4-0125-preview生成结果,并使用原始数据集中的答案对生成的结果进行验证,过滤掉错误答案,很大程度上保证了问题和答案的准确性。 本次发布了全量数据的25%,包含10K记录。 语言 中文和英文 数据集结构 每条数据代表一个完整的题目及答案,包含id、input、output、answer、dataset四个字段。 id:字符串,代表原始数据集中的题目id,与dataset字段结合可确定唯一题目。 input:字符串,代表问题。 output:字符串,代表gpt-4-0125-preview生成的答案。 answer:字符串,代表正确答案。… See the full description on the dataset page: https://huggingface.co/datasets/botp/Azure99_blossom-math-v4.texttext-generation10K<n<100K2 likes52 downloads2y agoHugging Face26botbotrobotics /MetaMathQA-40K-PTBRTradução do MetaMathQA-4k para portugues com NLLB 3.3b. text10K<n<100K4 likes48 downloads3y agoHugging Face27botintel-community /AVAINT-IMGimageimage-to-text100K<n<1M2 likes47 downloads2y agoHugging Face28Mihara-bot /olmo-igsm-arith OLMo iGSM-Easy Arithmetic This repository contains a frozen, evaluation-only release of the synthetic mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo evaluation code. It contains 750 examples: 250 examples at each target depth 2, 3, and 4. This is an i-GSM-style task variant, not a claim to be an official release of another dataset named iGSM. The olmo-igsm-arith name is used to make the implementation provenance explicit. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.tabularquestion-answeringn<1K0 likes41 downloads2mo agoHugging Face29botp /Azure99_blossom-wizard-v3 BLOSSOM WIZARD V3 介绍 Blossom Wizard V3是一个基于WizardLM_evol_instruct_V2衍生而来的中英双语指令数据集,适用于指令微调。 相比于blossom-wizard-v2,本版本完全使用GPT-4进行蒸馏。 本数据集从WizardLM_evol_instruct_V2中抽取了指令,首先将其翻译为中文并校验翻译结果,再使用指令调用gpt-4-0125-preview模型生成响应,并过滤掉包含自我认知以及拒绝回答的响应,以便后续对齐。此外,为了确保响应风格的一致性以及中英数据配比,本数据集还对未翻译的原始指令也进行了相同的调用,最终得到了1:1的中英双语指令数据。 相比直接对原始Wizard进行翻译的中文数据集,Blossom Wizard的一致性及质量更高。 本次发布了全量数据的50%,包含中英双语各10K,共计20K记录。 语言 以中文和英文为主。 数据集结构 每条数据代表一个完整的对话,包含id和conversations两个字段。… See the full description on the dataset page: https://huggingface.co/datasets/botp/Azure99_blossom-wizard-v3.texttext-generation10K<n<100K2 likes40 downloads2y agoHugging Face30botbotrobotics /chemistry-ptbr Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.texttext-generation10K<n<100K2 likes39 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.