CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agibot-world /AgiBotWorld-Betagated Key Features 🔑 1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. 100+ real-world scenarios across 5 target domains. Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots 200+ types of tasks: Contact-rich manipulation Long-horizon planning Multi-robot collaboration 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.textother100M<n<1B79 likes93k downloads1y agoHugging Face02ReactiveAI /Beta-Pre-Train-Corpus Reactive AI / Beta Pre-Train Corpus Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets, and code in different programming languages. 2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens Subsets & original datasets FineWeb-Edu fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.tabular100M<n<1B0 likes20k downloads7mo agoHugging Face03ReactiveAI /Beta-Hybrid-Interaction-SFTtext10M<n<100M0 likes17k downloads7mo agoHugging Face04microsoft /Updesh_beta 📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines. Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.textquestion-answering1M<n<10M16 likes8.6k downloads8mo agoHugging Face05betasecond /jimei-fire-smoke-yolo-datasetimage1 likes5.2k downloads6mo agoHugging Face063dllm /MMScan-betatext1M<n<10M1 likes4.3k downloads2y agoHugging Face07beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.1k downloads7mo agoHugging Face08beta3 /ASCII_Alphabet_Dataset_571_Fonts Dataset Description This dataset provides programmatically generated ASCII representations of the English alphabet rendered using 571 fonts from the PyFiglet library. Each letter (A–Z) is available in multiple typographic styles, resulting in a structured and high-variability dataset suitable for research, experimentation, and creative applications. The dataset was created to support tasks involving text-based pattern recognition, synthetic data generation, typography analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/beta3/ASCII_Alphabet_Dataset_571_Fonts.texttext-classification1 likes1.1k downloads7mo agoHugging Face09EvanSirius /LOVE-Agibot-Betaimage1M<n<10M0 likes730 downloads1y agoHugging Face10ReactiveAI /Beta-Code Reactive AI / Beta Code Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages. Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories. Original dataset It's created from codeparrot datasets: Python subsets from codeparrot/codeparrot-clean other subsets from codeparrot/github-code-clean texttext-generation1M<n<10M0 likes521 downloads9mo agoHugging Face11Bingchuan /BETA The BETA dataset Paper BETA: A Benchmark Database Towards BCI Application (Full Text) Summary The Brain–Computer Interface (BCI) provides an alternative means of communication and has sparked growing interest in the past two decades. Specifically, for Steady-State Visual Evoked Potential (SSVEP)-based BCI (SSVEP-BCI), significant improvements have been made in frequency recognition methods and data sharing. However, the number of public databases in this… See the full description on the dataset page: https://huggingface.co/datasets/Bingchuan/BETA.documentn<1K1 likes475 downloads1y agoHugging Face12latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K9 likes330 downloads2mo agoHugging Face13PersonaBias /Reverse-alpha-beta-no-outsidetabulartext-classification10K<n<100K0 likes295 downloads2mo agoHugging Face14ReactiveAI /finepdfs-edu-betatabular10M<n<100M0 likes233 downloads11mo agoHugging Face15ReactiveAI /Beta-Hybrid-SMAT Reactive AI / Beta Hybrid SMAT Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models text100K<n<1M0 likes199 downloads5mo agoHugging Face16ReactiveAI /beta-reasoningtext100K<n<1M0 likes173 downloads3mo agoHugging Face17beta3 /Historical_Data_of_Ecuador_Stock_ExchangeHistorical Data of Ecuador's Stock Exchange Unlock the latest financial trends with up-to-date data from the market Context The Guayaquil Stock Exchange (Bolsa de Valores de Guayaquil - BVG) and Quito Stock Exchange (Bolsa de Valores de Quito - BVQ) play a crucial role in Ecuador's financial markets, facilitating trading of stocks, bonds, and other securities. However, historical financial data from this exchange is often difficult to access in a structured and ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/beta3/Historical_Data_of_Ecuador_Stock_Exchange.tabular100K<n<1M3 likes172 downloads19d agoHugging Face18HPAI-BSC /Aloe-Beta-Medical-Collection Aloe-Beta-Medical-Collection Collection of curated datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available medical instruction tuning data sources (QA format). Most data samples correspond to single-turn QA pairs, while a small proportion contain multi-turn. All data sources are publicly available for… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-Medical-Collection.textquestion-answering100K<n<1M4 likes171 downloads1y agoHugging Face19nilc-nlp /nurc_tts_betaaudio100K<n<1M0 likes169 downloads7mo agoHugging Face20OALL /details_CausalLM__34b-beta Dataset Card for Evaluation run of CausalLM/34b-beta Dataset automatically created during the evaluation run of model CausalLM/34b-beta. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CausalLM__34b-beta.tabular100K<n<1M0 likes163 downloads2y agoHugging Face21happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5textn<1K0 likes159 downloads1y agoHugging Face22HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes154 downloads10mo agoHugging Face23kadirnar /VyvoTTS-EN-Beta-DPO-samples VyvoTTS EN-Beta — 2,000 automatic DPO pairs Exactly 2,000 unique target texts and chosen/rejected pairs, generated by Vyvo/VyvoTTS-EN-Beta at revision 70b37a5bfdbdc2f478515837081048aac63f909e. Each row embeds the actual 24 kHz reference, chosen and rejected audio, raw prompt and completion codec IDs, transcripts, WER/CER, DNSMOS P.835, sampling settings, seeds and waveform SHA-256 checksums. Both candidate waveforms are actual model outputs; no artificial corruption.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/VyvoTTS-EN-Beta-DPO-samples.audiotext-to-speech1K<n<10K1 likes138 downloads9d agoHugging Face24Asap7772 /arc-agi-impabs-dpolr1e-7-beta0.01-classifiersft5e-7tabular10K<n<100K0 likes130 downloads1y agoHugging Face25open-llm-leaderboard /HuggingFaceH4__zephyr-7b-beta-detailsgated Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.tabular10K<n<100K0 likes117 downloads2y agoHugging Face26VortexSamples /ReverseBass-Beta-Statustextn<1K0 likes117 downloads22d agoHugging Face27PersonaBias /alpha_beta_interventiontabulartext-classification10K<n<100K0 likes113 downloads4d agoHugging Face28happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_4textn<1K0 likes97 downloads1y agoHugging Face29SaProtHub /Dataset-Beta_Lactamase-PEER Description β-Lactamase Prediction studies the activity among first-order mutants of the TEM-1 beta-lactamase protein. Splits Protein Format: AA sequence The dataset is from PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence Understanding. We follow the original data splits, with the number of training, validation and test set shown below: Train: 4158 Valid: 520 Test: 520 Label The target y ∈ R is the experimentally tested fitness score… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Beta_Lactamase-PEER.text1K<n<10K2 likes96 downloads2y agoHugging Face30HPAI-BSC /Aloe-Beta-DPO Aloe-Beta-Medical-Collection Collection of curated DPO datasets used to align Aloe-Beta. Dataset Details Dataset Description The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data: Medical preference data: TsinghuaC3I/UltraMedical-Preference General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.textquestion-answering100K<n<1M2 likes90 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.