datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgiBotWorld-Beta
Key Features 🔑
1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours.
100+ real-world scenarios across 5 target domains.
Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots
200+ types of tasks:
Contact-rich manipulation
Long-horizon planning
Multi-robot collaboration
87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc.
Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.Beta-Hybrid-Interaction-SFTUpdesh_beta
📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages
NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines.
Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.jimei-fire-smoke-yolo-datasetMMScan-betaGridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.ASCII_Alphabet_Dataset_571_Fonts
Dataset Description
This dataset provides programmatically generated ASCII representations of the English alphabet rendered using 571 fonts from the PyFiglet library. Each letter (A–Z) is available in multiple typographic styles, resulting in a structured and high-variability dataset suitable for research, experimentation, and creative applications.
The dataset was created to support tasks involving text-based pattern recognition, synthetic data generation, typography analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/beta3/ASCII_Alphabet_Dataset_571_Fonts.LOVE-Agibot-BetaBeta-Code
Reactive AI / Beta Code
Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages.
Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories.
Original dataset
It's created from codeparrot datasets:
Python subsets from codeparrot/codeparrot-clean
other subsets from codeparrot/github-code-clean
BETA
The BETA dataset
Paper
BETA: A Benchmark Database Towards BCI Application (Full Text)
Summary
The Brain–Computer Interface (BCI) provides an alternative means of communication and has sparked growing interest in the past two decades. Specifically, for Steady-State Visual Evoked Potential (SSVEP)-based BCI (SSVEP-BCI), significant improvements have been made in frequency recognition methods and data sharing. However, the number of public databases in this… See the full description on the dataset page: https://huggingface.co/datasets/Bingchuan/BETA.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.Reverse-alpha-beta-no-outsidefinepdfs-edu-betaBeta-Hybrid-SMAT
Reactive AI / Beta Hybrid SMAT
Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models
beta-reasoningHistorical_Data_of_Ecuador_Stock_ExchangeHistorical Data of Ecuador's Stock Exchange
Unlock the latest financial trends with up-to-date data from the market
Context
The Guayaquil Stock Exchange (Bolsa de Valores de Guayaquil - BVG) and Quito Stock Exchange (Bolsa de Valores de Quito - BVQ) play a crucial role in Ecuador's financial markets, facilitating trading of stocks, bonds, and other securities. However, historical financial data from this exchange is often difficult to access in a structured and ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/beta3/Historical_Data_of_Ecuador_Stock_Exchange.Aloe-Beta-Medical-Collection
Aloe-Beta-Medical-Collection
Collection of curated datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available medical instruction tuning data sources (QA format). Most data samples correspond to single-turn QA pairs, while a small proportion contain multi-turn. All data sources are publicly available for… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-Medical-Collection.nurc_tts_betadetails_CausalLM__34b-beta
Dataset Card for Evaluation run of CausalLM/34b-beta
Dataset automatically created during the evaluation run of model CausalLM/34b-beta.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CausalLM__34b-beta.NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.VyvoTTS-EN-Beta-DPO-samples
VyvoTTS EN-Beta — 2,000 automatic DPO pairs
Exactly 2,000 unique target texts and chosen/rejected pairs, generated by
Vyvo/VyvoTTS-EN-Beta at revision 70b37a5bfdbdc2f478515837081048aac63f909e. Each row embeds the actual 24 kHz reference,
chosen and rejected audio, raw prompt and completion codec IDs, transcripts,
WER/CER, DNSMOS P.835, sampling settings, seeds and waveform SHA-256 checksums.
Both candidate waveforms are actual model outputs; no artificial corruption.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/VyvoTTS-EN-Beta-DPO-samples.arc-agi-impabs-dpolr1e-7-beta0.01-classifiersft5e-7HuggingFaceH4__zephyr-7b-beta-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.ReverseBass-Beta-Statusalpha_beta_interventionNEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_4Dataset-Beta_Lactamase-PEER
Description
β-Lactamase Prediction studies the activity among first-order mutants of the TEM-1 beta-lactamase protein.
Splits
Protein Format: AA sequence
The dataset is from PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence Understanding. We follow the original data splits, with the number of training, validation and test set shown below:
Train: 4158
Valid: 520
Test: 520
Label
The target y ∈ R is the experimentally tested fitness score… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Beta_Lactamase-PEER.Aloe-Beta-DPO
Aloe-Beta-Medical-Collection
Collection of curated DPO datasets used to align Aloe-Beta.
Dataset Details
Dataset Description
The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data:
Medical preference data: TsinghuaC3I/UltraMedical-Preference
General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.
