datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.Beta-Code
Reactive AI / Beta Code
Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages.
Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories.
Original dataset
It's created from codeparrot datasets:
Python subsets from codeparrot/codeparrot-clean
other subsets from codeparrot/github-code-clean
Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.Aloe-Beta-DPO
Aloe-Beta-Medical-Collection
Collection of curated DPO datasets used to align Aloe-Beta.
Dataset Details
Dataset Description
The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data:
Medical preference data: TsinghuaC3I/UltraMedical-Preference
General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.synthux-visual-tree-0.1-beta
SynthUX Visual Computer-Use Trajectories
This dataset contains synthetic computer-use episodes generated by SynthUX and executed inside lightweight web desktop environments instead of full VMs.
Each episode has:
goal
instruction
tree: the concrete Node Tree-style grammar expansion with executable terminal affordance leaves
trajectory: observed simulator recording with input events, state observations, alignment, and media refs
rendered screenshots referenced from trajectory… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-visual-tree-0.1-beta.3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.edgeimci-beta0-1k-multitask-enriched-2258-v1
EdgeIMCI Beta0-1K Multitask Enriched 2258
Dataset summary
This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow.
The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.Test
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/beta42ZH/Test.jimba-instuction-1k-betacyberagent/calm2-7b-chatの出力を人手でチェック・修正することで作成した日本語Instructionデータセットです。
詳しくはこちらの記事を御覧ください。
https://zenn.dev/kendama/articles/dc727218a2eae6
instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawizephyr-7b-beta-invoices
Zephyr-7B-Beta Customer Support Chatbot
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Introduction
Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.011_beta_search1Atma7-Beta
Dataset Card for Atma7-Beta
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma7-Beta.betawi-v0Synthetic Betawi Language dataset, generated by GPT-4o: Betawi v0 (Alpha)
Betawi v0 is a synthetic dataset created using GPT-4o, consisting of over 1,000 instruction-output pairs across a range of topics, structured in JSON format. It follows the Alpaca dataset format and is designed for fine-tuning large language models (LLMs) to enhance LLMs understanding of Bahasa Betawi.
Version Alpha
.hf-sanitized.hf-sanitized-uZg6qHEzHPlvxjDu8mVLI h1 { font-size: 36px; color: #000000;… See the full description on the dataset page: https://huggingface.co/datasets/risangpanggalih/betawi-v0.betadata8
