CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.1k downloads7mo agoHugging Face02ReactiveAI /Beta-Code Reactive AI / Beta Code Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages. Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories. Original dataset It's created from codeparrot datasets: Python subsets from codeparrot/codeparrot-clean other subsets from codeparrot/github-code-clean texttext-generation1M<n<10M0 likes521 downloads9mo agoHugging Face03latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K9 likes330 downloads2mo agoHugging Face04HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes154 downloads10mo agoHugging Face05HPAI-BSC /Aloe-Beta-DPO Aloe-Beta-Medical-Collection Collection of curated DPO datasets used to align Aloe-Beta. Dataset Details Dataset Description The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data: Medical preference data: TsinghuaC3I/UltraMedical-Preference General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.textquestion-answering100K<n<1M2 likes90 downloads1y agoHugging Face06jacob-valdez /synthux-visual-tree-0.1-beta SynthUX Visual Computer-Use Trajectories This dataset contains synthetic computer-use episodes generated by SynthUX and executed inside lightweight web desktop environments instead of full VMs. Each episode has: goal instruction tree: the concrete Node Tree-style grammar expansion with executable terminal affordance leaves trajectory: observed simulator recording with input events, state observations, alignment, and media refs rendered screenshots referenced from trajectory… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-visual-tree-0.1-beta.imagevisual-question-answering10K<n<100K0 likes84 downloads4mo agoHugging Face07beta3 /3M_Academic_Papers_Titles_and_Abstracts Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts 📋 Overview This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.texttext-classification1M<n<10M2 likes71 downloads1y agoHugging Face08Nini0la /edgeimci-beta0-1k-multitask-enriched-2258-v1 EdgeIMCI Beta0-1K Multitask Enriched 2258 Dataset summary This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow. The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.texttext-generation1K<n<10K0 likes38 downloads4d agoHugging Face09beta42ZH /Test Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/beta42ZH/Test.texttext-generation100K<n<1M0 likes37 downloads9mo agoHugging Face10Kendamarron /jimba-instuction-1k-betacyberagent/calm2-7b-chatの出力を人手でチェック・修正することで作成した日本語Instructionデータセットです。 詳しくはこちらの記事を御覧ください。 https://zenn.dev/kendama/articles/dc727218a2eae6 texttext-generation1K<n<10K6 likes31 downloads2y agoHugging Face11bryandts /instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawitexttext-generation100K<n<1M0 likes25 downloads2y agoHugging Face12erfanvaredi /zephyr-7b-beta-invoices Zephyr-7B-Beta Customer Support Chatbot This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Introduction Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.texttext-classification10K<n<100K2 likes20 downloads3y agoHugging Face13ExplodeMediaG /011_beta_search1gatedtext-classification10B<n<100B0 likes9 downloads2y agoHugging Face14HappyAIUser /Atma7-Beta Dataset Card for Atma7-Beta This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks. Dataset Description The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains: An instruction that specifies the task An optional input providing context A detailed output that addresses the instruction Usage This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma7-Beta.texttext-generation10K<n<100K0 likes8 downloads1y agoHugging Face15risangpanggalih /betawi-v0Synthetic Betawi Language dataset, generated by GPT-4o: Betawi v0 (Alpha) Betawi v0 is a synthetic dataset created using GPT-4o, consisting of over 1,000 instruction-output pairs across a range of topics, structured in JSON format. It follows the Alpaca dataset format and is designed for fine-tuning large language models (LLMs) to enhance LLMs understanding of Bahasa Betawi. Version Alpha .hf-sanitized.hf-sanitized-uZg6qHEzHPlvxjDu8mVLI h1 { font-size: 36px; color: #000000;… See the full description on the dataset page: https://huggingface.co/datasets/risangpanggalih/betawi-v0.texttext-generation1K<n<10K1 likes3 downloads2y agoHugging Face16Menoviar28 /betadata8text-generation100K<n<1M0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.