datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.Reverse-alpha-beta-no-outsideHistorical_Data_of_Ecuador_Stock_ExchangeHistorical Data of Ecuador's Stock Exchange
Unlock the latest financial trends with up-to-date data from the market
Context
The Guayaquil Stock Exchange (Bolsa de Valores de Guayaquil - BVG) and Quito Stock Exchange (Bolsa de Valores de Quito - BVQ) play a crucial role in Ecuador's financial markets, facilitating trading of stocks, bonds, and other securities. However, historical financial data from this exchange is often difficult to access in a structured and ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/beta3/Historical_Data_of_Ecuador_Stock_Exchange.alpha_beta_interventionDataset-Beta_Lactamase-PEER
Description
β-Lactamase Prediction studies the activity among first-order mutants of the TEM-1 beta-lactamase protein.
Splits
Protein Format: AA sequence
The dataset is from PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence Understanding. We follow the original data splits, with the number of training, validation and test set shown below:
Train: 4158
Valid: 520
Test: 520
Label
The target y ∈ R is the experimentally tested fitness score… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Beta_Lactamase-PEER.3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.github_fetch_huggingface_terminal_9134_x9v2m6_source_beta
Beta Support Conversations
Anonymized customer support conversation transcripts.
Overview
Dataset ID: SRC-BETA
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Author: Nova Data Engineering
Origin: Community tech-support forum public dump (2022-2024)
Records: 18,200
Product Line: Customer Feedback Analytics
License
Apache License 2.0
Contents
Anonymized support conversation transcripts with timestamps.
zephyr-7b-beta-invoices
Zephyr-7B-Beta Customer Support Chatbot
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Introduction
Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.ProtST-BetaLactamaseProtST-BetaLactamasebeta_1100_0.02_sentencesbeta_1100_0.5_sentencesProtST-BetaLactamasegemma_ours_beta_0.7_0.1dataset-20251216-beta
dataset-20251216-beta
Created on: 2025-12-16T04:30:36.804894+00:00
Session ID: 2025-12-16T04:30:36.804894+00:00-9747
beta-lactamase
Beta-Lactamase Fitness Prediction Dataset
This dataset is designed for training and evaluating machine learning models on protein fitness landscape prediction. It specifically focuses on the beta-lactamase enzyme, a primary driver of antibiotic resistance in bacteria.
Dataset Overview
The primary objective of this dataset is to map protein sequences to their functional activity. In the context of beta-lactamase, this functional activity—often referred to as… See the full description on the dataset page: https://huggingface.co/datasets/hazemessam/beta-lactamase.beta_1100_0.05_sentencesgemma_ours_beta_0.7_0.08autotrain-data-teste-betabeta_1100_0.08_sentencesgemma_ours_beta_0.7_0.03Mistral_7b_sft_beta_Samplingsthe-home-dataset-beta-1beta_1100_0.01_sentencesdataset-beta3dataset-20251223-beta-two
dataset-20251223-beta-two
Created on: 2025-12-23T12:51:51.166514+00:00
Session ID: 2025-12-23T12:51:51.166514+00:00-5184
beta_1100_0.2_sentencesbeta_1100_0.3_sentencesbeta_1100_1.0_sentences
