datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleSafetyTestsbert_fine_tune_medical_dataBERT-bitcoin-sentiment-assets
BERT-bitcoin-sentiment — large assets
Files that exceed GitHub's 100 MB limit, split out of the research repository at
https://github.com/Kosmosas. Fetch them into place with:
python scripts/download_assets.py
Contents
File
Size
What it is
backtesting/data/BTCUSDT-last.csv
~411 MB
1-minute BTCUSDT OHLCV + volume. The extended snapshot, running to Nov 2025; used by forecasting and backtesting.
weights_comparison_and_derive/BTCUSDT.csv
~381 MB
1-minute… See the full description on the dataset page: https://huggingface.co/datasets/Kosmosas/BERT-bitcoin-sentiment-assets.Realistic_LJP_BertSumtest-datasetbert-amz-csafety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.Bertbert-dataset
Road Traffic Act QA Dataset
This dataset is automatically generated question-answer pairs based on the official Road Traffic Act (Republic of Korea). The dataset is designed to support RAG (Retrieval-Augmented Generation) and legal NLP tasks.
Dataset Summary
Source: Road Traffic Act (English version)
Task: Question Answering (QA)
Type: Automatically generated by GPT-4o with custom multi-QA prompt
Size: 2,000+ QA pairs
Language: English
Format: CSV (Question, Answer)… See the full description on the dataset page: https://huggingface.co/datasets/YeahOuts/bert-dataset.TW-clasification-BERTbertqssp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173)
TW-Test-BERTmultilingual-bert-toc-95k-dataset
Dataset Details
Dataset Description
Contains line-by-line sequences from human-annotated legal/government documents and their corresponding labels.
Line-by-line examples derived from DocLayNet dataset
Dataset Creation
Notebook displaying how dataset was created can be accessed here
Training-Bert-Model-Analysiscustomer_feedback_analysis_bert_dataset
Customer Feedback Analysis
Description: Classify customer feedback based on sentiment and topic to identify improvement areas and strengthen customer engagement.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/customer_feedback_analysis_bert"
model = AutoModelForSequenceClassification.from_pretrained(model_name)… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/customer_feedback_analysis_bert_dataset.DebateLLMsFILT_BERT_TESTFILT_BERT_TRAINmaritime-berth-crane-productivity-coherence-risk-v0.1What this repo is for
Detect when berth use stops matching crane output.
You use it to flag:
hidden capacity loss
queue growth before official congestion
under-crewing or equipment drag
weather combined with resource mismatch
Why it matters
Berth looks busy long before throughput collapses.
Bert-distilbertbert-sentiment-analysisbert_englishBERT_trainevalVirBiCla-training
Dataset Card for VirBiCla-training
VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics.
This dataset is a support dataset for training the base ML model.
Dataset Details
Dataset Sources [optional]
Repository: GitHub repository for VirBiCla
Uses
This dataset is intended as support for training the base VirBiCla model
Dataset Structure
Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.advertisement_cap_on_banner_classification_bert_dataset
Advertisement Cap on Banner Classification
Description: Automatically classify and assign appropriate advertisement cap to banners to streamline manufacturing and delivery processes.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/advertisement_cap_on_banner_classification_bert"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/advertisement_cap_on_banner_classification_bert_dataset.autotrain-data-test-bertautotrain-data-nlp-bert-ner-testingbertCritiQ_BERT
