datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arkitscenes-preprocessedpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-Llama-3.2-3B-Instruct-bonsentiment_analysis_preprocessed_datasetBrief idea about dataset:
This dataset is designed for a Text Classification to be specific Multi Class Classification, inorder to train a model (Supervised Learning) for Sentiment Analysis.
Also to be able retrain the model on the given feedback over a wrong predicted sentiment this dataset will help to manage those things using Other Features.
Main Features
text
labels
This feature variable has all sort of texts, sentences, tweets, etc.
This target variable contains 3 types of… See the full description on the dataset page: https://huggingface.co/datasets/prasadsawant7/sentiment_analysis_preprocessed_dataset.preprocessed-full-aime_2023-n256-Qwen2.5-3B-Instruct-bonpreprocessed-full-MATH-500-n256-Phi-4-mini-instruct-bonpreprocessed-full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bonpreprocessed-full-gsm8k-private-n256-Phi-4-mini-instruct-bonpreprocessed-full-gsm8k-private-Qwen2.5-3B-Instruct-bonpreprocessed-full-aime_2026-n256-Qwen3-4B-Instruct-2507-bonpreprocessed-full-aime_2023-n256-Qwen3-4B-Instruct-2507-bonav_sql_preprocessed_data
Dataset Card for Preprocessed Text-to-SQL Benchmarks
This repository contains preprocessed data for several text-to-SQL benchmarks, as presented in the paper AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views.
The official code for the AV-SQL framework can be found on GitHub: pminhtam/AV-SQL.
Dataset Summary
This repository contains preprocessed data for several text-to-SQL benchmarks:
BIRD
KaggleDBQA
Spider
sciencebenchmark
BEAVER
Spider2-Lite… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/av_sql_preprocessed_data.preprocessed-full-aime_2024-n256-Qwen3-4B-Instruct-2507-bonpreprocessed-full-math-private-n256-Qwen2.5-3B-Instruct-bonPreprocessed_Neu3Dpreprocessed-full-aime_2024-n256-Qwen2.5-3B-Instruct-bonpreprocessed-full-aime_2026-n256-Phi-4-mini-instruct-bonmeajor_cleaned_preprocessed
Dataset Card for MeAJOR: Merged email Assets from Joint Open-source Repositories
The MeAJOR dataset was created with multi-source legitimate emails and malicious emails from different attack campaigns to enable a more reliable phishing detection. MeAJOR provides a total of 108685 data samples, in preprocessed format with engineered features suitable for the training, validation, and testing of ML and DL models.
Dataset Details
The dataset was created at GECAD… See the full description on the dataset page: https://huggingface.co/datasets/simlab-vs/meajor_cleaned_preprocessed.preprocessed-full-aime_2026-n256-Qwen2.5-3B-Instruct-bonpreprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.dpv2v-preprocessedsakuga_preprocessedpreprocessed-har-datasets
Ready-to-Use Preprocessed HAR Datasets
This dataset repository provides ready-to-use, preprocessed Human Activity Recognition (HAR) datasets. The released files are already windowed, split, and stored as NumPy arrays, so users can download them and run model comparisons directly.
The goal is to make HAR model comparison easier by releasing a consistent set of preprocessed windows, split files, labels, and metadata. The collection covers smartphone-based sensing, wearable IMU… See the full description on the dataset page: https://huggingface.co/datasets/shenjianmozhu/preprocessed-har-datasets.preprocessed-full-aime_2024-n256-Phi-4-mini-instruct-bonragbench-dual-clf-preprocessedpreprocessed_issues
Dataset Card for "preprocessed_issues"
More Information needed
preprocessed_stars
Dataset Card for "preprocessed_stars"
More Information needed
log-analysis-hdfs-preprocessedspeech_commands_mel_preprocessedDiscord-Dialogues-Preprocessed-Luna-Protocol
Discord-Dialogues-Preprocessed-Luna-Protocol is a preprocessed fork of mookiezi/Discord-Dialogues, adapted for fine-tuning Qwen2.5-family models as part of the Luna Protocol project.
This dataset contains anonymized Discord conversations for training and evaluating realistic conversational AI models in a ChatML-friendly format. It is derived directly from mookiezi/Discord-Dialogues with two targeted preprocessing steps applied (see below) — the underlying conversations, filtering pipeline… See the full description on the dataset page: https://huggingface.co/datasets/fox3000foxy/Discord-Dialogues-Preprocessed-Luna-Protocol.
