datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-roberta-baseecqa_model_generate_robertalamus-roberts-court-legal-arguments
LAMUS: Roberts Court Legal Arguments (2005-2025)
The Current Supreme Court Era - Chief Justice John Roberts
📋 Dataset Description
This dataset contains 362,891 sentences from U.S. Supreme Court opinions during the Roberts Court era (2005-2025), automatically labeled with legal argument categories. This represents the current Supreme Court under Chief Justice John G. Roberts Jr.
Why Roberts Court?
The Roberts Court is particularly significant for… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-roberts-court-legal-arguments.danish-car-marketplace-dataset
Danish used car listings — raw dataset
~28,000 car listings scraped from the danish market as of june 2026. this is the raw, messy version — unit suffixes mixed into values, danish decimal separators, empty fields, the whole thing. the point is to have something real to practice data cleaning and analysis on, not a tidy kaggle dataset that does half the work for you.
if you want to jump straight into modeling with this data, check out danish-used-car-price-prediction — that repo… See the full description on the dataset page: https://huggingface.co/datasets/robertcaliforniadk/danish-car-marketplace-dataset.fast-gift-532182
fast-gift-532182
Synthetic sensors test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Sable-Robert/fast-gift-532182.parsed-dataset-xlm-robertaroberta-leadership-dataset-finetunechess-roberta-pretraining-sansconfigs:
config_name: default
data_files:
split: train
path: train/*.csv
split: eval
path: eval/*.csv
ml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.xai_gab_multip_robertaroberta_embeddings_isearRoBERTa_eval_dataxlm-roberta-large-dfadolescent_behaviors_and_experiences_surveyThis dataset contains the underlying data for the Adolescent Behaviors and Experiences Survey dataset originally sourced from the CDC website.
Original dataset description can be found here: https://web.archive.org/web/20241222113221/https://www.cdc.gov/abes/data/index.html
RobertFrostRobertadatasettestStadium_RoBERTa_eval
