datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bls_cpi
Changelog
2025-01-18
I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code.
I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code.
I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site.
2025-01-12
While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.needle-threading
Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?
Dataset Summary
As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. Although the development of longer context models has seen rapid gains recently, our understanding of how effectively they use their context has not kept pace.
To address this, we conduct a set of retrieval experiments designed to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/needle-threading.chess-roberta-baseSTOP
🛑 STOP
This is the repository for STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions, a dataset comprised of 450 offensive progressions designed to target evolving scenarios of bias and quanitfy the threshold of appropriateness. This work was published in the 2024 Main Conference on Empirical Methods in Natural Language Processing and was honoured with the Social Impact Award.
Authors: Robert Morabito, Sangmitra Madhusudan, Tyler McDonald… See the full description on the dataset page: https://huggingface.co/datasets/Robert-Morabito/STOP.agent-traces
Agent Trace Dataset
Generated by build_hf_dataset.py. Each subset is one benchmark; rows are per-task trace records with score, trace, tool stats, and a link to the full trace files under trace_data/<benchmark>/<row_id>/.
roberta-largasstandard-chess-games-high-eloeval_caminbox1_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 1,
"total_frames": 3145,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/eval_caminbox1_smolvla.test_data_roberta_base_6_raceumbra02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra02.robertsjewelscamera_in_box4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 21,
"total_frames": 37097,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box4.lamus-roberts-court-legal-arguments
LAMUS: Roberts Court Legal Arguments (2005-2025)
The Current Supreme Court Era - Chief Justice John Roberts
📋 Dataset Description
This dataset contains 362,891 sentences from U.S. Supreme Court opinions during the Roberts Court era (2005-2025), automatically labeled with legal argument categories. This represents the current Supreme Court under Chief Justice John G. Roberts Jr.
Why Roberts Court?
The Roberts Court is particularly significant for… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-roberts-court-legal-arguments.danish-car-marketplace-dataset
Danish used car listings — raw dataset
~28,000 car listings scraped from the danish market as of june 2026. this is the raw, messy version — unit suffixes mixed into values, danish decimal separators, empty fields, the whole thing. the point is to have something real to practice data cleaning and analysis on, not a tidy kaggle dataset that does half the work for you.
if you want to jump straight into modeling with this data, check out danish-used-car-price-prediction — that repo… See the full description on the dataset page: https://huggingface.co/datasets/robertcaliforniadk/danish-car-marketplace-dataset.test_data_roberta_baseretrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
fast-gift-532182
fast-gift-532182
Synthetic sensors test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Sable-Robert/fast-gift-532182.augmented-roberts-jewelstwitter_ae_xlm_roberta_sentiment_stratifiedretrieval_verification_bm25_roberta
Dataset Card for "retrieval_verification_bm25_roberta"
More Information needed
umbra1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 5,
"total_frames": 5410,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra1.camera_in_box2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 9,
"total_frames": 18426,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box2.CSIC_RoBERTa_FT
Dataset Card for "CSIC_RoBERTa_FT"
More Information needed
Thunderbird_RoBERTa_FT
Dataset Card for "Thunderbird_RoBERTa_FT"
More Information needed
umbra03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 50,
"total_frames": 34838,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra03.PKDD_RoBERTa_FT
Dataset Card for "PKDD_RoBERTa_FT"
More Information needed
hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
bbq_roberta_large_race_custom_loss_lamda_07_predictionscamera_in_box_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 101,
"total_frames": 179197,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box_merged.camera_in_box_merged24This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 30,
"total_frames": 55523,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box_merged24.
