datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.religious-artwork-analysis-data
Data
Download from Kaggle (needs an API token from https://www.kaggle.com/settings):
pip install kaggle
python data/download.py
Expected layout after download:
data/artwork_metadata.csv 3,997 rows — filename, religion (1,000 each of
buddhism / christianity / hinduism; 997 islam),
sub_religion, artist, title, year, place,
source, source_id, source_url, image_url
data/images/ the… See the full description on the dataset page: https://huggingface.co/datasets/cvikl/religious-artwork-analysis-data.routing_analysis-code-data
routing_analysis — code and data
This dataset repository stores the non-checkpoint portion of the
routing_analysis filesystem snapshot as individual files under their
original relative paths. Files are uploaded directly; they are not packed into
split tar archives.
The selection is enumerated from the filesystem and does not consult
.gitignore. Ignored files and .gitignore files themselves are therefore
included whenever they belong to the code/data selection.
Training… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-code-data.DataMind-Analysis-SFT-DataThis repository contains the data presented in Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
Code: https://github.com/zjunlp/DataMind
sentiment_analysis_data
Dataset Card for "sentiment_analysis_data"
More Information needed
open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.math_benbench_data_leak_analysis
Dataset description
This is a math dataset mixed from four open-source data. It was used to analyze the contamination test on the MATH and contains 1M samples.
Dataset fields
question
question from open-source data
solution
the answer corresponding to question
5grams
5-gram list of f"{question} {answer}"
test_question
the most relevant question from MATH
test_solution
the answer corresponding to test_question
test_5grams
5-gram list of… See the full description on the dataset page: https://huggingface.co/datasets/newsbang/math_benbench_data_leak_analysis.global-commodity-shocks-analysis-dataautotrain-data-ukrainian-telegram-sentiment-analysis
AutoTrain Dataset for project: ukrainian-telegram-sentiment-analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project ukrainian-telegram-sentiment-analysis.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "\u0421\u043e\u0432\u043e\u043a",
"target": 1
},
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/dmytrobaida/autotrain-data-ukrainian-telegram-sentiment-analysis.text-analysis-context-cased-case-data1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.data-analysis-datasets
CANNS Analysis Datasets
This repository contains example datasets for the CANNS (Continuous Attractor Neural Networks) data analysis package.
Datasets
ROI_data.txt (703 KB)
Description: 1D CANN ROI data for bump analysis
Format: Text file with neural activity measurements
Usage: 1D CANN analysis, MCMC bump fitting
Example: Used in 1D CANN analysis tutorials
grid_1.npz (8.7 MB)
Description: Grid cell spike data with position information
Format:… See the full description on the dataset page: https://huggingface.co/datasets/canns-team/data-analysis-datasets.Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data
Dataset Summary
The dataset is a collection of Youtube Comments and it was captured using the YouTube Data API.
The data set consists of 1500 nostalgic and non-nostalgic comments in English.
Languages
The language of the data is English.
Citation
If you find this dataset usefull for your study, please cite the paper as followed:
@article{postalcioglu2020comparison,
title={Comparison of Neural Network Models for Nostalgic Sentiment Analysis of YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Senem/Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data.livebench-data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/lthn/livebench-data_analysis.tcc-sentiment-analysis-dataautotrain-data-books-rating-analysis
AutoTrain Dataset for project: books-rating-analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project books-rating-analysis.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"feat_Unnamed: 0": 1976,
"feat_user_id": "792500e85277fa7ada535de23e7eb4c3",
"feat_book_id": 18243288… See the full description on the dataset page: https://huggingface.co/datasets/LewisShanghai/autotrain-data-books-rating-analysis.1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.HS_data_analysis_teaching_0829-freedrive-cam4instruct_code_for_data_analysistennis-momentum-analysis-data
Dataset Card for Tennis Momentum Analysis Dataset
This dataset contains detailed point-by-point data from a 2023 Wimbledon Championship tennis match between Carlos Alcaraz and Nicolas Jarry, designed to support momentum analysis and sports analytics research. It captures 100 data points with 47 features covering match progression, player performance metrics, and technical details of each point.
Dataset Details
Dataset Description
The Tennis Momentum… See the full description on the dataset page: https://huggingface.co/datasets/sarahzhoo620/tennis-momentum-analysis-data.analysis-data
EvoLen — token analysis input data
Derived interval files needed to reproduce the token analyses in Section 4 of
EvoLen: Evolution-Guided Tokenization for DNA Language Model
(arXiv:2604.08698).
Analysis code lives in the evolen repository
under analysis/.
Contents
region_beds/
source/ INPUT to the P4 enrichment analysis -- the four genomic
regions, merged and cleaned:… See the full description on the dataset page: https://huggingface.co/datasets/EvoLenTokenizer/analysis-data.sentiment_analysis_financial_news_dataTwitter-mulitlingual-synthetic-data-sentiment-analysis
Dataset Card for Twitter-mulitlingual-synthetic-data-sentiment-analysis
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Paul-HF/Twitter-mulitlingual-synthetic-data-sentiment-analysis/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/Paul-HF/Twitter-mulitlingual-synthetic-data-sentiment-analysis.autotrain-data-sentiment_analysis
AutoTrain Dataset for project: sentiment_analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project sentiment_analysis.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "I grew up (b. 1965) watching and loving the Thunderbirds. All my mates at school watched. We played \"Thunderbirds\" before… See the full description on the dataset page: https://huggingface.co/datasets/raghuram13/autotrain-data-sentiment_analysis.Data_Analysis_Workflow_20240509_134737autotrain-data-mnist-analysis
AutoTrain Dataset for project: mnist-analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project mnist-analysis.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<28x28 L PIL image>",
"target": 1
},
{
"image": "<28x28 L PIL image>",
"target": 0
}
]
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/realzdlegend/autotrain-data-mnist-analysis.code_completion_for_data_analysisData_Analysis_Workflow_20240509_121811claw-analysis-dataautotrain-data-imdb-sentiment-analysis
AutoTrain Dataset for project: imdb-sentiment-analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project imdb-sentiment-analysis.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "Me neither, but this flick is unfortunately one of those movies that are too bad to be good and too good to be… See the full description on the dataset page: https://huggingface.co/datasets/linktimecloud/autotrain-data-imdb-sentiment-analysis.
