datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
analytics
NuBerea Analytics
Derived analytical tables over a curated biblical corpus estate, part of the
NuBerea corpus estate of biblical and patristic texts. It gathers computed
research outputs — comparative, structural, and simulation-style summaries
spanning the corpora — into a set of ready-to-load configurations.
Attribution
Analytical outputs are original to the NuBerea project. Underlying data derive
from NuBerea datasets licensed CC BY 4.0 or in the public domain.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/analytics.analytics-pauline
NuBerea Research: Pauline Authorship
Computational stylometric analysis of the Pauline epistles, addressing two
classic questions of New Testament scholarship: the authorship of the disputed
letters (Ephesians, Colossians, 2 Thessalonians, and the Pastorals) and the
composite-letter hypothesis for 2 Corinthians (the Semler and Bornkamm
partition theories). The dataset gathers derived analytical summary tables —
register profiles, calibrated leave-one-out comparisons, and… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/analytics-pauline.analytics-synoptic
NuBerea Research: Synoptic Problem
Computational analysis of the synoptic problem — the question of the literary
relationships among the Gospels of Matthew, Mark, and Luke. The dataset gathers
derived analytical tables on pericope order agreement (the Markan-priority /
Lachmann argument), on the minor agreements (readings shared by Matthew and
Luke against Mark), and on register profiles relevant to the Q and Farrer
hypotheses. Part of the NuBerea curated corpus estate of… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/analytics-synoptic.analytical_reasoningringside-analytics
Ringside Analytics — Pro Wrestling Match Archive
A relational snapshot of professional wrestling history from 1980 to the present:
292K matches, 611K wrestler-match participations, 35K events, and 12.8K
wrestlers across WWE, AEW, WCW, ECW, NXT, TNA, and others. Sourced from
public Cagematch.net scrapes and the alexdiresta profightdb dump, normalized
into a Postgres schema, and exported as parquet files that preserve the
relational structure (one file per table, joinable by id).
This… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/ringside-analytics.OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains.
football_analytics
⚽ Football Analytics Database
Overview
This dataset contains a processed SQLite database generated from StatsBomb Open Data.
The objective of this dataset is to provide a ready-to-use relational database for football analytics, eliminating the need to parse and transform thousands of raw JSON files.
The database was created as part of the Football Analytics Dashboard project and is intended for:
Football Analytics
Sports Data Science
SQL Practice
Data Engineering… See the full description on the dataset page: https://huggingface.co/datasets/Maiyarasu/football_analytics.cloudflare-domain-traffic-analytics
Cloudflare Domain Traffic Analytics Snapshot
This repository is a structured snapshot of Cloudflare data across 200 zones. It combines domain inventory, redirect configuration, DNS metadata, Web Analytics setup, HTTP time series, request dimension groups, real-user page-load and Web Vitals metrics, DNS analytics, Speed API availability, and firewall aggregates where Cloudflare exposes them into analysis-ready Parquet tables.
The collection window for traffic tables is… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/cloudflare-domain-traffic-analytics.SweSetupBench-liteThis repository contains the data presented in SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks.
analytical_reasoning_1analytics-parquetsOrcaAgentInstruct-analyticalreasoningcricket_analytics_datasetf1-visual-analyticsdev-sent-llm-use-analyticsamazon-beauty-analytics-ready
Amazon Reviews 2023 — Beauty & Personal Care (Analytics / ML Ready)
Stage 2 of 2. ต่อยอดจาก star schema ใน
Madnesss/amazon-beauty-star-schema
โดยทำความสะอาดข้อความ สร้าง feature และแบ่ง train/val/test ตามเวลาให้เรียบร้อย
โหลดแล้วเทรนโมเดลได้เลย
ไฟล์ในชุดนี้
ไฟล์
จำนวนแถว
ใช้ทำอะไร
reviews/split=train/
2,581,670
เทรนโมเดล
reviews/split=val/
467,957
จูน hyperparameter
reviews/split=test/
320,143
วัดผลครั้งสุดท้าย
product_features.parquet
998,708… See the full description on the dataset page: https://huggingface.co/datasets/Madnesss/amazon-beauty-analytics-ready.customer-transcript-analytics
Customer Transcript Analytics
Curated customer-support and meeting transcripts mapped to a single fixed
"analyze this transcript → compact JSON" prompt, for benchmarking batched
offline LLM inference on realistic workloads.
Motivation and intended use
This dataset provides a realistic transcript-analytics workload for batched
offline-inference experiments: throughput benchmarking and
predicted-vs-observed throughput validation. Rows range from short support
chats… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-analytics.ecommerce-funnel-analyticsanalytic-combinatorics-llm-samples-v1
analytic-combinatorics-llm-samples-v1 (PARTIAL)
Status: In progress — Llama-3.1-8B Phase 1 sampling on DGX Spark.
Current state
Rows: 2000 (of 15,000 target for Llama-3.1-8B)
Cells complete: open_qa T=0.1 (1K), open_qa T=0.7 (1K)
Truncation: 0%
Stochasticity: Verified — 995/1000 unique full outputs per cell
History
v1 (previous) was deleted due to a determinism bug in sample_vllm.py:
passing fixed seed=42 to vLLM produced 50 identical outputs per prompt.… See the full description on the dataset page: https://huggingface.co/datasets/latkes/analytic-combinatorics-llm-samples-v1.user-analytics-privatemultilingual-hate-detection-dataset_v38bmultilingual-hate-detection-dataset_v38agretel-pii-masking-en-v1-ner-coarseclinical-trials-sponsor-analytics
Clinical Trial Sponsor Analytics
⬇️ This is a 10% sample. Get the full dataset (10,200+ sponsors) on Gumroad →
10,200+ clinical trial sponsors profiled with trial volume, completion rates, phase distribution, therapeutic focus, and pipeline activity. Built for competitive intelligence, due diligence, and portfolio analysis.
Why This Dataset Exists
Answering questions like "What's Pfizer's completion rate in oncology?" or "Which biotech sponsors have the most active… See the full description on the dataset page: https://huggingface.co/datasets/wapplewhite4/clinical-trials-sponsor-analytics.multilingual-hate-detection-dataset_v38fmultilingual-hate-detection-dataset_v38gai4privacy-pii-masking-en-v1-nermultilingual-hate-detection-dataset_v39gwikipedia_crystallography_analyticalgretel-pii-masking-en-v1
