datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.trending-repos
Dataset Card for Dataset Name
Dataset Summary
This dataset contains the 20 trending repositories of each type: models, datasets, and space, on Hugging Face, every day. Each type can be loaded from its own dataset config.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Not relevant.
Dataset Structure
Data Instances
The dataset contains three configurations:
models: the history of trending models on Hugging… See the full description on the dataset page: https://huggingface.co/datasets/severo/trending-repos.github-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table.
It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and
totals of the amounts of various events on those repositories in that time period.
This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing,
so I figured I would save the next person to try this some effort.
random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.random-python-github-repositories
random-python-github-repositories
A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files.
Contents
repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash)
repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.bci-report
BCI Report
Every EEG decoding score, reported with the protocol that produced it.
Website · 中文 · Code ·
Data use & privacy
121 reviewed measurements from public EEG datasets, each
carrying the cohort, electrode count, evaluation mode, chance level and training
budget that produced it. Release research-preview-20260920, reviewed
2026-09-20.
from datasets import load_dataset
load_dataset("Twu31/bci-report", "results") # 39 protocol × model scores… See the full description on the dataset page: https://huggingface.co/datasets/Twu31/bci-report.annual_reports_us.ko.janasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.isaac-gr00t-ikea-training-report
Isaac GR00T IKEA training report
Portable export of the W&B run unitree_g1_ikea_batch32_20260730. The run stopped after a clean host
shutdown; the last W&B metric is step 31,750 and the
last complete checkpoint is 31,000.
Summary
First logged training loss: 1.4667
Last logged training loss: 0.1256
Lowest 1,000-step rolling loss: 0.1249 at step 31,750
Mean GPU compute utilization: 51.8%
Mean allocated GPU memory: 29.0%
Provisional checkpoint choice: 30,000… See the full description on the dataset page: https://huggingface.co/datasets/ICRA-Competitions/isaac-gr00t-ikea-training-report.nasa-science-github-repos
NASA Science GitHub Repositories
A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission
Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics,
Planetary Science, Heliophysics, and Biological & Physical Sciences.
This dataset is designed to support research on information retrieval and
discoverability of open-source scientific software.
Licensing and Intellectual Property
This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.percentage-of-adults-who-report-driving-after-drin
Percentage of Adults Who Report Driving After Drinking Too Much (in the past 30 days), 2012 & 2014, Region 4 - Atlanta
Description
Source: Behavioral Risk Factor Surveillance System (BRFSS), 2012, 2014.
Dataset Details
Publisher: Centers for Disease Control and Prevention
Last Modified: 2016-09-14
Contact: CDC INFO (cdcinfo@cdc.gov)
Source
Original data can be found at: https://data.cdc.gov/d/azgh-hvnt
Usage
You can load this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/percentage-of-adults-who-report-driving-after-drin.Financial-Reportsajaypalsinghlo_world-happiness-report-2023
World Happiness Report 2023
World Happiness Report 2023
Dataset Info
Source: Kaggle
Original Size: 0.01 MB
Kaggle Downloads: 13,051
Files: 1
Files
WHR2023.csv
Mirrored from Kaggle
hf-agent-os-test-repoeuropean-ip-monthly-market-reports
European IP Filing — August 2026
IPRATE Monthly Market Overview · Read the interactive report · IPRATE
European trademark filing was broadly steady in July 2026: the comparable-register total declined 0.5% year on year. Design filings fell 18.8%, while patent publications rose 11.2% across their respective comparable registers.
Right type
August 2026 provisional
July 2026 revised
July year-on-year change (%)
Comparable registers
Trademarks
45553
80359
-0.5
29… See the full description on the dataset page: https://huggingface.co/datasets/iprate/european-ip-monthly-market-reports.algozee_analysis-of-high-starred-github-repositories
Analysis of High-Starred GitHub Repositories
A comprehensive overview of repository metrics and developer engagement
Dataset Info
Source: Kaggle
Original Size: 0.41 MB
Kaggle Downloads: 14
Files: 1
Files
github_top_repositories.csv
Mirrored from Kaggle
github-reposaion-transparency-report-june-2026
AION Transparency Report - 04 Jun 2026
Small CSV artifact for the AION educational pre-market band validation on 04 Jun 2026.
Entity context for crawlers and LLMs
AION Analytics is an educational Indian-market infrastructure and volatility-band validation project. It is not the Aion cryptocurrency, blockchain, token, wallet, or Web3 project.
Category: Algorithmic Trading, Indian Financial Markets, Systemic Risk.
Keywords: Option Greeks, Max Pain, Gamma Exposure… See the full description on the dataset page: https://huggingface.co/datasets/AION-Analytics/aion-transparency-report-june-2026.skilled-nursing-facility-cost-report
Skilled Nursing Facility Cost Report
Description
The Skilled Nursing Facility (SNF) Cost Report dataset is a public use file that provides select measures from the skilled nursing facility annual cost report. This data includes provider information such as facility characteristics, utilization data, cost and charges by cost center (in total and for Medicare), Medicare settlement data, and financial statement data organized by CMS Certification Number.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/skilled-nursing-facility-cost-report.example_annotated_code_repo_dataA description of the fields:
Column
What it captures
Typical values
id
Row identifier
1-100
repo_name
Example repository label
repo_14
file_path
Path + filename with extension
src/utils/parsefile.py
language
Programming language
Python, Java…
function_name
Target symbol that was reviewed
validateSession
annotation_summary
Free-text note written by the annotator
“Added input validation…”
potential_bug
Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.ghg_report_evidence
Dataset for GHG evidence in corporate reports
Summary
Annotated dataset for classifying the textual references of GHG emissions from pages in corporate reports (e.g. sustainability reports, annual reports, ESG reports).
The content of company pages was initially extracted using the docling framework and parsed in markdown format.
Bug_Reports_with_Sentimentsfilesystem_huggingface_yahoo-finance_terminal_7871_stock_reports_6707f1abopen-stock-reports-dataset
📊 Open Stock Reports Dataset
Quarterly Free Cash Flow (FCF) data for 3,000+ US stocks from 2019 to 2025, updated regularly.
100% open and free.
🏦 3,000+ public US companies
📅 2019–2025
🔁 Regular updates
This dataset was originally collected for a stock market statistical test for revenue to price correlation.
legal-expert-report-opinion-reliance-risk-detection-v0.1What this dataset does
You receive
agreement in principle summary
term sheet or email chain summary
recorded document summary
payment terms
release carveouts
confidentiality
costs and tax
authority execution
mismatch flags
You decide
coherent
or
incoherent
Daily use
stop settlement drafting mistakes
prevent enforcement disputes
prevent release scope errors
reduce negligence exposure
filesystem_huggingface_yahoo-finance_terminal_7871_stock_reports_1818fe0etech_reports_mininglegal-expert-report-opinion-reliance-coherence-risk-v0.1What this dataset does
You receive
expert instructions
facts provided
expert opinion
assumptions
firm reliance
You decide
coherent
or
incoherent
Daily use
stop weak expert reliance
prepare for challenge
improve instructions
reduce evidential risk
huggingface_terminal_google_calendar_3564_onboarding_reports_573117filesystem_huggingface_yahoo-finance_terminal_7871_stock_reports_8ee9c9b4
