datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audit-reportsRepoBench-neo4jopen-github-major-repos🌐 AdhyanshVerma's Open GitHub Major Repos
An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel.
📖 Introduction
Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants.
It utilizes the Lazy Pointer Pattern: instead of bloating your storage… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.repo3multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.repo2rlenv-swe-flow
Repo2RLEnv SWE-Flow
Trace passing repository tests, select exercised source functions, remove their implementations and generate a reconstruction instruction. Restore the original functions as the reference solution and use the retained tests for deterministic reward.
Contains 100 Harbor tasks generated with the owned
swe_flow recipe in Repo2RLEnv.
Browse the complete task bundles in Harbor Visualiser or
open the task folders. Each folder is a runnable Harbor task:… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-swe-flow.prof_report__wavymulder-Analog-Diffusion__multi__24
Dataset Card for "prof_report__wavymulder-Analog-Diffusion__multi__24"
More Information needed
repo-imagescar-dataset-repo-v3financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.car-dataset-repoesg_reports_v2
Vidore Benchmark 2 - ESG Restaurant Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "french" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_v2.Stack-RepoThis is the Stack-Repo datasetesg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.economics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.blackbox-repository
Blackbox Repository
This dataset contains hyperparameter optimization (HPO) evaluations from several paper:
fcnet: Tabular benchmarks for joint architecture and hyperparameter optimization. Klein, A. and Hutter, F. 2019.
icml-deepar, icml-xgboost: A quantile-based approach for hyperparameter transfer learning. Salinas, D., Shen, H., and Perrone, V. 2021.
lcbench: Auto-PyTorch: Multi-Fidelity MetaLearning for Efficient and Robust AutoDL. Lucas Zimmer, Marius Lindauer, Frank… See the full description on the dataset page: https://huggingface.co/datasets/synetune/blackbox-repository.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
prof_report__22h-vintedois-diffusion-v0-1__multi__24
Dataset Card for "prof_report__22h-vintedois-diffusion-v0-1__multi__24"
More Information needed
G1-Pretrain-Data-Repohfcache_test_target_repo
Records
59783 records in total. Only 100 records shown.
id
width
height
rating
tags
file_size
file_url
created_at
75352
832
1216
e
1boy 1girl :o after_ejaculation after_footjob ahoge bangs bare_legs barefoot black_ribbon black_skirt blue_eyes blue_hair blush book boots boots_removed braid brown_cape cape commentary crossed_bangs cum cum_on_body cum_on_feet cum_string erection feet grey_shirt hair_between_eyes hair_flaps hair_intakes hair_ribbon hair_whiskers hay hetero… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/hfcache_test_target_repo.repo2rlenv-cli-gym
Repo2RLEnv CLI-Gym
Start with a healthy repository, synthesize a disruption and recovery, verify that damage breaks the original tests and recovery restores them, then export an environment-repair instruction and a deterministic verifier.
Contains 25 Harbor tasks generated with the owned
cli_gym recipe in Repo2RLEnv.
Browse the complete task bundles in Harbor Visualiser or
open the task folders. Each folder is a runnable Harbor task:
tasks/<task_id>/
├── task.toml… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-cli-gym.Dyna_Repo_Inria_MDPosit_Files
DynaRepo MDPosit
This repository mirrors Molecular Dynamics (MD) project metadata and files laid out per accession as exposed by Dynarepo. It includes global and per-accession manifests, project-level JSON, and the actual files (structures, trajectories, derived PCA binaries, and analysis screenshots). Field names and shapes follow Dynarepo’s documented model.
Overview
How to consume
Stream-parse: Read the file line-by-line and json.loads each… See the full description on the dataset page: https://huggingface.co/datasets/DataQuests/Dyna_Repo_Inria_MDPosit_Files.repojus
RepoJus: Jurisprudência Brasileira de Inteiro Teor
O RepoJus é um corpus consolidado de decisões judiciais e administrativas brasileiras obtidas de fontes públicas oficiais. O snapshot v2 reúne 4.106.391 registros canônicos, sem duplicatas de normalized_hash, em 4.666 shards Parquet determinísticos (39.290.819.138 bytes).
O corpus foi estruturado para pesquisa jurídica, recuperação de informação, RAG, classificação, sumarização, extração de entidades e treinamento ou avaliação… See the full description on the dataset page: https://huggingface.co/datasets/andrebadini/repojus.the_stack_v2_python_repos_pretraining_dataset_imported_context-datasetMixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.apt-cti-reports
APT CTI Reports Dataset
A collection of 2,710 PDF reports on Advanced Persistent Threats (APT) and Cyber Threat Intelligence (CTI).
Structure
├── apt_groups/ (937 files) - Reports attributed to specific APT groups
└── other/ (1773 files) - Multi-attribution, unattributed, and general CTI reports
Filename Format
<GROUP/TYPE>__<YEAR>__<TITLE>.pdf
Examples:
APT28__2019__Fancy_Bear_Campaign.pdf
MULTI__2023__Global_Threat_Report.pdf… See the full description on the dataset page: https://huggingface.co/datasets/hackerman700000/apt-cti-reports.audit_assistant_reportsRepoZero-Py2JS
RepoZero-Py2JS
RepoZero-Py2JS is a benchmark subset for evaluating zero-shot Python to JavaScript repository-level code translation.
Overview
Property
Value
Source Language
Python
Target Language
JavaScript (Node.js ESM)
Number of Libraries
24
Total Test Files
400
HuggingFace Dataset
jessezhaoxizhang/RepoZero-Py2JS
Task Description
Given a Python library implementation, the model must translate it… See the full description on the dataset page: https://huggingface.co/datasets/jessezhaoxizhang/RepoZero-Py2JS.Report-Generator-Data
