datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
browsecomp-plus-selected-tools-analysis-v1
BrowseComp-Plus: Selected Tools Analysis
Side-by-side view of selected tool calls from a reference trajectory alongside the new agent trajectory conditioned on those steps.
Retrieval model: Qwen3-Embedding-8BAgent model: gpt-oss-120bRun: traj_summary_ext_selected_tools_gpt-oss-120b_seed0
Columns
Column
Description
query_id
Query identifier
rationale
GPT rationale for why these k steps were selected from the reference trajectory
selected_indices
Step indices… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/browsecomp-plus-selected-tools-analysis-v1.source-analysis
NuBerea Source Analysis
Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.translation-analysis
NuBerea Translation Verse Texts
Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts.
Attribution
Upstream Data Sources
Source
License
Clementine Vulgate, NOCR
Public Domain
Luther Bible 1545, NOCR
Public Domain
Matthew's Bible 1537, Textus Receptus Bibles
Public Domain
NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.pseudepigrapha-analysis
NuBerea Pseudepigrapha Analysis
Derived linguistic datasets over pseudepigraphal literature, part of the NuBerea curated corpus estate. Covers the Greek and Latin witnesses of these texts along with a multilingual view across the available witness languages.
License
CC BY 4.0.
Attribution
Source
Link
License
NuBerea project
https://huggingface.co/NuBerea
CC BY 4.0
septuagint-analysis
NuBerea Septuagint Textual Analysis
Curated datasets for study of the Septuagint (the ancient Greek translation of the
Hebrew Bible), part of the NuBerea corpus estate of biblical and patristic texts.
It gathers Septuagint verse texts, apparatus notes, and edition-comparison material
into a set of ready-to-load configurations.
Attribution
This dataset derives from the following upstream sources, which require attribution:
Source
License
Rahlfs… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-analysis.lxx-analysis
NuBerea Research: LXX Translation-Technique Noise Model
Quantitative study of Septuagint translation technique: verse-by-verse measurements of
where the ancient Greek translation (LXX) diverges from the Hebrew Masoretic Text, with
book-level statistical summaries. The material lets researchers distinguish a translator's
habitual working style — free versus literal rendering — from genuine textual anomalies
worth close scholarly attention, putting on a measurable footing what LXX… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/lxx-analysis.ticker_analysis_articlesdetails_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.ticker_analysis_pricesgretel-financial-risk-analysis-v1
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated by fine-tuning Phi-3-mini-128k-instruct on 14,306 SEC filings (10-K, 10-Q, and 8-K) from 2023-2024, utilizing differential privacy. It is designed for training models to extract key risk factors and generate structured summaries from financial documents while demonstrating the application of differential privacy to safeguard sensitive information.
This dataset showcases… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-financial-risk-analysis-v1.QuRatedPajama-1B_tokens_for_analysis
QuRatedPajama
Paper: QuRating: Selecting High-Quality Data for Training Language Models
This dataset is a 1B token subset derived from princeton-nlp/QuRatedPajama-260B, which is a subset of cerebras/SlimPajama-627B annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria:
Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers
Facts & Trivia - how much factual and trivia knowledge the text… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-1B_tokens_for_analysis.sentiment-analysis-in-commodity-market-gold
Dataset Card for Sentiment Analysis of Commodity News (Gold)
This is a news dataset for the commodity market which has been manually annotated for 10,000+ news headlines across multiple dimensions into various classes. The dataset has been sampled from a period of 20+ years (2000-2021).
The dataset was curated by Ankur Sinha and Tanmay Khandait and is detailed in their paper "Impact of News on the Commodity Market: Dataset and Results." It is currently published by the authors on… See the full description on the dataset page: https://huggingface.co/datasets/SaguaroCapital/sentiment-analysis-in-commodity-market-gold.isom5240-td-traffic-analysisharvey-labs-llm-artifact-analysis
Harvey Labs LLM artifact analysis
This dataset contains artifacts from a non-LLM analysis of the Harvey Labs DOCX corpus.
The analysis used filename similarity, document extraction heuristics, a small manually
labeled seed set, CatBoost native text features, and a native CatBoost embedding feature
built from a mean Word2Vec representation. It was designed to find documents where an
LLM refused the requested task and returned a safe alternative instead.
Source… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/harvey-labs-llm-artifact-analysis.details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1isom5240-td-application-traffic-analysis
Split:
application: 38 samples
Class Distribution:
car (ID: 3): 67 (46.2%)
motorcycle (ID: 4): 14 (9.7%)
airplane (ID: 5): 62 (42.8%)
truck (ID: 8): 2 (1.4%)
Annotation Files:
Latest: application/application_labels.json
Timestamped: application/application_labels_20250309_212205.json
Split:
application: 49 samples
Class Distribution:
car (ID: 3): 120 (60.0%)
motorcycle (ID: 4): 14 (7.0%)
airplane (ID: 5): 62 (31.0%)
truck (ID: 8):… See the full description on the dataset page: https://huggingface.co/datasets/slliac/isom5240-td-application-traffic-analysis.details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1_v2.open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.gene-expression-tcga-cptac-survival-analysisomnimind-colab-baseline-analysis
OmniMind — Colab baseline analysis
Unified analysis outputs from the Colab run on 2026-09-02.
Datasets analyzed
brain_organoids_four_protocols — 69,794 cells, 4 protocols.
10x_visium_dlpfc — 47,681 Visium spots from 12 DLPFC sections.
linnarsson_astrocyte — 155,025 human astrocytes.
linnarsson_upper_layer_intratelencephalic — 455,006 human upper-layer cortical neurons.
Structure
brain_organoids/ — protocol × annotation crosstabs, QC stats, NEST… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-colab-baseline-analysis.ami-speaker-analysis_full_run_silero_Final_atlasttweets_pt_sentiment_analysis
Dataset Card for "tweets_pt_sentiment_analysis"
More Information needed
text-analysis-context-cased-case-datahoike_normal_expression_GTEx_Analysis_v10_log2tpmplus1
Hōʻike - Normal GTEx Gene Expression Data in Log2(TPM+1) Format
These data are for use in the Hoike gene expression data generation models as the normal_profile input.
These were obtained from the GTEx_Analysis_v10_RNASeQCv2.4.2_gene_tpm.gct.gz file from the GTEx Portal on 6/10/2026.
review-analysistwitter-sentiment-analysis
Twitter Sentiment Analysis: Prabowo's First 100 Days
Dataset Overview
This dataset contains tweets related to President Prabowo Subianto's first 100 days in office in Indonesia (2024-2029). The tweets have been preprocessed and classified into three sentiment categories using a fine-tuned BERT model for Indonesian language (IndoBERT).
Dataset Details
Language: Indonesian
Source: Twitter/X
Time period: First 100 days of President Prabowo's… See the full description on the dataset page: https://huggingface.co/datasets/KidzRizal/twitter-sentiment-analysis.clone-of-gretel-financial-risk-analysis-v1
⚠️🔴 IMPORTANT NOTICE 🔴⚠️
This dataset is directly cloned from gretelai/gretel-financial-risk-analysis-v1 on Hugging Face. No modifications have been made to the original dataset, it is only for archival.
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated using differential privacy guarantees, trained on 14,306 SEC (10-K, 10-Q, and 8-k) filings from 2023-2024. The dataset is designed for training models to extract… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/clone-of-gretel-financial-risk-analysis-v1.log-analysis-hdfs-preprocessedmyanmar-social-media-sentiment-analysis-dataset
Myanmar Social Media Sentiment Analysis Dataset
A Myanmar language dataset for sentiment analysis of social media content, translated from an English source dataset.
Dataset Description
This dataset contains social media text with sentiment annotations translated into Myanmar language. It is derived from the original Social Media Sentiments Analysis Dataset on Kaggle, with texts professionally translated to Myanmar language while preserving the sentiment labels.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-social-media-sentiment-analysis-dataset.Trial_Meta-Analysis
Automatically Extracting Numerical Results from Randomized Controlled Trials with LLMs
Dataset Description
Links
Homepage:
Github
Repository:
Github
Paper:
arXiv
Leaderboard:
PapersWithCode
Contact (Original Authors):
Hye Sun Yun (yun.hy@northeastern.edu)
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
The human-annotated data is available in the data folder as both csv and json formats. The dev set has 10… See the full description on the dataset page: https://huggingface.co/datasets/araag2/Trial_Meta-Analysis.
