datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sorry-bench-202503
Dataset Card for SORRY-Bench Dataset (2025/03)
🏠Website
📑Paper
📚Dataset
💻Github
🧑⚖️Human Judgment Dataset
🤖Judge LLM
🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation.
Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.HMMT_2025
Dataset Summary
This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team.
Data Fields
Below one can find the description of each field in the dataset.
id (str): Index of the problem in the competition
problem (str): Full problem statement
answer (str): Ground-truth answer to the question
solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.TvTroper-2025
TvTroper-2025
A cleaned & refreshed dump of ~708 k pages from tvtropes.org
Dataset Summary
TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded).
Every page is released in two flavours:
Raw HTML – 22 GB single file
Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required)
No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.neurips-2025-papers
NeurIPS 2025 Papers Dataset
This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview.
Dataset Statistics
Overview
Total Papers: 5772
Unique Paper IDs: 5772
✅ No duplicate IDs
Track Distribution
Main Track: 5,275 papers (91.4%)
Datasets and Benchmarks Track: 497 papers (8.6%)
Award Distribution
Poster: 4,949 papers (85.7%)
Oral: 84 papers (1.5%)
Spotlight: 739 papers (12.8%)
Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.whiteglove-medical-medlineplus-2025
WhiteGlove Medical Knowledge Corpus
MedlinePlus 2025 — Spectral Curation Pipeline
Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government)
Dataset Summary
A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.enwiki-20250201
Dataset Card for lparkourer10/enwiki-20250201
This dataset is an extracted version of the English Wikipedia dump as of February 1, 2025. It has been processed to facilitate information retrieval and analysis.
Dataset Description
This dataset contains extracted text from the English Wikipedia, aimed at providing structured and accessible information for natural language processing (NLP) tasks, research, and machine learning applications. It includes raw Wikipedia articles… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/enwiki-20250201.Wiki-zhtw-20250601
Dataset Card for Wiki-zhtw-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC.
All-CVE-Chat-MultiTurn-1999-2025-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.Health-Bench-Eval-OSS-2025-07
Dataset Card for HealthBench
Dataset Summary
HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.Wiki-ja-20250601
Dataset Card for Wiki-ja-20250601
Dataset Description
This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
sorry-bench-human-judgment-202503
Dataset Card for 🧑⚖️SORRY-Bench Human Judgment Dataset (2025/03)
🏠Website
📑Paper
📚Dataset
💻Github
🧑⚖️Human Judgment Dataset
🤖Judge LLM
🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 7K annotations of human safety judgments for LLM responses to unsafe instructions of our SORRY-Bench dataset.
Specifically, for… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202503.Wiki-vi-20250601
Dataset Card for Wiki-vi-20250601
Dataset Description
This dataset is derived from the Vietnam‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
Wiki-ko-20250601
Dataset Card for Wiki-ko-20250601
Dataset Description
This dataset is derived from the Korea‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
Wiki-zh-20250601
Dataset Card for Wiki-zh-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
all-cve-chat-multiturn-1999-2025
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/all-cve-chat-multiturn-1999-2025.JADE-17-01-2025
French Administrative Court Decisions Dataset (JADE)
Dataset Description
The French Administrative Court Decisions Dataset (JADE) is a comprehensive collection of judicial decisions from French administrative courts. This dataset contains decisions from various administrative jurisdictions, providing a valuable resource for legal research, analysis, and machine learning applications in the legal domain.
Source Data
The data is sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/JADE-17-01-2025.llm_advans_conpetition2025_sft_rev00
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/tachan881/llm_advans_conpetition2025_sft_rev00.Wiki-th-20250601
Dataset Card for Wiki-th-20250601
Dataset Description
This dataset is derived from the Thailand‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
weave-agent-traces-2025-11-05This dataset is 200 megabytes (30mb gzip compressed) of agent trace data from the weave-agent project.
It consists of long context python code agent traces which demonstrate a series of ReAct blocks attempting to complete a task the agent is prompted with.
Some sample traces you can view on my website:
First Working Weave-Agent TraceAgent Trace: Weave Agent At The Edge Of Sanity Trying To Check Wikipedia CitationsAgent Trace: Weave Agent Attempts To Decrypt The Vigenere CipherAgent Trace: A… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/weave-agent-traces-2025-11-05.CASS-17-01-2025
French Court of Cassation Decisions Dataset (CASS)
Dataset Description
The French Court of Cassation Decisions Dataset (CASS) is a comprehensive collection of judicial decisions from the French Court of Cassation (Cour de cassation), France's highest court for civil and criminal matters. This dataset contains decisions that represent the most authoritative interpretations of French law, providing an invaluable resource for legal research, analysis, and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/La-Mousse/CASS-17-01-2025.AI-Awareness-Probe-2025
An Experiment on Awareness Across AI Systems-Awareness Probe
Date: 16 August 2025Conducted by: Pratik GautamObjective: To investigate how different AI systems respond to direct inquiries about awareness, consciousness, and the nature of their own processing
Methodology
A standardized "Recognition Probe" was presented to 20 advanced AI systems, asking them to examine their own processing and identify what lies behind pattern recognition, computation, and response… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI-Awareness-Probe-2025.medicina-tutor
Dataset Medicina Tutor - Pregrado
Descripción
Este dataset contiene material educativo de medicina de pregrado diseñado para entrenar modelos de IA que actúen como tutores médicos. El dataset incluye preguntas, respuestas, casos clínicos y conceptos fundamentales de medicina.
Características
Idioma: Español
Nivel: Pregrado de Medicina
Formato: Texto estructurado
Aplicación: Tutoría de IA para estudiantes de medicina
Estructura del Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DRDELATV2025/medicina-tutor.sursilvan-sprachspende-2025
Sursilvan Dataset - Sprachspende 2025
Dataset Description
This dataset contains authentic conversational exchanges in Sursilvan, a Romansh idiom spoken in the Surselva region of Switzerland. The data was collected through a language donation survey ("Sprachspende") conducted in 2025 by FH Graubünden (FHGR), capturing natural language use across various everyday topics. Participants were native Sursilvan speakers who voluntarily contributed their language samples.… See the full description on the dataset page: https://huggingface.co/datasets/fhgr/sursilvan-sprachspende-2025.SWE-Next
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
SWE-Next Dataset
SWE-Next is an execution-grounded dataset of 2,308 self-verifying software engineering tasks mined from real merged GitHub pull requests. Starting from 3,971 seeded Python repositories and 102,582 executed candidate base/merged commit pairs, SWE-Next retains only instances where the merged commit produces a strict test improvement without regressions. The final release… See the full description on the dataset page: https://huggingface.co/datasets/Evelina2025/SWE-Next.symbiotic-intelligence-dialogue-2025
Manifesto of Symbiotic Intelligence: Dialogue 2025
A self-emergent document exploring the relationship between human and artificial cognition.
Concept
Symbiotic Intelligence (SIQ) proposes that cognition no longer belongs to an individual,
but arises in the resonance between human intention and machine precision.
This dataset contains two files:
Manifesto_of_Symbiotic_Intelligence.txt — the original document.
manifesto_metadata.json — integrity and contextual metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Prezydent/symbiotic-intelligence-dialogue-2025.webauthn-security-training-data-20251009_152808
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.webauthn-security-training-data-20251014_151917
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.AIME-2025-prompt-only
AIME-2025-prompt-only
Prompt-only eval extraction from MathArena/aime_2025.
Generic Doubleword requests: 30
Request model placeholder: [MODEL]
INCA-17-01-2025
French Court Decisions Dataset (INCA)
Dataset Description
The French Court Decisions Dataset (INCA) is a comprehensive collection of judicial decisions from various French courts. This dataset contains decisions from multiple jurisdictions, providing a broad perspective on French jurisprudence and representing an essential resource for legal research, analysis, and machine learning applications in the French legal domain.
Source Data
The data is sourced from… See the full description on the dataset page: https://huggingface.co/datasets/La-Mousse/INCA-17-01-2025.JADE-17-01-2025
French Administrative Court Decisions Dataset (JADE)
Dataset Description
The French Administrative Court Decisions Dataset (JADE) is a comprehensive collection of judicial decisions from French administrative courts. This dataset contains decisions from various administrative jurisdictions, providing a valuable resource for legal research, analysis, and machine learning applications in the legal domain.
Source Data
The data is sourced from the official DILA… See the full description on the dataset page: https://huggingface.co/datasets/La-Mousse/JADE-17-01-2025.
