datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-copyacrobat-patches-v3
ACROBAT Registered H&E-IHC Patches v3
H&E-IHC patch pairs extracted from registered ACROBAT whole-slide images. All IHC slides were warped into H&E space by VALIS — patches at the same (x, y) coordinates in HE and IHC are perfectly aligned.
Key Features
205 patients — full ACROBAT breast cancer dataset
9,658 HE patches at 1024×1024 px, 0.92 µm/px (10X)
31K IHC pairs: ER (8,385), PGR (8,453), HER2 (5,543), KI67 (8,439)
WSI thumbnails — 512×512 low-res H&E per patient for… See the full description on the dataset page: https://huggingface.co/datasets/ahmedayman4a/acrobat-patches-v3.acrft-annot-noprop
acrft-annot-noprop
RLT annotation for AC-RFT critic training: raw memmaps (.dat) + meta.json, read directly by
scripts/train_rlt_critic.py and scripts/eval_rlt_critic.py in the openpi fork.
Shapes are in meta.json: rl_token [T, D], base_action [T, N, H, A], action_chunk [T, H, A],
reward/mc_return/done/episode_index/frame_index [T], base_action_heldout [T, num_heldout, H, A].
dtype and reward_scheme are in meta.json. Load with numpy.memmap.
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/acrft-annot-noprop.ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/acrever/ACSE-Eval.AAD-Across-All-Domains-TestAcrobot-v1
Acrobot-v1 - Imitation Learning Datasets
This is a dataset created by Imitation Learning Datasets project.
It was created by using Stable Baselines weights from a DQN policy from HuggingFace.
Description
The dataset consists of 1,000 episodes with an average episodic reward of -69.852.
Each entry consists of:
obs (list): observation with length 6.
action (int): action (0, 1 or 2).
reward (float): reward point for that timestep.
episode_returns (bool): if that state was… See the full description on the dataset page: https://huggingface.co/datasets/NathanGavenski/Acrobot-v1.repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
acro-yamalex-llmjp-4-math-cot
acro-yamalex-llmjp-4-math-cot(データセット)
日本語数学推論のためのChain-of-Thought (CoT) データセットです。
StackMathQAの問題に対して、DeepSeek V3を用いて日本語CoT形式の解法を生成しました。
本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。
データセット概要
項目
値
データ件数
306,366件
言語
日本語
ソース
StackMathQA
生成モデル
DeepSeek V3 (deepseek-chat)
フォーマット
JSONL(マルチターン対話形式)
データ作成手法
OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。
Step 1: 日本語CoT解法の生成
StackMathQAの問題に対してDeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-cot.acro-yamalex-llmjp-4-math-tir
acro-yamalex-llmjp-4-math-tir(データセット)
日本語数学推論のためのTool-Integrated Reasoning (TIR) データセットです。
自然言語による推論とPythonコード実行を組み合わせたマルチターン形式のデータセットで、OpenWebMathから抽出・生成した問題に対してDeepSeek V3を用いてTIR形式の解法を生成しました。
本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。
データセット概要
項目
値
データ件数
134,834件
言語
日本語
ソース
OpenWebMath
生成モデル
DeepSeek V3 (deepseek-chat)
フォーマット
JSONL(マルチターン対話形式)
データ作成手法
OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-tir.EverythingLM-V3-ShareGPT
EverythingLM V3 Data converted to ShareGPT format.
acrac
Dataset Card for the ACR Appropriateness Criteria Corpus
This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help clinicians order appropriate diagnostic imaging studies for patients. The corpus is formatted similarly to the corpuses introduced in MedRAG by Xiong et al. (2024), and can therefore be similarly used for medical Retrieval-Augmented Generation (RAG).… See the full description on the dataset page: https://huggingface.co/datasets/michaelsyao/acrac.arxiv-acronym-gen
AcronymGen-Titles 🧠
AcronymGen-Titles is a dataset of over 8,000 high-quality acronym–title pairs extracted from scientific papers on arXiv. Each entry contains a scientific acronym and the corresponding paper title in which it appears. This variant contains title-only pairs — no abstracts — offering a compact and focused dataset for acronym generation, matching, or compression tasks.
📚 Dataset Overview
📦 Size: 8,070 examples
🧠 Source: Extracted from arXiv.org… See the full description on the dataset page: https://huggingface.co/datasets/sjmoran/arxiv-acronym-gen.pii_ner_instructions_alpaca_styleritos_igrejaacrmeu_casamento
