datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
acme-home-inbox
ACME Home Inbox
Public dataset: https://huggingface.co/datasets/Mitchins/acme-home-inbox
The v0.2.0 checkpoint contains the complete synthetic dataset plus the first
canonical OCR/vision deployment bake-off. Benchmark results are an auditable
research checkpoint, not a claim that any tested routing policy is ready for
unattended household use.
ACME Home Inbox is a fully synthetic, reproducible stress test for a practical
systems question:
Does graphical evidence change the… See the full description on the dataset page: https://huggingface.co/datasets/Mitchins/acme-home-inbox.AcmeTrace
Acme Trace
This repository hosts the public releases of Acme traces from the Shanghai AI Lab, encompassing workloads spanning from March 2023 to August 2023. We encourage anyone to use the traces for academic purposes, and if you had any questions, feel free to send an email to us, or file an issue on Github.
Furthermore, we have conducted a thorough analysis of the Acme workloads, detailed in our NSDI '24 paper titled Characterization of Large Language Model Development in the… See the full description on the dataset page: https://huggingface.co/datasets/Qinghao/AcmeTrace.Anonymous_ACMMM_2025_Submission
🗂️ Anonymous_ACMMM_2025_Submission Dataset
This dataset is prepared for the Anonymous ACMMM 2025 submission, containing multi-view event-based data designed for dynamic 3D scene reconstruction tasks.
📁 Dataset Structure
Each subfolder corresponds to a distinct synthetic or real-world scene, such as:
lego_6_views/
capsule_6_views/
garage_6_views/
Restroom_6_views/
Cubes_6_views/
Hinge_6_views/
MC-Toy_6_views/
Rubik’s-Cube_6_views/
Each scene folder contains 6 views… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-ACMMM-2025-Submission/Anonymous_ACMMM_2025_Submission.acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-3
Annotates the rollouts of: iter-2 rollouts (…-train-rollouts-…-epoch2)
One .npz per (question, rep) trajectory · 736 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch3.acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-2
Annotates the rollouts of: iter-1 rollouts (…-train-rollouts-…-epoch1)
One .npz per (question, rep) trajectory · 849 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch2.acme-sentiment-511pin4k
Customer Sentiment Corpus
Samples: 25000
License: mit
Language: en
Description
Sentiment-labeled customer feedback corpus.
Usage
Intended for sentiment classification of customer feedback.
acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1
Teacher (Qwen3.5-397B-A17B) top-20 forward-KL log-prob annotations for offline on-policy
distillation (OPD) of Qwen3.5-9B on BrowseComp-Plus train680 (MemTool regime).
Trains: OPD iter-1
Annotates the rollouts of: base model rollouts (react+memtool ×4)
One .npz per (question, rep) trajectory · 1145 files.
Schema (per file, numpy.load)
key
shape
dtype
meaning
input_ids
(L,)
int32… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-teacher-logprobs-qwen3.5-9b-epoch1.Acmmm2025_video_benchmarkacm-browsecompplus-train-rollouts-qwen3.5-9b
lixiaochuan2020/acm-browsecompplus-train-rollouts-qwen3.5-9b
Base Qwen3.5-9B student rollouts on BrowseComp-Plus bcp_train_680 (680 questions),
pass@4 (4 runs), two agent regimes:
memtool/ — context-managed (MemTool: manage_context + query_memory), 131K / 100 turns
react/ — ReAct baseline
Each run{1..4}/ holds full per-question trajectories (run_*.json) and the GPT-5 grade
file (gpt5_eval.json). These are the example Stage-1 rollouts for the BrowseComp-Plus OPD… See the full description on the dataset page: https://huggingface.co/datasets/lixiaochuan2020/acm-browsecompplus-train-rollouts-qwen3.5-9b.acm-browsecompplus-train-rollouts-qwen3.5-9b-epoch1
lixiaochuan2020/acm-browsecompplus-train-rollouts-qwen3.5-9b-epoch1
BrowseComp-Plus train680 pass@4 rollouts (MemTool regime) from qwen3.5-9b-opd_iter1 — one row per (question, rep) in rollouts.jsonl.
Fields: question, gold, final_answer, correct_gpt5 (GPT-5 judge), num_turns, tool_call_counts, and full trajectories (raw_history + folded history + mem_operations + token_trajectory).
Stats: 2720 rollouts / 680 questions / reps [1, 2, 3, 4] · pass@1 70.0% · pass@4 83.2% (GPT-5).
multi_domain_ai_human_text
multi_domain_ai_human_text — Datasheet
Balanced, multi-domain AI-vs-human text detection benchmark with dedicated
out-of-distribution and adversarial evaluation panels. Built by
scripts/build_paper_dataset.py from an 11-corpus unified aggregation.
Splits
Split
AI
Human
Total
Purpose
train
300,000
300,000
600,000
training (balanced, English, clean)
validation
2,996
2,999
5,995
model selection
test
4,991
4,999
9,990
in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
AcmeTrace
Acme Trace
This repository hosts the public releases of Acme traces from the Shanghai AI Lab, encompassing workloads spanning from March 2023 to August 2023. We encourage anyone to use the traces for academic purposes, and if you had any questions, feel free to send an email to us, or file an issue on Github.
Furthermore, we have conducted a thorough analysis of the Acme workloads, detailed in our NSDI '24 paper titled Characterization of Large Language Model Development in the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingAGree/AcmeTrace.RecSys_ACM_2026_Hallucinated
RecSys 2026 ACM challenge, Team: Hallucinated
Step 1: Download the full challenge datasets to obtail the following structure
data/talkpl-ai/TalkPlayData-Challenge-Blind-A
data/talkpl-ai/TalkPlayData-Challenge-Blind-B
data/talkpl-ai/TalkPlayData-Challenge-Dataset
data/talkpl-ai/TalkPlayData-Challenge-Track-Embeddings
data/talkpl-ai/TalkPlayData-Challenge-Track-Metadata # Both all_tracks and test_tracks
data/talkpl-ai/TalkPlayData-Challenge-User-Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/NicoloLocatelli/RecSys_ACM_2026_Hallucinated.stocks-ACMESOLAR-1D-candlesACM
Audio-Centric Multimodal Benchmark (ACM)
This dataset packages the ACM benchmark introduced by
Efficient and High-Fidelity Omni Modality Retrieval.
The paper is available at arXiv:2603.02098.
Dataset repo: chuonghm/ACM
ACM contains four HuggingFace subsets. Each subset uses a single test split
so query and candidate tables can keep their natural schemas:
composed_audio_retrieval_queries: flattened AT2A query rows.
composed_audio_retrieval_candidates: AT2A audio candidate pool.… See the full description on the dataset page: https://huggingface.co/datasets/chuonghm/ACM.acme10-he-ragapp-embeddingsD_llm2_run0_gen5_WXS_doc1000_synt64_lr1e-04_acm_MPP75pcLASTghostbuster_wphf10938-acme-rl01-data-01
Acme Sentiment Reviews (Multilingual)
Multilingual product reviews labeled for sentiment (English, French, German, Spanish).
Provenance: derived from the Acme Sentiment Corpus (Roy229/hf10938-acme-zz9q-data-01).
Compliance
License corrected from odc-by to cc-by-4.0 on 2026-08-24 to comply with the Acme Data Catalog license compliance policy. Derived datasets must carry the same license as the upstream corpus (cc-by-4.0).
filesystem-huggingface-9883-acme-support-tickets-0q1cs8jkghostbuster_reuteracm-icaif-2025_chunk_rankingjailbreaks_dataset_with_perplexity_bigcode_starcoder2-3b_bigcode_starcoder2-7beval_acm2_sorting_fin_50_3ghostbuster_essayhf10938-acme-zz9q-data-01
Acme Sentiment Corpus
The canonical Acme Sentiment Corpus contains product reviews labeled with sentiment polarity. It is the upstream source for several derivative datasets in the Acme Data Catalog.
Provenance: original dataset. All derivative datasets must inherit its license (cc-by-4.0).
hf10938-acme-zz9q-data-02
Acme Sentiment Reviews (English)
English product reviews labeled for sentiment, used for training text classifiers.
Provenance: derived from the Acme Sentiment Corpus (Roy229/hf10938-acme-zz9q-data-01).
Compliance
License corrected from unknown to cc-by-4.0 on 2026-08-24 to comply with the Acme Data Catalog license compliance policy. Derived datasets must carry the same license as the upstream corpus (cc-by-4.0).
hf10938-acme-rl01-data-02
Acme E-Commerce Feedback
Customer feedback from the Acme e-commerce platform, labeled by sentiment.
Provenance: derived from the Acme Sentiment Corpus (Roy229/hf10938-acme-zz9q-data-01).
