datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-Knowledgedroid-trajectories-5k
DROID Visual Trajectories (5K Episodes) — LeRobot Format
Visual object trajectories extracted from 4506 episodes of the
DROID dataset using
Gemini Robotics-ER + SAM2.
How to load images/videos
Images are not embedded in this dataset. Each row contains reference paths
to the source videos in cadene/droid_1.0.1:
from datasets import load_dataset
from huggingface_hub import hf_hub_download
ds = load_dataset("ASethi04/droid-trajectories-5k", split="train")
row =… See the full description on the dataset page: https://huggingface.co/datasets/ASethi04/droid-trajectories-5k.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.pie-synthetic
PIE synthetic dataset
Repo: https://github.com/awasthiabhijeet/PIE
Paper: https://aclanthology.org/D19-1435.pdf
ImageNet-SampledBSDSmerlin
MERLIN corpus
Project URL: https://merlin-platform.eu/C_mcorpus.php
Dataset URL: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/6
The MERLIN corpus is a written learner corpus for Czech, German, and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus contains learner texts produced in standardized language certifications covering CEFR levels A1-C1. The MERLIN annotation… See the full description on the dataset page: https://huggingface.co/datasets/aseifert/merlin.mcqa_greek_asep
Dataset Card for Multiple Choice QA Greek ASEP
Dataset Details
Dataset Description
The Multiple Choice QA Greek ASEP dataset is a set of 2346 multiple choice questions in Greek. The questions were extracted and converted from questions available at the website of the Greek Supreme Council for Civil Personnel Selection (Ανώτατο Συμβούλιο Επιλογής Προσωπικού, ΑΣΕΠ-ASEP). The dataset combines materials from the 1Γ/2025 examination and the updated 2026… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/mcqa_greek_asep.robot-trajectories-datasetA.S.E
A.S.E
AICGSecEval (A.S.E) is a project-level benchmark for evaluating the security of AI-generated code. It is designed to assess the security performance of AI-assisted programming by simulating real-world development workflows:
Code Generation Tasks – Derived from real-world GitHub projects and authoritative CVE patches
Project-Level Context – Each sample includes retrieval-based code context (relevant files + function summaries) to simulate realistic AI programming scenarios… See the full description on the dataset page: https://huggingface.co/datasets/tencent/A.S.E.ASearcher_en_no-math_Qwen3-8B-reject-sampleSee our blog for details:
Cut the Bill, Keep the Turns: Affordable Multi-Turn Search RL
This dataset is originally from https://huggingface.co/datasets/inclusionAI/ASearcher-train-data
We do filtering to original data:
Remove Chinese samples
Our wiki server does not handle Chinese retrieval well and may return garbled text;
There are about 2k Chinese-related samples in the ASearcher dataset, which we remove entirely.
Remove math problems
Use formula patterns / specific regexes to filter… See the full description on the dataset page: https://huggingface.co/datasets/aidenjhwu/ASearcher_en_no-math_Qwen3-8B-reject-sample.asean-geolocation
ASEAN Geolocation
ASEAN Geolocation is an image dataset for country-level geolocation across the 11 ASEAN countries. It is designed for image classification, visual geolocation, and explainable geolocation research.
The dataset contains 4850 images organized by country label. Images are provided at their original aspect ratios; model-specific resizing and normalization should be handled during preprocessing.
Dataset summary
Task: country-level image… See the full description on the dataset page: https://huggingface.co/datasets/0xRafie/asean-geolocation.Tiriyo-New-Testament-CorpusThis dataset contains a translation of the New Testament into the Tiriyó language (spoken in Brazil and Suriname, in South America). Here is the format I follow:
every Now Testament book is placed in a .txt file, with the filename identifying the book (e.g., Matthew.txt)
inside every file, each line starts with XX:YY, two two-digit numbers separated by a colon, with XX identifying the chapter and YY the verse (sometimes XX:YY-ZZ, when two or more verses were collapsed into one in the Tiriyó… See the full description on the dataset page: https://huggingface.co/datasets/Asehpe/Tiriyo-New-Testament-Corpus.khp-youth-mental-health-guardrail
KHP Youth Mental Health Safety Guardrail Dataset
A synthetic multi-turn conversational dataset for training and evaluating input guardrails for AI assistants serving youth in mental distress. The dataset targets 9 distress signal categories and provides a binary high-risk label for classification, decoupled from the individual signal flags.
GitHub repository: Aser97/Guardrail-For-Agents
Dataset Summary
Split
Rows
High-risk
Low-risk
Train
~1578
~789
~789… See the full description on the dataset page: https://huggingface.co/datasets/AserLompo/khp-youth-mental-health-guardrail.cis5190-fox-nbc-headlines
License
Code and metadata under MIT. Headlines retain their original copyrights.
aym-bireysel-basvuru-kararlari
AYM Bireysel Başvuru Kararları
A complete corpus of individual-application (bireysel başvuru) decisions
of the Turkish Constitutional Court (Anayasa Mahkemesi, AYM), scraped from
kararlarbilgibankasi.anayasa.gov.tr,
parsed into structured fields, and enriched with citation links between
decisions.
Dataset summary
17,015 AYM individual-application decisions (full corpus as of the
scrape date).
Each decision is split into a facts section (vakia) and an
evaluation… See the full description on the dataset page: https://huggingface.co/datasets/aselimgul/aym-bireysel-basvuru-kararlari.chris_er_traces_top1_v1asearcher_short_form_rlvr_with_system_promptCAD-SResumeCredibilityAssessmentDataset
CAD-S: Resume Credibility Assessment Dataset
📌 Overview
CAD-S (Credibility Assessment Dataset ) is the first openly available dataset designed specifically for resume credibility assessment using Natural Language Processing techniques.
The dataset supports supervised learning for detecting inconsistencies between claimed skills and supporting evidence (e.g., projects, experience) within resumes.
This dataset is intended for:
Resume verification systems
Natural… See the full description on the dataset page: https://huggingface.co/datasets/aselasperera/CAD-SResumeCredibilityAssessmentDataset.FineTome-100k-week02-splitkorean-petitionsFineTome-100k-week02-eval500distilabel-aset-qa-30ragpapersdatasetASearcher-SmallPartyNLI
PartyNLI 🥳
PartyNL is an extension of LIveNLI designed to study label and explanation variation --- in both humans and LLMs. In addition to the original 122 LiveNLI items, PartyNLI also features new label annotations and free-text explanations from a variety of LLMs.
architectural-spatial-blindspots-smolvlm
Architectural and Spatial Blind Spots of SmolVLM
Model Tested: HuggingFaceTB/SmolVLM-BaseParameter Count: 2.2B
1. How I loaded the model (Python Code)
I loaded the model in a Google Colab environment using the transformers library with 4-bit quantization to fit within a free-tier GPU.
from transformers import AutoProcessor, AutoModelForVision2Seq
import torch
model_id = "HuggingFaceTB/SmolVLM-Base"
processor = AutoProcessor.from_pretrained(model_id)
model =… See the full description on the dataset page: https://huggingface.co/datasets/AsefaH/architectural-spatial-blindspots-smolvlm.SynGlassaseanllama2-without-emojisaseanllama2-instruct-without-emojiscpp_emb_nl_to_code
