datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-Knowledgesampled-local-resumes
sampled-local-resumes
This dataset contains synthetic resume data sampled from local folders (20% sample from each folder).
License
This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details.
Attribution
Copyright 2025 Fairly AI Inc. dba Asenion
This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0.
You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.ASearcher-test-dataaseem00r_storedroid-trajectories-5k
DROID Visual Trajectories (5K Episodes) — LeRobot Format
Visual object trajectories extracted from 4506 episodes of the
DROID dataset using
Gemini Robotics-ER + SAM2.
How to load images/videos
Images are not embedded in this dataset. Each row contains reference paths
to the source videos in cadene/droid_1.0.1:
from datasets import load_dataset
from huggingface_hub import hf_hub_download
ds = load_dataset("ASethi04/droid-trajectories-5k", split="train")
row =… See the full description on the dataset page: https://huggingface.co/datasets/ASethi04/droid-trajectories-5k.A-secret-data-repoASearcher-train-data
.hf-sanitized.hf-sanitized-6N1eYoZHVQFjCw0x7dytd .no-border-table table, .hf-sanitized.hf-sanitized-6N1eYoZHVQFjCw0x7dytd .no-border-table th, .hf-sanitized.hf-sanitized-6N1eYoZHVQFjCw0x7dytd .no-border-table td { border: none !important; }aiocr_asistantHumanStereoPreview
HumanStereo Preview
This repository contains a non-sharded preview subset of the HumanStereo training split for NeurIPS review.
Sessions: 1438
Staged size: 3.801 GiB
Source dataset: asergiu/HumanStereo
Manifest: manifest/sessions.csv
Croissant metadata: metadata.json
The full dataset is available separately as asergiu/HumanStereo.
ASearcher
ASearcher: An Open-Source Large-Scale
Reinforcement Learning Project for Search Agents
| 📰 Paper | 🤗 Datasets | 🤗 Models |
Introduction
ASearcher is an open-source framework designed for large-scale online reinforcement learning (RL) training of search agents. Our mission is to advance Search Intelligence to expert-level performance. We are fully committed to open-source by releasing model weights, detailed training methodologies, and data synthesis pipelines. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/ASearcher.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.wifi-csi-human-activities
WiFi CSI Human Activities Dataset
This repository contains WiFi Channel State Information (CSI) measurements, intermediate processing outputs, extracted features, visualizations, and experimental results used for WiFi-based Human Activity Recognition (HAR).
The dataset accompanies the Master's thesis:
Assel Ussenova
WiFi Sensing Through Digital Receive Beamforming and CSI
MSc in ICT and Internet Engineering
Università degli Studi di Roma Tor Vergata (2024/2025)… See the full description on the dataset page: https://huggingface.co/datasets/aselya9185/wifi-csi-human-activities.pie-synthetic
PIE synthetic dataset
Repo: https://github.com/awasthiabhijeet/PIE
Paper: https://aclanthology.org/D19-1435.pdf
UR10_RTDE_SpaceMouse_EE_pick_and_place_object_tabletopThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur10_rtde",
"total_episodes": 26,
"total_frames": 7557,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:26"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/UR10_RTDE_SpaceMouse_EE_pick_and_place_object_tabletop.roboverse-mujoco-franka-pick_cube-E200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 251,
"total_frames": 29080,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:251"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/roboverse-mujoco-franka-pick_cube-E200.ImageNet-SampledBSDSmerlin
MERLIN corpus
Project URL: https://merlin-platform.eu/C_mcorpus.php
Dataset URL: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/6
The MERLIN corpus is a written learner corpus for Czech, German, and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus contains learner texts produced in standardized language certifications covering CEFR levels A1-C1. The MERLIN annotation… See the full description on the dataset page: https://huggingface.co/datasets/aseifert/merlin.mcqa_greek_asep
Dataset Card for Multiple Choice QA Greek ASEP
Dataset Details
Dataset Description
The Multiple Choice QA Greek ASEP dataset is a set of 2346 multiple choice questions in Greek. The questions were extracted and converted from questions available at the website of the Greek Supreme Council for Civil Personnel Selection (Ανώτατο Συμβούλιο Επιλογής Προσωπικού, ΑΣΕΠ-ASEP). The dataset combines materials from the 1Γ/2025 examination and the updated 2026… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/mcqa_greek_asep.robot-trajectories-datasetroboverse-mujoco-franka-pick_cube-E100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 121,
"total_frames": 13779,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:121"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/roboverse-mujoco-franka-pick_cube-E100.roboverse-pick_cube-test-ep50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 62,
"total_frames": 7218,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:62"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/roboverse-pick_cube-test-ep50.UR5_RTDE_SpaceMouse_EE_pick_and_place_object_DifficultThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_rtde",
"total_episodes": 140,
"total_frames": 43568,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:140"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/UR5_RTDE_SpaceMouse_EE_pick_and_place_object_Difficult.UR5_RTDE_SpaceMouse_EE_pick_and_place_object_Easy_Difficult_filteredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_rtde",
"total_episodes": 110,
"total_frames": 34032,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 15,
"splits": {
"train": "0:110"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/UR5_RTDE_SpaceMouse_EE_pick_and_place_object_Easy_Difficult_filtered.roboverse-pick_cube-mujocoE100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"panda_finger_joint1",
"panda_finger_joint2",
"panda_joint1",
"panda_joint2",
"panda_joint3"… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/roboverse-pick_cube-mujocoE100.UR5_RTDE_SpaceMouse_EE_pick_and_place_object_Easy_Difficult_filtered_plus_last20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_rtde",
"total_episodes": 130,
"total_frames": 40226,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1,
"fps": 15,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/UR5_RTDE_SpaceMouse_EE_pick_and_place_object_Easy_Difficult_filtered_plus_last20.roboverse-pick_cube-genesis-E100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"panda_finger_joint1",
"panda_finger_joint2",
"panda_joint1",
"panda_joint2",
"panda_joint3"… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/roboverse-pick_cube-genesis-E100.PandaPickCubeSpacemouseRandom2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 2706,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ases200q2/PandaPickCubeSpacemouseRandom2.A.S.E
A.S.E
AICGSecEval (A.S.E) is a project-level benchmark for evaluating the security of AI-generated code. It is designed to assess the security performance of AI-assisted programming by simulating real-world development workflows:
Code Generation Tasks – Derived from real-world GitHub projects and authoritative CVE patches
Project-Level Context – Each sample includes retrieval-based code context (relevant files + function summaries) to simulate realistic AI programming scenarios… See the full description on the dataset page: https://huggingface.co/datasets/tencent/A.S.E.Aseel_Arabic_Dataset
Arabic Speech Dataset
Contains 10000 24kHz Mono WAV clips.
