datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edacc
EdAcc: The Edinburgh International Accents of English Corpus
The Edinburgh International Accents of English Corpus (EdAcc) is a new automatic speech recognition (ASR) dataset
composed of 40 hours of English dyadic conversations between speakers with a diverse set of accents. EdAcc includes a
wide range of first and second-language varieties of English and a linguistic background profile of each speaker.
Results on latest public, and commercial models show that EdAcc highlights… See the full description on the dataset page: https://huggingface.co/datasets/edinburghcstr/edacc.dlr_edan_shared_control_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "dlr_edan",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 10,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/dlr_edan_shared_control_lerobot.eda2-159mhz
EDA2 159 MHz Southern-Sky Maps
This dataset contains four products from the EDA2 159 MHz southern-sky
survey: maps made with and without a large-scale sky prior, a noise map,
and a spectral-index map against the Haslam 408 MHz survey. The two
intensity maps and the noise map cover the complete HEALPix grid; the
spectral-index map has 279 pixels carrying the source HEALPix UNSEEN value.
Intended uses
The maps support low-frequency sky and foreground studies… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/eda2-159mhz.pinecone_test
Dataset Card for "pinecone_test"
More Information needed
dlr_edan_shared_controlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 14,
"total_videos": 104,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_edan_shared_control.HARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.e-daic-ai-controllededacc-l1cls
EdAcc
The Edinburgh International Accents of English Corpus (EdAcc) is a speech dataset designed to evaluate automatic speech recognition (ASR) systems on a wide range of global English accents.
Citation
@inproceedings{sanabria23edacc,
title="{The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR}",
author={Sanabria, Ramon and Bogoychev, Nikolay and Markl, Nina and Carmantini, Andrea and Klejch, Ondrej and Bell… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/edacc-l1cls.EvolveR-NQ-HotpotQAThis repository contains the dataset associated with the paper EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.
EvolveR is a framework designed to enable LLM agents to self-improve through a complete, closed-loop experience lifecycle. This lifecycle comprises two key stages: (1) Offline Self-Distillation, where interaction trajectories are synthesized into reusable principles, and (2) Online Interaction, where the agent retrieves these principles to guide… See the full description on the dataset page: https://huggingface.co/datasets/Edaizi/EvolveR-NQ-HotpotQA.edacc
Draft conversion of EdAcc
Final dataset will be moved to the edinburghcstr organisation.
edacc_testedacc_dataset_resampled_16Kedacc_test_cleanOXE_dlr_edan_shared_control_embeddingsLanguage Table (LeRobot) — Embedding-Only Release
(DINOv3 + SigLIP2 image features; EmbeddingGemma task-text features)
This repository packages a re-encoded variant of IPEC-COMMUNITY/dlr_edan_shared_control_lerobot where raw videos are replaced by fixed-length image embeddings, and task strings are augmented with text embeddings. All indices, splits, and semantics remain consistent with the source dataset while storage and I/O are substantially lighter. To make the dataset practical to… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/OXE_dlr_edan_shared_control_embeddings.benchmark
EditJudge-Bench
EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as
automated judges for image-edit verification. Each row contains a source image,
an edited image, a factual edit instruction, counterfactual instructions, and
ground-truth scene parameters produced by a controlled Blender/Infinigen
generation pipeline.
This repository is an anonymous review release for a NeurIPS Evaluations and
Datasets submission.
Dataset Contents
1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.africa-namibia-namibia-accessibility-indicators-edad722d
Namibia - Accessibility Indicators | Africa (Namibia official open data)
6,600 rows - 1 Africa country - 2020-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Namibia as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Namibia - Accessibility Indicators
Publisher:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-namibia-namibia-accessibility-indicators-edad722d.fr_crawler2test-repoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 3464,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Edautel/test-repo.psych_aimirea-tl-eda
RTU MIREA Telegram Channel EDA
The dataset contains 22 metrics, describing post user engagement, linguistic units features, readability, AI messages generation score, semantic topic label, and model labelling confidence for the whole available time period at the time of analysis (over 10k text messages).
Overview
The dataset is the result of EDA performed on the data from RTU MIREA Telegram channel, retreived via aiogram. The full code for the data processing… See the full description on the dataset page: https://huggingface.co/datasets/complicat9d/mirea-tl-eda.edaiceval_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 419,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Edautel/eval_test.edaic-ponly-vad-w2v-erpscyh_ai_birlesmiseval_pickup-cube-demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 575,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/eda-amd/eval_pickup-cube-demo.EDA_SREDA_SR_Vietnameseedacc_processeddlr_edan_shared_control_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "dlr_edan",
"total_episodes": 104,
"total_frames": 8928,
"total_tasks": 10,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:104"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/FedorX8/dlr_edan_shared_control_lerobot.pickup-cube-demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 7,
"total_frames": 4193,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/eda-amd/pickup-cube-demo.
