datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dv_syn_speech_md
Dataset Card for Dhivehi Speech - Medium
Dataset Summary
This small dataset contains audio, and sentence in dhivehi.
Supported Tasks and Leaderboards
Automatic Speech Recognition
Text-to-Speech
Languages
Dhivehi
Dataset Structure
TODO
Data Instances
A typical data point comprises the path to the audio file and its sentence.
Data Fields
TODO
PDM-Lite-DVS
PDM-Lite-DVS
PDM-Lite-DVS is an independently collected synthetic CARLA 0.9.15 event-camera
dataset generated with the rule-based PDM-Lite expert on route configurations
published by carla_garage. The release lineage is 5,545 public route XMLs →
5,503 recordings in the frozen local pool → 257 rejected recordings → 5,246
published recordings. “Public route release” refers only to the route
configurations: the sensor measurements are an independent collection. This
is not the… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/PDM-Lite-DVS.LEAD-DVS
LEAD-DVS
LEAD expert-driving data collected in CARLA 0.9.15 with synchronized RGB,
depth, semantic/instance segmentation, LiDAR, radar, HD map, metadata, 3D
bounding boxes, and a forward-facing DVS event camera.
Code release
Dataset construction, preprocessing, training, and evaluation code will be
published in SamanthaZhang-stu/ReflexWorldModel.
This GitHub repository is the designated code-release location for the project.
Contents
1,579… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/LEAD-DVS.LEGIT-VIPER-Jigsaw-Toxic-Comment-Perturbeddv_syn_speech_md
Dataset Card for Dhivehi Speech - Medium
Dataset Summary
This small dataset contains audio, and sentence in dhivehi.
Supported Tasks and Leaderboards
Automatic Speech Recognition
Text-to-Speech
Languages
Dhivehi
Dataset Structure
TODO
Data Instances
A typical data point comprises the path to the audio file and its sentence.
Data Fields
TODO
umlsdv_syn_speech_sm
Dataset Card for Dhivehi Speech - Small
Dataset Summary
This small dataset contains audio, and sentence in dhivehi.
Supported Tasks and Leaderboards
Automatic Speech Recognition
Text-to-Speech
Languages
Dhivehi
Dataset Structure
TODO
Data Instances
A typical data point comprises the path to the audio file and its sentence.
Data Fields
TODO
dv-synthetic-errors-mixedDV Text Errors
Dhivehi text error correction dataset containing correct sentences and synthetically generated errors.
The dataset aims to test Dhivehi language error correction models and tools.
About Dataset
Task: Text error correction
Language: Dhivehi (dv)
Dataset Structure
Input-output pairs of Dhivehi text:
correct: Original correct sentences
incorrect: Sentences with synthetic errors
Note: This is replica of alakxender/dv-synthetic-errors: added more synthetic errors. x5
dv-synthetic-errors
DV Text Errors
Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools.
About Dataset
Task: Text error correction
Language: Dhivehi (dv)
Dataset Structure
Input-output pairs of Dhivehi text:
correct: Original correct sentences
incorrect: Sentences with synthetic errors
Statistics
Train set: {train_examples} examples… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors.dv_subjectFor more information about this data refer the main repository for the supplementary material of the manuscript Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research.
ASL-DVS
ASL-DVS
This is a denoised, windowed derivative of the ASL-DVS event-camera dataset.
Each row is a 1 second asynchronous event window from the original DAVIS240C
recordings.
The original events are preserved in the source sensor coordinate system
(240x180). Events are not converted to frames and are not cropped.
For convenience, each row includes a recommended 128x128 crop location as
metadata only.
Upstream Dataset Credit
The original ASL-DVS dataset was… See the full description on the dataset page: https://huggingface.co/datasets/papo1011/ASL-DVS.transito-hn-retrieval-eval
Tránsito HN Retrieval Eval
A small, source-verifiable benchmark for article-level retrieval in Honduran law.
Given a masked excerpt from a Supreme Court ruling and the version of the Traffic Act in force on that date, can an embedding model retrieve an article the court cited?
Why we built it
General-purpose leaderboards help shortlist embedding models, but they cannot tell us which compact models work well on Honduran legal text.
We built this benchmark while… See the full description on the dataset page: https://huggingface.co/datasets/DVSGlobal/transito-hn-retrieval-eval.dv_syn_speech_sm
Dataset Card for Dhivehi Speech - Small
Dataset Summary
This small dataset contains audio, and sentence in dhivehi.
Supported Tasks and Leaderboards
Automatic Speech Recognition
Text-to-Speech
Languages
Dhivehi
Dataset Structure
TODO
Data Instances
A typical data point comprises the path to the audio file and its sentence.
Data Fields
TODO
dv-synthetic-errors-lg
Dhivehi Correction Dataset
Dataset Description
This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts.
Dataset Summary
The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of:
A correct Dhivehi sentence
The same sentence with synthetic errors
Data Splits
The dataset is split into:
Train: 80% (~5.8M pairs)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors-lg.LEGIT
Dataset Card for "LEGIT-2023"
Label key:
0 or 1: word 0 or 1 is more legible, other unknown
2: both words are equally legible
3: neither word is legible
dv-synthetic-errors
DV Text Errors
Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools.
About Dataset
Task: Text error correction
Language: Dhivehi (dv)
Dataset Structure
Input-output pairs of Dhivehi text:
correct: Original correct sentences
incorrect: Sentences with synthetic errors
Statistics
Train set: {train_examples} examples… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors.dvsr-1786923253dv-synthetic-errors-lg
Dhivehi Correction Dataset
Dataset Description
This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts.
Dataset Summary
The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of:
A correct Dhivehi sentence
The same sentence with synthetic errors
Data Splits
The dataset is split into:
Train: 80% (~5.8M pairs)… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors-lg.dv-syn-female2-for-tts
