art
Datasets
All datasets matching “art”multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.VintixDatasetII
Vintix II Cross-Domain ICRL Dataset
Dataset Summary
This dataset is a large-scale cross-domain benchmark for in-context reinforcement learning and continuous control. It was introduced with Vintix II and covers a diverse set of tasks spanning robotic manipulation, dexterous control, locomotion, energy management, industrial process control, autonomous driving, and other control settings.
The training set contains 209 tasks across 10 domains, totaling 3.8M episodes and… See the full description on the dataset page: https://huggingface.co/datasets/artfawl/VintixDatasetII.artem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.VaaniVAANI is an India-representative multi-modal multi-lingual dataset.
The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages.
From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts.
Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.art
Dataset Card for "art"
Dataset Summary
ART consists of over 20k commonsense narrative contexts and 200k explanations.
The Abductive Natural Language Inference Dataset from AI2.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
anli
Size of downloaded dataset files: 5.12 MB
Size of the generated dataset: 34.36 MB
Total amount of disk used: 39.48… See the full description on the dataset page: https://huggingface.co/datasets/allenai/art.OCR-Synthetic-Multilingual-v1
OCR-Synthetic-Multilingual-v1
Overview
Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al.
This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/Arturito1/OCR-Synthetic-Multilingual-v1.
