datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.doom-dataset-largedoom-rnd-largeVietnam-Celeb
unofficial mirror of Vietnam-Celeb dataset
official announcement:
https://www.isca-archive.org/interspeech_2023/pham23b_interspeech.html
https://github.com/Vietnam-Celeb/Vietnam-Celeb
https://huggingface.co/datasets/hustep-lab/Vietnam-Celeb
official download:
Part 0: https://drive.google.com/file/d/1pMuT3DFzSwib7SVcRS8VkDwPuLTsemSG/view?usp=share_link
Part 1: https://drive.google.com/file/d/1xayHt2HRqE1aJ4HvtUT40_9XlgvfDfRY/view?usp=share_linkPart 2:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/Vietnam-Celeb.crowd-code-dataset-1.0
Install crowd-code 2.0 to help crowd-source the next-generation coding dataset.
crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.vlsp2020_vinai_100h
unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset
official announcement:
tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/
in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/
VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020
official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing
contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.doocs-leetcode-solutions
Doocs LeetCode Solutions
A comprehensive dataset of LeetCode problems and solutions created from the Doocs LeetCode repository. This dataset is designed for fine-tuning large language models to understand programming problems and generate code solutions.
Description
Repository: Doocs LeetCode Solutions
Total Problems: 3500+
Total Solutions: 15,000+ (across multiple languages)
Size: ~60 MB (Parquet format)
Languages:
C
Cangjie
C++
C#
Dart
Go
Java
JavaScript
Kotlin
Nim
PHP… See the full description on the dataset page: https://huggingface.co/datasets/olegshulyakov/doocs-leetcode-solutions.doom-dataset-1LSVSC
unofficial mirror of LSVSC dataset (novel large-scale Vietnamese speech corpus)
official announcement: https://www.mdpi.com/2079-9292/13/5/977
official download: https://drive.google.com/drive/folders/1tiPKaIOC7bt6isv5qFqf61O_2jFK8ZOI
100h, 57k samples
pre-process: see my code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/clean-lsvsc.py
need to do: check misspelling, restore foreign words phonetised to vietnamese
usage with HuggingFace:
# pip install -q… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/LSVSC.gpt_oss_20b_doorkey_boundary_activationsSEED-Bench-2-Plusfrom https://huggingface.co/datasets/AILab-CVC/SEED-Bench-2-plus
SEED-Bench-2-Plus Card
Benchmark details
Benchmark type: SEED-Bench-2-Plus is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 2.3K multiple-choice questions with precise human annotations, spanning three broad categories: Charts, Maps, and Webs, each of which covers a wide spectrum of text-rich scenarios in the real world.
Benchmark date: SEED-Bench-2-Plus was collected in April 2024.… See the full description on the dataset page: https://huggingface.co/datasets/doolayer/SEED-Bench-2-Plus.DoorBench
DoorBench
1000 procedural articulated doors for robot simulation, with MJCF, URDF and USD exports, Blender appearances, and reference motion.
Interactive catalogue · Source and tools · Release guide
Version v2026.09.05 contains 1014 saved Blender images, including 14 images rendered with the higher sample preset, and reference clips/native trajectories for 1000 doors. The release manifest, complete per-file SHA256 inventory, source hashes and immutable Hub revision identify the… See the full description on the dataset page: https://huggingface.co/datasets/adamraudonis/DoorBench.LawBenchhttps://github.com/open-compass/LawBench/blob/main/README_EN.md
infore1_25hours
unofficial mirror of InfoRe Technology public dataset №1
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
25h, 14.9k samples, InfoRe paid a contractor to read text
official download: magnet:?xt=urn:btih:1cbe13fb14a390c852c016a924b4a5e879d85f41&dn=25hours.zip&tr=http%3A%2F%2Foffice.socials.vn%3A8725%2Fannounce
mirror: https://files.huylenguyen.com/datasets/infore/25hours.zip
unzip password: BroughtToYouByInfoRe
pre-process: see… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore1_25hours.klue-mrc-bm25base4-mobile-door-eef-merged-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
10
],
"names": {
"motors": [
"joint1",
"joint2",
"joint3",
"joint4",
"joint5"… See the full description on the dataset page: https://huggingface.co/datasets/L5vel/base4-mobile-door-eef-merged-v30.crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.hate_speech_labeledfpt_fosd
unofficial mirror of FPT Open Speech Dataset (FOSD)
released publicly in 2018 by FPT Corporation
100h, 25.9k samples
official link (dead): https://fpt.ai/fpt-open-speech-data/
mirror: https://data.mendeley.com/datasets/k9sxg2twv4/4
DOI: 10.17632/k9sxg2twv4.4
pre-process:
remove non-sense strings: -N \r\n
remove 4 files because missing transcription:
Set001_V0.1_008210.mp3
Set001_V0.1_010753.mp3
Set001_V0.1_011477.mp3
Set001_V0.1_011841.mp3
need to do: check misspelling
usage… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/fpt_fosd.door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805
Outdoor Door Handle Push/Pull Training Set
Synthetic outdoor dataset staged in an alleyway to train a robot to detect door handles and determine whether a door must be pushed or pulled. 20 renders at 1024x1024 with albedo and metric depth passes, per-frame annotations, and authored lighting.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/door-opening-manipulation-training-pack-next-pack-fbe147bb-03e76805.modern_music_reDatasets for Relation Extraction TaskSource from Wikipedia (CC-BY-2.0)Contributors : Doohae Jung, Hyesu Kim, Bosung Kim, Isaac Park, Miwon Jeon, Dagon Lee, Jihoo Kim
base4-mobile-door-01-BC-FV-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
10
],
"names": {
"motors": [
"joint1",
"joint2",
"joint3",
"joint4",
"joint5"… See the full description on the dataset page: https://huggingface.co/datasets/maskjp/base4-mobile-door-01-BC-FV-v30.base4-mobile-door-02-BC-FV-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
10
],
"names": {
"motors": [
"joint1",
"joint2",
"joint3",
"joint4",
"joint5"… See the full description on the dataset page: https://huggingface.co/datasets/maskjp/base4-mobile-door-02-BC-FV-v30.doodlecraft_cleanbase4-mobile-door-04-BC-FV-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
10
],
"names": {
"motors": [
"joint1",
"joint2",
"joint3",
"joint4",
"joint5"… See the full description on the dataset page: https://huggingface.co/datasets/maskjp/base4-mobile-door-04-BC-FV-v30.outdoor-door-handle-push-pull-training-set-next-pack-1c340bea-fea6c091
Door Handle Detection & Push/Pull Direction — Training Pack
Training dataset to teach a robot to locate door handles and infer whether each door must be pushed or pulled to open. Renders staged in home and office environments (office, home entrance, break area) that frame doors with various handle types. Outputs RGB with metric depth and world-space normals for contact geometry, per-frame annotations (bounding boxes, camera pose), authored lighting, at 1024×1024. Targets object… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/outdoor-door-handle-push-pull-training-set-next-pack-1c340bea-fea6c091.doom-dense-arnold
DoomDiT dense Arnold recordings
Lossless, per-tic recordings of the Arnold agent (Lample and Chaplot, AAAI 2017) playing ViZDoom deathmatch
with 8 bots on Freedoom assets, made for training action-conditioned world models. Every engine tic
(35 per second) is stored with the executed control vector, so the data can be used at any frame stride.
Recorded September 2026 at CMU for the DoomDiT project (Rohan Nagabhirava, Keerthana Chirumamilla).
What is here… See the full description on the dataset page: https://huggingface.co/datasets/RohanNaga/doom-dense-arnold.VietMed_labeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) labeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the labeled set: 9.2k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-labeled.py
need to do: check misspelling, restore foreign words phonetised to vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_labeled.dooggies
Dataset Card
Disclaimer
All rights belong to their owners.
Models and datasets can be removed from the site at the request of the copyright holder.
Dataset Summary
NFT images dataset for unconditional generation.
NFT collection available here.
Model is available here.
Check Space: link.
Supported Tasks and Leaderboards
More Information Needed
How to use
How to load this dataset directly with the datasets library:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/huggingnft/dooggies.
