datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.VietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/scarlettlin/VietPET-RoI.CS50-rawpi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.ovos-wake-word-bench-picovoice-view-glass
OVOS wake_word bench — picovoice-view-glass
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.vietprofs
VietProfs
VietProfs is a community-maintained directory of Vietnamese and Vietnamese-diaspora academics at universities and eligible public or nonprofit scholarly research institutes worldwide.
This repository publishes the current complete, unmodified roster from VietProfs. It is a JSON array whose records describe a person's current academic appointment and may include research areas, education, honors, public profile links, and portrait source URLs. Portrait image files are… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhvuh/vietprofs.vietquill-qcpg
Introduction
This is the primary training data source used to fine-tune VietQuill, the Vietnamese controlled paraphrase generation model presented in our paper.
The dataset is designed to support research on Vietnamese paraphrase generation, particularly controlled paraphrase generation with lexical, syntactic, and semantic transformations.
Citation
If you use VietQuill-QCPG in your research, please cite the corresponding VietQuill-QCPG paper or repository.… See the full description on the dataset page: https://huggingface.co/datasets/ngwgsang/vietquill-qcpg.vietnamese-evidence-retrieval-indexes
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f.
Rows: 52,605
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal
row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.vietnamese-evidence-retrieval-indexes-v2-r1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e.
Rows: 63,699
Source embedding shards: 13
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.traceweave-viewer-test
Agent Traces
Coding-agent sessions collected with TraceWeave,
rehydrated into the Claude Code JSONL schema
consumed by the Hugging Face Agent Trace Viewer.
Format
Each *.jsonl file at the dataset root is one session. Events use:
{"type":"user","message":{"role":"user","content":"..."},"uuid":"...","parentUuid":null,"sessionId":"...","timestamp":"..."}
{"type":"assistant","message":{"role":"assistant","content":[{"type":"text","text":"..."}]},"uuid":"..."… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/traceweave-viewer-test.Vietnamese-Openorca-Multiplechoice-gg-translatedvietnamese-evidence-retrieval-indexes-v3-1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
aiMy144/vietnamese-evidence-corpus-embeddings-e5-large-v3-1 at revision ea0826c0a9d4273eec267eb66b6dbbfe64d5fb11.
Rows: 53,747
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/vietnamese-evidence-retrieval-indexes-v3-1.vietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.vietnam-real-estate-listings
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
vietnam_real_estate_listings
This dataset contains over one million Vietnamese real estate listings, primarily featuring detailed textual descriptions of properties for sale or rent. Each entry includes structured attributes such as location (province, district, ward, street), property specifications (area, price, floor count, room counts), and directional orientation. The… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/vietnam-real-estate-listings.vietnamese-evidence-corpus-chunked
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.ViePAWSvietnamese-evidence-corpus-embeddings-e5-large-v2vietquill-qcpg-100k-synthesis-sentencevietquill-qcpg-100k-synthesis-questionvietnamese-evidence-corpus-chunked-e5-v2
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.viettelsecurity-ai__security-llama3.2-3b-details
Dataset Card for Evaluation run of viettelsecurity-ai/security-llama3.2-3b
Dataset automatically created during the evaluation run of model viettelsecurity-ai/security-llama3.2-3b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/viettelsecurity-ai__security-llama3.2-3b-details.deepfake-real-frame-views-hardneg-v2short_video_ocr_viewer_test
Minimal Dataset Viewer test
This repository intentionally contains one tiny JSON data file. The YAML
configuration explicitly tells Hugging Face Dataset Viewer which file and split
to index.
viereckvietnamese_ultrafeedback_binarizedAIDetection_Vietnamese_HumanData
