dom
Datasets
All datasets matching “dom”dom-pi-pdfs-2025
DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí)
PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí,
organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi
extraído por OCR/parsing. 41.617 PDFs · ~70 GB.
Dataset de texto derivado (carregável, com limpeza e tiers de qualidade):
gutoportelaa/dom-pi-corpus-2025.
Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.DOM
Dynamic Object Manipulation (DOM)
Project Page | Paper | Code
TL;DR: DOM is a large-scale dynamic manipulation dataset with 200K episodes, 2,800+ scenes, and 206 objects for training and evaluating VLA models.
Introduction
The Dynamic Object Manipulation (DOM) benchmark is designed to address the challenges of rapid perception and temporal anticipation in robotics. It includes:
200K synthetic episodes across 2,800+ scenes and 206 objects.
Support for evaluating VLA… See the full description on the dataset page: https://huggingface.co/datasets/hzxie/DOM.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.doMinhTri2003TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.Dome-Objaverse
Dome-Objaverse
Multi-view renders of 83,296 Objaverse objects — 48 views each, on a camera dome of 4 elevations × 12 azimuths.
The 48 views in order: three azimuth rings at elevations 0°, 30° and 60°, plus a top-down ring at 90°. The highlighted camera on the dome is the one that took the image on the left.
Rendering is the computationally demanding bottleneck of multi-view 3D datasets — typically over 100,000 CPU… See the full description on the dataset page: https://huggingface.co/datasets/zeyuanyin/Dome-Objaverse.
