CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vladimir-io /healthspend-datatextn<1K0 likes1.6k downloads27d agoHugging Face02UCSC-VLAA /MedReason MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs 📃 Paper |🤗 MedReason-8B | 📚 MedReason Data ✨ Latest News [05/27/2025] 🎉 MedReason wins 3rd prize🏆 in the Huggingface Reasoning Datasets Competition! ⚡Introduction MedReason is a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs). We utilize a structured medical knowledge… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedReason.textquestion-answering10K<n<100K88 likes1k downloads1y agoHugging Face03ganlinyang /Vlasertabular10K<n<100K0 likes930 downloads6mo agoHugging Face04yuuu94 /W2-VLA-CoT World-to-Wrist: Offline CoT Labels This dataset contains frame-aligned offline chain-of-thought annotations used to train W²-VLA policies on LIBERO, RoboTwin, and four real-world manipulation tasks. Matching LeRobot action data is available in W2-VLA-Training-Data. Dataset Structure W2-VLA-CoT/ ├── libero/ │ ├── libero_10_no_noops_1.0.0_lerobot/ │ ├── libero_goal_no_noops_1.0.0_lerobot/ │ ├── libero_object_no_noops_1.0.0_lerobot/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/yuuu94/W2-VLA-CoT.textn<1K2 likes803 downloads21d agoHugging Face05Thirteen13tj /Robot-VLA-R1text100K<n<1M2 likes603 downloads1y agoHugging Face06UCSC-VLAA /ClinSeek-Bench ClinSeek-Bench ClinSeek-Bench is the evaluation suite introduced in ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning. It evaluates clinical reasoning under two paired settings with the same task definitions and answer labels: Curated Input: the model answers from the evidence package provided by the source benchmark. Automated Evidence-Seeking: the curated context is removed, and the model must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.tabular1K<n<10K2 likes300 downloads24d agoHugging Face07Falcon110120 /vla-reasoningtext100K<n<1M0 likes250 downloads1y agoHugging Face08shreethar /thinkflow-vla-features-b2tabular1K<n<10K0 likes131 downloads2mo agoHugging Face09vladimirbesk /tsiolkovsky-papers Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky (1857–1935), who derived the rocket equation and described the multistage rocket decades before anyone could test either. The archive had been scanned and put online, but without a catalogue you could query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.tabulartext-generation10K<n<100K2 likes99 downloads1mo agoHugging Face10UCSC-VLAA /VLM-CapCurriculum-Perception-Data VLM-CapCurriculum-Perception (D_perc) Stage-1 visual perception data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). Each sample is a 4-way multiple-choice question over an image where the question can be answered from a fine-grained image caption but is missed by a strong VLM looking only at the image — by construction, these samples isolate perception failures from… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-Perception-Data.textvisual-question-answering1K<n<10K0 likes98 downloads4mo agoHugging Face11VLAI-AIVN /ViVQA-X Dataset Card for ViVQA-X Dataset Description ViVQA-X is the Vietnamese version of the VQA-X dataset, designed for tasks in VQA-NLE in the Vietnamese language. This dataset was created to support research in Vietnamese VQA, offering translated examples with explanations, in alignment with the original English dataset. The ViVQA-X dataset is generated using a multi-stage pipeline for translation and evaluation, including both automated translation and post-processing, with… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/ViVQA-X.textvisual-question-answering10K<n<100K13 likes71 downloads2y agoHugging Face12i-am-shaurya05 /robotrace-vla-robustness-traces RoboTrace Evidence Bundle This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines. The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run. What this bundle is for Use this repository to inspect evidence from RoboTrace: action-trace stability metrics visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.imageroboticsn<1K0 likes70 downloads3mo agoHugging Face13THULab /vlabench_primitive_ft_lerobot_video VLABench Primitive Tasks — LeRobot v3.0 (TsFile) Apache TsFile version of VLABench/vlabench_primitive_ft_lerobot_video. Overview This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 and the RLDS versions, this release stores the visual observations in a video-compressed format rather than as individual image files, giving better storage efficiency and data-loading… See the full description on the dataset page: https://huggingface.co/datasets/THULab/vlabench_primitive_ft_lerobot_video.tabularroboticsn<1K0 likes65 downloads1mo agoHugging Face14Vladimir0-1 /Air-Chat Format Each entry contains: question: the user’s question answer: the assistant’s response Example {"question": "Hello!", "answer": "Hello to you too!"} File architecture dialogue_dataset/ ├── README.md ├── dataset_infos.json ├── data/ │ ├──ru/ │ │ └── train-ru-00001-of-00001.jsonl │ └──en/ │ └── train-en-00001-of-00001.jsonl └── .gitattributes Load Russian version dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Vladimir0-1/Air-Chat.textn<1K2 likes56 downloads1mo agoHugging Face15Mayank022 /urban-vla-expert-v1 Urban VLA Expert v1 Urban VLA Expert v1 is a simulator dataset for language-conditioned urban driving. Each frame pairs a 256 x 256 front-camera image with ego state, a natural-language instruction, and continuous driving controls. This is a small research dataset, not evidence that a policy is ready for a real vehicle. The expert is a deterministic simulator controller, and the language prompts are curated paraphrases rather than speech collected from drivers. What… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/urban-vla-expert-v1.tabularn<1K1 likes50 downloads3mo agoHugging Face16UCSC-VLAA /VLM-CapCurriculum-TextReasoning-Data VLM-CapCurriculum-TextReasoning (D_text) Stage-2 textual-reasoning data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.texttext-generation10K<n<100K0 likes42 downloads4mo agoHugging Face17vladvlasov256 /opensubs-collocations OpenSubtitles Collocations NPMI-scored bigram collocations extracted from the OpenSubtitles parallel corpus. Three languages, three relation types, ~43K bigrams total. Languages & Corpus Size Language Code Corpus lines Bigrams English en ~100M 15,000 Dutch nl ~105M 15,000 Serbian sr ~50M 13,586 Relation Types ADJ+NOUN — adjective-noun pairs: "slim contract", "kreditan kartica" VERB+ADP — phrasal verbs / verb-preposition: "come on", "houden… See the full description on the dataset page: https://huggingface.co/datasets/vladvlasov256/opensubs-collocations.tabularfeature-extraction10K<n<100K1 likes34 downloads6mo agoHugging Face18UCSC-VLAA /ViLReward-73KProcess Reward Data for ViLBench: A Suite for Vision-Language Process Reward Modeling Paper | Project Page There are 73K vision-language process reward data sourcing from five training sets. imageimage-text-to-text10K<n<100K2 likes30 downloads1y agoHugging Face19VLAI-AIVN /DAM-QA-annotations DAM-QA Unified Annotations 22,675 question-answer pairs from 6 major VQA benchmarks, unified for the DAM-QA framework. This collection consolidates annotations from InfographicVQA, TextVQA, VQAv2, DocVQA, ChartQA, and ChartQA-Pro into standardized JSONL formats. 📖 Paper: Describe Anything Model for Visual Question Answering on Text-rich Images⚠️ Note: Images not included - obtain from original sources with proper licensing Repository Structure DAM-QA-annotations/… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/DAM-QA-annotations.textquestion-answering10K<n<100K0 likes26 downloads1y agoHugging Face203CTeam /action-evidence-vla-phase-state-cachetabularn<1K0 likes25 downloads10d agoHugging Face21VladHong /Alpha-Instruct Alpha-Instruct A synthetic instruction-tuning dataset for quantitative finance, covering formulaic alphas, technical indicators, and academic factor definitions. Designed to fine-tune language models on the vocabulary and reasoning patterns of quant researchers. Dataset Summary 336 rows of instruction–response pairs in chat format, generated from three distinct quant finance source corpora and post-processed to remove noise and near-duplicates. Each example is a messages… See the full description on the dataset page: https://huggingface.co/datasets/VladHong/Alpha-Instruct.texttext-generationn<1K0 likes22 downloads7mo agoHugging Face22vladtest88888888 /edupres Dataset Card for Edupres.ru Presentations Dataset Summary This dataset contains metadata about 44,210 presentations from the edupres.ru platform, with 21,941 presentations available in their original format. The dataset includes information such as presentation titles, descriptions, authors, publication dates, and file sizes. The presentations are primarily in Russian and cover various educational topics. Languages The dataset is multilingual, with Russian… See the full description on the dataset page: https://huggingface.co/datasets/vladtest88888888/edupres.texttext-classification10K<n<100K0 likes18 downloads7mo agoHugging Face23kyne0127 /vla-evaluation-v3imagen<1K0 likes14 downloads4mo agoHugging Face24bujangkiray /vlangtextn<1K1 likes13 downloads1y agoHugging Face25siulhin-vlad37 /chatbottextn<1K0 likes9 downloads3y agoHugging Face26VladdyAI123 /Vladdy_dataset_slimtext100K<n<1M0 likes8 downloads1y agoHugging Face27siulhin-vlad37 /finaltextn<1K0 likes7 downloads3y agoHugging Face28VladdyAI123 /Vladdy_Dataset_Refactoredtext100K<n<1M0 likes7 downloads4mo agoHugging Face29VladdyAI123 /Vladdy_Awarenesstext1K<n<10K0 likes7 downloads4mo agoHugging Face30VladShash /MNLP_M2_mcqa_datasettext10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.