datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sonic-o1
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
🎯 What is SONIC-O1?
The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts.
Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.omnimcp_agentops_vector_reranker_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentops_vector_reranker_teaser.HumaniBench
HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
**HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks:
✅ Open/closed VQA
🌍 Multilingual QA
📌 Visual grounding
💬 Empathetic captioning
🧠 Robustness, reasoning, and ethics
Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.omnimcp_semantic_vector_cache_resolver_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_semantic_vector_cache_resolver_teaser.arXiv-AI-papers-multi-vector
Overview
This is a dataset containing individual pages from the top-40 most cited AI papers on arXiv](https://arxiv.org/abs/2412.12121) from the period 2023-01-01 to 2024-09-30.
Only the first 10 pages from each paper is included.
The dataset includes an image of each page as well as a multi-vector embedding using vidore/colqwen2-v1.0.
autonomous-db-internals-vector-search-suite
⚡ Autonomous Database Internals, Vector Search Engines & Distributed Storage Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Database Kernel & Vector Retrieval LLMs
💼 Get Full 12,500-Row Enterprise Suite on Gumroad →
Full 10,000 SFT + 2,500 DPO Rows • 254.6 MB Pre-Indexed SQLite DB • RLVR/GRPO Sandboxed Testbed • Commercial License
⚡ Overview & Industry Problem
Deploying autonomous AI… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-db-internals-vector-search-suite.vectorstore-mental_health
Vectorstore Dataset: Mental Health
Overview
This dataset contains pre-computed vector embeddings for the mental health domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-mental_health.vectorstore-women_health
Vectorstore Dataset: Women Health
Overview
This dataset contains pre-computed vector embeddings for the women health domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-women_health.vectorstore-accounting
Vectorstore Dataset: Accounting
Overview
This dataset contains pre-computed vector embeddings for the accounting domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-accounting.vectorstore-legal_business
Vectorstore Dataset: Legal Business
Overview
This dataset contains pre-computed vector embeddings for the legal business domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-legal_business.vectorstore-academic_tutoring
Vectorstore Dataset: Academic Tutoring
Overview
This dataset contains pre-computed vector embeddings for the academic tutoring domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-academic_tutoring.ibm-hls-burn-vectorizedmy-recipe-chat-fine-tuning-data
Dataset Card for my-recipe-chat-fine-tuning-data
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data.vectorstore-economics
Vectorstore Dataset: Economics
Overview
This dataset contains pre-computed vector embeddings for the economics domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.
What… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-economics.wikipedia_movies_dump_vector
Dataset Card for Dataset Name
This data contains movie vectors (embedding of approx of 168k) (embeddings of movies), each vector has title, cast, director, year etc and plot.
All of this information taken from wikipedia dump.
Embedded using all-MiniLM-L6-v2.
Metadata dump file contains dataframe of movies which used for vector creation.
vectorstore-general_health
Vectorstore Dataset: General Health
Overview
This dataset contains pre-computed vector embeddings for the general health domain, ready for use in Retrieval-Augmented Generation (RAG) applications, semantic search, and knowledge base systems. The embeddings are generated from high-quality source documents using state-of-the-art sentence transformers, making it easy to build production-ready RAG applications without the computational overhead of embedding generation.… See the full description on the dataset page: https://huggingface.co/datasets/meetara-lab/vectorstore-general_health.Vector-dataset
