datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic-stsb
Arabic STSB Structure
The Arabic Version of the the Semantic Textual Similarity Benchmark (Cer et al., 2017)
it is a collection of sentence pairs drawn from news headlines, video and image captions, and natural language inference data.
Each pair is human-annotated with a similarity score from 1 to 5. However, for this variant, the similarity scores are normalized to between 0 and 1.
Examples:
{
"sentence1": "طائرة ستقلع",
"sentence2": "طائرة جوية ستقلع",
"score": 1.0
}
{… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-stsb.Arabic-NLi-Triplet
Arabic NLI Triplet
Dataset Summary
The Arabic Version of SNLI and MultiNLI datasets. (Triplet Subset)
Originally used for Natural Language Inference (NLI),
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Triplet Subset
Columns: "anchor", "positive", "negative"
Column types: str, str, str
Examples:
{
"anchor": "شخص على حصان يقفز فوق طائرة معطلة",
"positive": "شخص في الهواء الطلق، على حصان.",
"negative":… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Triplet.Arabic-NLi-Pair-Score
Arabic NLI Pair-Score
Dataset Summary
The Arabic Version of SNLI and MultiNLI datasets. (Pair-Score Subset)
Originally used for Natural Language Inference (NLI),
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair-Class Subset
Columns: "sentence1", "sentence2", "score"
Column types: str, str, float
Arabic Examples:
{
"sentence1": "شخص على حصان يقفز فوق طائرة معطلة",
"sentence2": "شخص يقوم… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Pair-Score.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.csa-clinical-stage-asset-intelligence-sample
CSA — Clinical-Stage Asset Intelligence · Free Sample
Clinical trials, FDA, and SEC — linked to the drug asset and the listed sponsor, with a
forward catalyst calendar. This is a free 150-row sample of the nearest-term
catalysts; the full snapshot carries 2,221 forward catalysts (955 linked to
124 listed sponsors) and 1,890 resolved assets.
Data, not investment advice. CSA is information, not a recommendation to buy, sell,
or hold any security. Estimated catalyst dates (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/csa-clinical-stage-asset-intelligence-sample.Arabic-NLi-Pair-Class
Arabic NLI Pair-Class
Dataset Summary
The Arabic Version of SNLI and MultiNLI datasets. (Pair-Class Subset)
Originally used for Natural Language Inference (NLI),
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair-Class Subset
Columns: "premise", "hypothesis", "label"
Column types: str, str, class with {"0": "entailment", "1": "neutral", "2": "contradiction"}
Arabic Examples:
{
"premise": "شخص… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Pair-Class.Spatial_Intelligence_UnderstandingArabic-NLi-Pair
Arabic-NLI-PAir
Dataset Summary
The Arabic Version of SNLI and MultiNLI datasets. (Pair Subset)
Originally used for Natural Language Inference (NLI),
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair Subset
Columns: "anchor", "positive"
Column types: str, str
Examples:
{
"anchor": "كيف أكون جيولوجياً جيداً؟",
"positive": "ماذا علي أن أفعل لأكون جيولوجياً عظيماً؟"
}
Disclaimer
Please note… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Pair.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.us-industrial-facility-intelligence-sample
US Industrial Facility Intelligence — Free Sample
This is a free 100-record sample. It is a subset of the full 1,464-record commercial dataset, provided so you can evaluate the data before deciding whether the full toolkit is useful to you.
An independent, unofficial dataset by NeuroLab Works. Not affiliated with, sponsored by, or endorsed by the U.S. EPA.
What this is
100 real, deduplicated US industrial facilities regulated under EPA's Toxics Release Inventory… See the full description on the dataset page: https://huggingface.co/datasets/NeuroLabWorks/us-industrial-facility-intelligence-sample.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Uzaib52/indian-tech-career-intelligence-2026.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Jidnesh298/indian-tech-career-intelligence-2026.cyber-threat-intelligence-custom-dataADeLe_battery_v1dot0
Dataset Card for ADeLe
Dataset Summary
ADeLe (Annotated-Demand-Levels) battery is a single, unified test set whose every item is labelled with the level (0-5+) it demands on 18 general ability dimensions (e.g. attention and scan, logical reasoning, various knowledge areas) plus an “unguessability” dimension. It is produced by applying the DeLeAn rubrics, via GPT-4o annotators, to AI benchmarks.
Version 1.0 contains 16 108 items drawn from 63 tasks spread across a diverse… See the full description on the dataset page: https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0.ds_benchmark_edicom_edicom_3Artificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.Arab3M-Triplets
Arab3M-Triplets
This dataset is designed for training and evaluating models using contrastive learning techniques, particularly in the context of natural language understanding. The dataset consists of triplets: an anchor sentence, a positive sentence, and a negative sentence. The goal is to encourage models to learn meaningful representations by distinguishing between semantically similar and dissimilar sentences.
Dataset Overview
Format: Parquet
Number of rows: 3.03… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arab3M-Triplets.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/ShaikFayaz042/indian-tech-career-intelligence-2026.Arabic-Cohere-include-base-44-mmlu-style
The Refined Arabic Cohere INCLUDE Base 44 Dataset as MMLU-Style
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark spanning 44 languages that evaluates multilingual LLMs in the actual linguistic environments where they are deployed. The original dataset contains 22,637 4-option multiple-choice questions (MCQs) extracted from academic and professional exams, covering 57 topics, including regional knowledge.
When we reviewed the Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-Cohere-include-base-44-mmlu-style.Arabic-Quora-Duplicates
Arabic-Quora-Duplicates
Dataset Summary
The Arabic Version of the Quora Question Pairs Dataset
It contains the Quora Question Pairs dataset in four formats that are easily used with Sentence Transformers to train embedding models.
The data was originally created by Quora for this Kaggle Competition.
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair Subset
Columns: "anchor", "positive"
Column types: str, str… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-Quora-Duplicates.Medical_Intelligence_Dataset_40k_Rows_of_Disease_Info_Treatments_and_Medical_QAMedical Intelligence Dataset: 40k+ Rows of Disease Info, Treatments, and Medical Q&ACreated by: Huzefa Nalkheda Wala
Unlock a valuable dataset containing 40,443 rows of detailed medical information. This dataset is ideal for patients, medical students, researchers, and AI developers. It offers a rich combination of disease information, symptoms, treatments, and curated medical Q&A for students, as well as dialogues between patients and doctors, making it highly versatile.
What's… See the full description on the dataset page: https://huggingface.co/datasets/huzaifa525/Medical_Intelligence_Dataset_40k_Rows_of_Disease_Info_Treatments_and_Medical_QA.Arabic-finanical-rag-embedding-dataset
Arabic Version of The Finanical Rag Embedding Dataset
This dataset is tailored for fine-tuning embedding models in Retrieval-Augmented Generation (RAG) setups. It consists of 7,000 question-context pairs translated into Arabic, sourced from NVIDIA's 2023 SEC Filing Report.
The dataset is designed to improve the performance of embedding models by providing positive samples for financial question-answering tasks in Arabic.
This dataset is the Arabic version of the original… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-finanical-rag-embedding-dataset.ai-failure-intelligence
AI Failure Intelligence Dataset
Structured, analyst-grade dataset of real-world AI failures.Built for researchers, red-teamers, and enterprise AI security teams.
Dataset Description
This is a FREE sample of 50 curated AI failure cases from the full
AI Failure Intelligence dataset (5,000+ cases, updated daily).
Each case is enriched with machine intelligence including:
Failure classification across 14 categories
Severity scoring (0-100)
Root cause analysis
Risk pattern… See the full description on the dataset page: https://huggingface.co/datasets/aifi-intelligence/ai-failure-intelligence.awesome_chatgpt_prompts_ar
📦 Awesome Arabic Chatgpt Prompts
📝 Overview
This repository contains a collection of Arabic prompts designed for use with AI language models (such as ChatGPT).
The goal is to provide a lightweight dataset that helps Arabic-speaking users quickly get started with generative AI.
🔗 Website / Demo
Check out the live demo site:omarnj-lab.github.io/awesome_chatgpt_prompts_ar
✨ Features
Entirely in Arabic 🕌
Suitable for educational and… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/awesome_chatgpt_prompts_ar.career-intelligence-benchmarks
Psychometric Career Intelligence Benchmarks
Benchmark dataset of 20 student psychometric assessment cases with individual scores for personality trait, cognitive ability, interest alignment, motivation clarity, strength discovery, and career readiness.
Built by Psychometric.fyi.
Dataset Description
This dataset contains benchmark data for an educational resource exploring the science of psychometric assessments and career decision making — helping students… See the full description on the dataset page: https://huggingface.co/datasets/psychometric-fyi/career-intelligence-benchmarks.serp-intelligence-benchmarks
SERP Intelligence Benchmarks
Benchmark dataset of 20 SERP intelligence cases with individual scores for SERP visibility, search intent, ranking pattern, competitor visibility, SERP feature, and content opportunity signals.
Built by SERPChecker.fyi.
Dataset Description
This dataset contains benchmark data for a research focused SERP intelligence framework helping SEO researchers, content strategists, and digital marketers analyze search results, understand ranking… See the full description on the dataset page: https://huggingface.co/datasets/serpchecker-fyi/serp-intelligence-benchmarks.alpaca-cleaned-annotatedreview-intelligence-benchmarks
Review Removal Intelligence Benchmarks
Benchmark dataset of 20 review management cases with individual scores for review risk, authenticity, policy compliance, issue detection, platform coverage, and workflow efficiency.
Built by ReviewRemoval.Services.
Dataset Description
This dataset contains benchmark data for an automated review management system helping businesses and reputation management teams monitor reviews, identify potential violations, and manage… See the full description on the dataset page: https://huggingface.co/datasets/review-removal-services/review-intelligence-benchmarks.alpaca-zulualgozee_ai-driven-global-market-intelligence-dataset
AI-Driven Global Market Intelligence Dataset
Global Financial Market Data for Risk, Trend, and Investment Analysis
Dataset Info
Source: Kaggle
Original Size: 9.04 MB
Kaggle Downloads: 166
Files: 1
Files
global_market_ai_dataset.csv
Mirrored from Kaggle
