datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InfoSeek_emb_qwen3vle_2baim-activation-informed-merging
AIM: does activation-informed merging change what makes a merge work?
Headline
AIM does exactly what it claims, the targeting is what makes it work — and it changes
nothing about what predicts a good merge.
AIM is exactly what it says on the tin, and that is verifiable from public artefacts alone.
The published with-AIM checkpoints are recovered, to R² = 0.9992, as a closed-form
per-input-channel shrinkage of their baseline twins toward the base model, with ω̂ =… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/aim-activation-informed-merging.InformationCapacity
AI-Flow-Information Capacity
🏆 Leaderboard |
🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Information Capacity evaluates an LLM's efficiency based on text compression performance relative to computational complexity, leveraging the inherent correlation between compression and intelligence.
Larger models can predict the next token more accurately, leading to higher… See the full description on the dataset page: https://huggingface.co/datasets/TeleAI-AI-Flow/InformationCapacity.15000-Famous-People-Marriage-Divorce-Info
🎯 Accurate Marriage and Divorce Info for 15,000 Famous People 🌍
Includes information about marriage type, spouse, marriage/divorce dates, and outcomes.
High credibility data verified using reliable, paid API services.
Covers notable figures including Politicians, Authors, Singers, Actors, and more.
📜 About
This dataset was created to support research into validating principles of Vedic astrology.
Each record contains detailed information in a structured format… See the full description on the dataset page: https://huggingface.co/datasets/vedastro-org/15000-Famous-People-Marriage-Divorce-Info.d-info-2005-names
German Name Frequencies by State & District (D-Info 2005)
Regional frequency of surnames and forenames in Germany, from the D-Info
2005 telephone-directory CD-ROM (klickTel, data status 02.06.2005), at two
administrative levels aligned with census-2022 geography:
State = Bundesland — the 16 federal states.
District = Landkreis / kreisfreie Stadt — the 400 districts, keyed by
their 5-digit Kreisschlüssel (AGS).
For every name each table gives its number of 2005 telephone… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/d-info-2005-names.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.vn-provinces-informal-employment-rate
Vietnam provinces informal employment rate
Provincial and regional share of employed persons in informal employment (percent). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-informal-employment-rate.ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark
🛸 ZeroTwin-UAV-Synthetic
Multi-Agent Physics-Informed Degradation Benchmark for Autonomous UAV Swarms
═══════════════════════════════════════════════════════════════════════════════════════
P H I L A B • P E N E L O P E I N C . R E S E A R C H D I V I S I O N
═══════════════════════════════════════════════════════════════════════════════════════
🏛️ Provenance & Institutional Trademarks
This open-source benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark.distillation-llm-rawanimals-infoDense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.Shamela_Books_info
Shamela Books information
This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text
Dataset Features
The dataset includes the following cleaned and standardized features:
Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
africa-synth-employment-informal-sector-employment-africa-all
Africa Synth Employment Informal Sector Employment Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-employment-informal-sector-employment-africa-all.spa-bench-data-information
Spa-Bench data information
This repository exposes two reviewer-friendly tables through Hugging Face Data Studio:
Model Evaluation Data Information: 2,892 scored physical rollouts across five checkpoints.
Training Data Information: the 1,200-row prompt and compositional-design manifest.
Use the configuration selector in the viewer to switch between the two tables. Task names are normalized to Counting, Ordinal Position, Physical State, Referential Description, Relational… See the full description on the dataset page: https://huggingface.co/datasets/Spa-Bench/spa-bench-data-information.synthetic-confidential-information-injected-business-excerpts
Synthetic Confidential Information Injected Business Excerpts
This dataset aims to provide business report excerpts which contain relevant confidential/sensitive information.
This includes mentions of :
1. Internal Marketing Strategies.
2. Proprietary Product Composition.
3. License Internals.
4. Internal Sales Projections.
5. Confidential Patent Details.
6. others.
The dataset contains around 1k business excerpt - Reasons pairs. The Reason field contains the… See the full description on the dataset page: https://huggingface.co/datasets/Rohit-D/synthetic-confidential-information-injected-business-excerpts.Medical_Intelligence_Dataset_40k_Rows_of_Disease_Info_Treatments_and_Medical_QAMedical Intelligence Dataset: 40k+ Rows of Disease Info, Treatments, and Medical Q&ACreated by: Huzefa Nalkheda Wala
Unlock a valuable dataset containing 40,443 rows of detailed medical information. This dataset is ideal for patients, medical students, researchers, and AI developers. It offers a rich combination of disease information, symptoms, treatments, and curated medical Q&A for students, as well as dialogues between patients and doctors, making it highly versatile.
What's… See the full description on the dataset page: https://huggingface.co/datasets/huzaifa525/Medical_Intelligence_Dataset_40k_Rows_of_Disease_Info_Treatments_and_Medical_QA.africa-synth-urbanization-urban-land-informal-development-all
Africa Synth Urbanization Urban Land Informal Development All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-urbanization-urban-land-informal-development-all.qna-quantum-information
Q&A Quantum Information Dataset
This dataset is created by digesting 500 different papers from the quantum information directory on arXiv,
the papers are located based on their relevance to the quantum information keyword.
Data retrieval
The data is extracted from the .pdf files using PyMuPDF package proxied from langchain.
Then the Q&A pair is generated by:
Generate N questions per page of the .pdf document based on its content.
We will feed each question to an LLM… See the full description on the dataset page: https://huggingface.co/datasets/CoAILab/qna-quantum-information.InfoVQA_Question_Completion_valHindi-Non-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains.
The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.Visual_information_retrieval
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.InfoVQA_VALinformal-art-9ccaba
informal-art-9ccaba
Synthetic weather test data: 31 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Silver-Eli/informal-art-9ccaba.informatics_kazclinical-handover-information-continuity-coherence-risk-v0.1What this repo is for
Detect when
handover information
and
actual followthrough
lose continuity
before
missed care
and avoidable deterioration.
informal-bug-44e49c
informal-bug-44e49c
Synthetic products test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jerry317/informal-bug-44e49c.InfoVQATestinfosec_harmful_behaviors
Infosec Harmful Behaviors
Offensive-security instruction prompts for refusal-direction research and abliteration of code/security models.
Dataset Details
This dataset contains infosec-domain harmful prompts intended to elicit refusal behavior from aligned instruction models. It is designed as the harmful side of a harmful/harmless contrast pair, analogous to mlabonne/harmful_behaviors but focused on offensive-security and malicious-coding requests.
Rows:
train:… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/infosec_harmful_behaviors.
