datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled.
guru-RL-92k-extra-info-compressed
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Note for this extra-info-compressed data version!
The dataset provided in this repository is specifically intended for use with the latest release of VeRL (v0.4.0). Since VeRL rl_dataset.py processes datasets as datasets.Dataset, it is essential that the structure of all Parquet files remains fully consistent. This repository is designed to meet that requirement.
In this repo, the… See the full description on the dataset page: https://huggingface.co/datasets/IFM/guru-RL-92k-extra-info-compressed.infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled.
vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.InfoSeek_emb_qwen3vle_2bFZI-AURA
FZI-AURA
A multimodal autonomous-driving dataset featuring the largest LiDAR sensor
suite of any public autonomous-driving dataset.
Download |
Python SDK |
Data format |
License
2,473 scenes |
13.70 hours |
8 cameras |
up to 12 LiDARs |
4.11 million 3D-box annotations |
30.01 billion segmented LiDAR points
Public preview release: 1,081 out of the 2,473 FZI-AURA scenes are currently available. The remaining scenes will be added soon.
Pin a… See the full description on the dataset page: https://huggingface.co/datasets/fzi-forschungszentrum-informatik/FZI-AURA.eu-hydro-master-skeleton
EU-Hydro Master Skeleton
Per-basin GeoParquet shards derived from the Copernicus EU-Hydro v1.3 GeoPackages. Four layers are published — river centerlines, river-surface polygons, inland-water polygons (lakes + wide waters), and river-basin polygons — all reprojected to a common CRS and stripped of admin-only columns for easier querying.
Contents
eu_hydro_master_skeleton_geoparquet/
├── river_lines/ # River_Net_l MultiLineString ~1.3 M features
├──… See the full description on the dataset page: https://huggingface.co/datasets/InfoVis-Project-Group-19/eu-hydro-master-skeleton.aim-activation-informed-merging
AIM: does activation-informed merging change what makes a merge work?
Headline
AIM does exactly what it claims, the targeting is what makes it work — and it changes
nothing about what predicts a good merge.
AIM is exactly what it says on the tin, and that is verifiable from public artefacts alone.
The published with-AIM checkpoints are recovered, to R² = 0.9992, as a closed-form
per-input-channel shrinkage of their baseline twins toward the base model, with ω̂ =… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/aim-activation-informed-merging.tiktok-videos-users-info
TikTok Data, post + poster (user) info, ~1 Million
Dataset name: EinzzCookie/tiktok-videos-users-info
This dataset contains a large collection of TikTok video records paired with detailed creator/user information, stored in a single Parquet file (tiktok_video_user_data.parquet, ~3.87 GB). It is derived from TikTok’s internal video (“aweme”) data model and includes both post-level metadata/engagement stats and nested author profile data.
Source
Collected by… See the full description on the dataset page: https://huggingface.co/datasets/EinzzCookie/tiktok-videos-users-info.infovqa_beirThis is a copy of https://huggingface.co/datasets/jinaai/infovqa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/infovqa_beir.kbot_cappuccino_countThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kbot_right_arm_follower",
"total_episodes": 224,
"total_frames": 166988,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:224"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/infoslack/kbot_cappuccino_count.arabic_infographicsvqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar_beir.InformationCapacity
AI-Flow-Information Capacity
🏆 Leaderboard |
🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Information Capacity evaluates an LLM's efficiency based on text compression performance relative to computational complexity, leveraging the inherent correlation between compression and intelligence.
Larger models can predict the next token more accurately, leading to higher… See the full description on the dataset page: https://huggingface.co/datasets/TeleAI-AI-Flow/InformationCapacity.SO-101-ACT-bigThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 80,
"total_frames": 46835,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/infoslack/SO-101-ACT-big.d-info-2005-names
German Name Frequencies by State & District (D-Info 2005)
Regional frequency of surnames and forenames in Germany, from the D-Info
2005 telephone-directory CD-ROM (klickTel, data status 02.06.2005), at two
administrative levels aligned with census-2022 geography:
State = Bundesland — the 16 federal states.
District = Landkreis / kreisfreie Stadt — the 400 districts, keyed by
their 5-digit Kreisschlüssel (AGS).
For every name each table gives its number of 2005 telephone… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/d-info-2005-names.simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing.
SO-101-SMOLVLA-BIGThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 250,
"total_frames": 195461,
"total_tasks": 7,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:250"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/infoslack/SO-101-SMOLVLA-BIG.IndustryCorpus2_other_information_services_information_security
IndustryCorpus2: Information Services
This repository contains the IndustryCorpus2: Information Services domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_other_information_services_information_security.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
An_Infoton_Dose_of_RxRx
An Infoton Dose of RxRx — 735 Human Proteins
Version: 1.0.0
Released: 2026-09-06
Author: Walker, January Natatia — Infoton
DOI: 10.5281/zenodo.18210355
ORCID: 0009-0000-6843-2051
Overview
The Infoton Identity Physics Engine computes in seconds per protein and aligns with what cellular imaging measures. The coordinates in the An Infoton Dose of RxRx dataset were derived before correlation was run with timestamp proof.
An Infoton Dose of RxRx provides Infoton… See the full description on the dataset page: https://huggingface.co/datasets/Infoton/An_Infoton_Dose_of_RxRx.Institutional-Information-of-Bangladesh
Institutional-Information-of-Bangladesh Dataset
This Dataset contains all verified and authorized Institutional information in Bangladesh
Description
I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
http://data.gov.bd/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.vn-provinces-informal-employment-rate
Vietnam provinces informal employment rate
Provincial and regional share of employed persons in informal employment (percent). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-informal-employment-rate.details_informatiker__Qwen2-7B-Instruct-abliterated
Dataset Card for Evaluation run of informatiker/Qwen2-7B-Instruct-abliterated
Dataset automatically created during the evaluation run of model informatiker/Qwen2-7B-Instruct-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_informatiker__Qwen2-7B-Instruct-abliterated.SO-101-SMOLVLA-500This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 202,
"total_frames": 157501,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:202"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/infoslack/SO-101-SMOLVLA-500.ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark
🛸 ZeroTwin-UAV-Synthetic
Multi-Agent Physics-Informed Degradation Benchmark for Autonomous UAV Swarms
═══════════════════════════════════════════════════════════════════════════════════════
P H I L A B • P E N E L O P E I N C . R E S E A R C H D I V I S I O N
═══════════════════════════════════════════════════════════════════════════════════════
🏛️ Provenance & Institutional Trademarks
This open-source benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/ZeroTwin-UAV-Synthetic_Physics-Informed-Multi-UAV-Fault-Telemetry-Benchmark.quantum-information-and-complexity-theory
Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage
A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.animals-infoicra_infonova_randomview_sort_20260919_181016This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/saipuneethgottam/icra_infonova_randomview_sort_20260919_181016.icra_infonova_randomview_sort_20260919_183136This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/saipuneethgottam/icra_infonova_randomview_sort_20260919_183136.
