datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.ink_test01Music-info-datasets
Music Info Datasets
Music-info-datasets 是一个围绕歌曲元数据、歌词、评论语义、本地音频匹配、音频特征和 MERT 表征整理的音乐信息数据集。数据以网易云音乐 song_id 作为统一主键,适合用于个人音乐库分析、标签体系构建、音乐检索、推荐实验、歌词/评论语义分析和音乐信息检索原型。
本数据集不提供可播放音频文件。表中的 local_audio_path、file_path 等字段来自作者本地音乐库匹配结果,仅用于说明匹配关系和特征来源,通常不能在其他环境中直接访问。
Dataset Summary
当前包含两个子数据集:
子数据集
说明
主表歌曲数
标签表歌曲数
本地音频匹配数
音频特征数
MERT 索引数
ink_bai_liked
作者网易云喜欢音乐/个人曲库样本
1,488
1,488
870
870
870
middle_ages
主题歌单样本
219
219
12
12
12
额外文件:
类型
数量/说明
原始 JSON 快照… See the full description on the dataset page: https://huggingface.co/datasets/Ink-bai/Music-info-datasets.InkubaLM-benchmarking
Dataset Card for Evaluation run of lelapa/InkubaLM-0.4B
Dataset automatically created during the evaluation run of model lelapa/InkubaLM-0.4B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/InkubaLM-benchmarking.allura-org__MN-12b-RP-Ink-details
Dataset Card for Evaluation run of allura-org/MN-12b-RP-Ink
Dataset automatically created during the evaluation run of model allura-org/MN-12b-RP-Ink
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allura-org__MN-12b-RP-Ink-details.lrobot4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 4,
"total_frames": 6214,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ink-swpfy/lrobot4.allura-org__L3.1-8b-RP-Ink-details
Dataset Card for Evaluation run of allura-org/L3.1-8b-RP-Ink
Dataset automatically created during the evaluation run of model allura-org/L3.1-8b-RP-Ink
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allura-org__L3.1-8b-RP-Ink-details.eval_lrobot2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 5,
"total_frames": 8864,
"total_tasks": 1,
"total_videos": 5,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ink-swpfy/eval_lrobot2.Inkuba_isixhosa_dev_scored_v1InkBench-ocr-results
OCR Bench Results: InkBench-ocr
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
zai-org/GLM-OCR
0.9B
1706
1614–1858
29
6
5
72%
2
lightonai/LightOnOCR-2-1B
1B
1622
1535–1740
25
11
4
62%
3
deepseek-ai/DeepSeek-OCR
4B
1527
1428–1631
20
17
3
50%
4
FireRedTeam/FireRed-OCR
2.1B
1382
1268–1474
13
27
0
32%
5… See the full description on the dataset page: https://huggingface.co/datasets/NealCaren/InkBench-ocr-results.Inkuba_english_dev_scored_v1Triangle104__L3.1-8B-Dusky-Ink_v0.r1-details
Dataset Card for Evaluation run of Triangle104/L3.1-8B-Dusky-Ink_v0.r1
Dataset automatically created during the evaluation run of model Triangle104/L3.1-8B-Dusky-Ink_v0.r1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__L3.1-8B-Dusky-Ink_v0.r1-details.Inkuba_xhosa_dev_sliced_v1Inkuba_xhosa_train_sliced_v1Inkuba_isizulu_train_sliced_v1Inkuba_isizulu_dev_scored_v1Triangle104__L3.1-8B-Dusky-Ink-details
Dataset Card for Evaluation run of Triangle104/L3.1-8B-Dusky-Ink
Dataset automatically created during the evaluation run of model Triangle104/L3.1-8B-Dusky-Ink
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__L3.1-8B-Dusky-Ink-details.Inkuba_isizulu_dev_sliced_v1Inkuba_english_train_sliced_v1Inkuba_isizulu_train_scored_v1Inkuba_isixhosa_train_scored_v1Inkuba_english_train_scored_v1Inkuba_english_dev_sliced_v1
