datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguistic-similaritylibero_lingbot_va
LIBERO datasets pre-encoded for LingBot-VA
This repository mirrors three community-preprocessed LIBERO suites for easier migration and training with the official Robbyant/LingBot-VA pipeline.
Contents
libero_spatial/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents.
libero_goal/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents.
libero_object/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents.
Each suite contains… See the full description on the dataset page: https://huggingface.co/datasets/Wjjjh/libero_lingbot_va.humanego_serve_bread_lingbot_lerobot_with_latents
HumanEgo Serve Bread LingBot LeRobot With Latents
This dataset contains LeRobot-format robot demonstrations for the task:
pick up the bread and place it on the plate
The repository has two standalone LeRobot-style roots:
humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos.
humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos.
Each split includes:
data/: episode parquet files.
videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.coyo-700m
Dataset Card for COYO-700M
Dataset Summary
COYO-700M is a large-scale dataset that contains 747M image-text pairs as well as many other meta-attributes to increase the usability to train various models. Our dataset follows a similar strategy to previous vision-and-language datasets, collecting many informative pairs of alt-text and its associated image in HTML documents. We expect COYO to be used to train popular large-scale foundation models
complementary to other… See the full description on the dataset page: https://huggingface.co/datasets/lingao123/coyo-700m.IndicMMLU-Pro
IndicMMLU Dataset
This dataset contains the following languages:
punjabi
hindi
urdu
telugu
gujrati
kannada
tamil
marathi
bengali
UPLOAD
Cite our work.
This dataset is also described in IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding.
@dataset{kj2024indicmmlupro,
author = {Kj, Sankalp and Kumar, Ashutosh and Balaji, Laxmaan and Kotecha, Nikunj and Jain, Vinija and Chadha, Aman and Bhaduri, Sreyoshi},
title =… See the full description on the dataset page: https://huggingface.co/datasets/LinguaLift/IndicMMLU-Pro.tiktok-video-engagement-200k
TikTok Creator and Video Engagement (200K)
This release contains 209,543 TikTok videos from 1,872 creators with daily engagement and follower statistics, covering videos posted from 2024-06-24 to 2024-11-09.
The release contains derived video-content fields. Audio transcripts, screenshots, textual data, and music metadata were used to generate short video summaries; topic labels and emotion scores were then derived from those summaries using machine learning models. These… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-200k.linguabase
Linguabase
A large, sense-aware map of English word meaning — released free to the
public domain by IDEA.org, a nonprofit.
Linguabase is a set of flat tables describing how English words relate to
one another: which words are associated, which are opposed, which share a
root, what senses a word has, how to define or clue it, and how to filter
it for an audience. It is built for people who want rich semantic data they
can build on — game makers first, but also educators… See the full description on the dataset page: https://huggingface.co/datasets/linguabase/linguabase.tiktok-video-engagement-1m
TikTok Creator and Video Engagement (1M)
This release contains 1,035,817 TikTok videos from 4,926 creators with daily engagement and follower statistics, covering videos posted from 2024-06-09 to 2025-03-20.
Github: https://github.com/lingbowzd/tiktok-creator-video-trend-data
Cite this dataset: When does Trend-following Pay off? Evidence from Trending Content and Hashtag use
Uses
This dataset supports research on TikTok creator behavior, content strategy, trend… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-1m.lingcomp-qa-es
License & Attribution
MTEB-format derivative of somosnlp/LingComp_QA (Spanish computational-linguistics QA). Query = question; corpus = answer. Licensed under Apache-2.0 (same as source).
screenplay-features-linguistic
Screenplay Features - Linguistic Categories
This dataset reorganizes the features from screenplay-features into theoretically-motivated linguistic categories.
Dataset Structure
The dataset contains 837 features organized into 10 linguistic categories:
1. SURPRISAL (57 features)
Language model predictability features measuring cognitive processing difficulty.
bert_surprisal (15)
surprisal (5) - GPT-2 surprisal
gpt2_char_surprisal (6)
ngram_surprisal (5)… See the full description on the dataset page: https://huggingface.co/datasets/Ishaank18/screenplay-features-linguistic.cross-lingual-pitfalls
Cross-Lingual Pitfalls
Cross-Lingual Pitfalls is a fixed, failure-focused dataset from the ACL 2025 paper "Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models." It contains 6,713 bilingual English-to-target-language question pairs across 16 target languages. The paper's search-based multilingual LLM evaluation method uses beam search and LLM-based simulation to discover cases where a model answers correctly in English but fails… See the full description on the dataset page: https://huggingface.co/datasets/xzx34/cross-lingual-pitfalls.tiktok-trending-hashtags-music
TikTok Trending Hashtags and Music (2024 - 2025)
This release contains the top 100 daily trending hashtags and music records from TikTok Creative Center, covering the period from 2024-05-23 to 2025-07-09. It includes 13,399 unique hashtags and 11,157 unique songs.
The data comes from TikTok Creative Center:
https://ads.tiktok.com/business/creativecenter/inspiration/popular/hashtag/pc/en
The trending hashtag and music rankings are not personalized and are updated daily by TikTok.… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-trending-hashtags-music.IndicMMLU
IndicMMLU Dataset
This dataset contains the following languages:
bengali
gujarati
hindi
kannada
marathi
punjabi
tamil
telugu
urdu
HinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.ling-3.0-tiny-atlas
ling-3.0-tiny-atlas
browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2
browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2
DiscoveryBench ctxgraph-8b-clean Qwen3-8B ctxgraph fair config repeat 2/3; strict 0.0646, vista job 932507. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 157/239 answered, mean HMS 0.0983 over answered / 0.0646 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24).
Dataset Info
Rows: 157
Columns: 10
Columns
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2.ChiEngMixBench-Dataset
ChiEngMixBench v0.2.0
Paper: arXiv:2601.16217Code and frozen release: GitHubGitHub release: v0.2.0
ChiEngMixBench evaluates terminology-form choice in Chinese AI/CS discourse. It contains controlled Chinese-English minimal pairs, item-level model outputs, anonymized human ratings, and auditable analysis code.
This release deliberately separates two views:
Paired terminology choice: whether an open-weight model assigns higher length-normalized sequence likelihood to an English… See the full description on the dataset page: https://huggingface.co/datasets/AI-Ling00/ChiEngMixBench-Dataset.multi-lingual-greetings-test
Multi-Lingual-Greetings-Test
This dataset was generated using NeMo Data Designer, a comprehensive framework for creating high-quality synthetic datasets from scratch or using seed data.
Custom Description
This dataset is a test dataset for multi-lingual greetings.
About NeMo Data Designer
NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides:
Diverse data generation using… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/multi-lingual-greetings-test.grading-question-triage-datasetMedCalc-Bench-Verified
Updates
Updates to MedCalc-Bench Verified will be made on this page going forward.
Here is the github link for our repository: https://github.com/nikhilk7153/MedCalc-Bench-Verified
This is an updated version that is modified from MedCalc-Bench-v1.2.
While we have audited MedCalc-Bench Verified on mulitple occasions, should there by any corrections or enhancements, we will update with a new release and specify any changes.
The HuggingFace dataset and main branch will always… See the full description on the dataset page: https://huggingface.co/datasets/lingyi0101/MedCalc-Bench-Verified.browsecomp-ctxgraph-30b-rl-sft-v3-corpus
SFT-v3 training corpus (clean) — 138 trajectories
Best-of-pool trajectory per synth DiscoveryBench task, from 21 runs across 4 sources
(8b-fold 58 / 8b-ctxgraph 44 / 8b-react 30 / 30b-ctxgraph 6), all rendered under the
ctxgraph code_graph prompt. 62/200 synth tasks were EXCLUDED by a data audit
(gold hypothesis references constant/missing columns in the task's own data —
see audit_synth_tasks.py); corpus covers all 138 clean tasks.
Selection (threshold-light): binary gates… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-sft-v3-corpus.LingzuV2_pick_box
LingzuV2 Pick Box Dataset
A robot manipulation dataset for picking up objects, created using LeRobot v3.0 format.
Dataset Statistics
Property
Value
Total Episodes
7
Total Frames
5,628
FPS
30
Codebase Version
v3.0
Robot Type
single_arm_7dof
Arm Side
right
Data Features
Observations
2 Camera Streams (480×640×3 RGB video @ 30 FPS, AV1 codec):
observation.images.desk - Desk camera view
observation.images.wrist - Wrist… See the full description on the dataset page: https://huggingface.co/datasets/XXXXXXDU/LingzuV2_pick_box.browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3
browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3
DiscoveryBench fold-8b Qwen3-8B fold baseline repeat 3/3; strict 0.0728, vista job 932461. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 150/239 answered, mean HMS 0.1160 over answered / 0.0728 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24).
Dataset Info
Rows: 150
Columns: 10
Columns
Column
Type
Description
task_id
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3.browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2
browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2
DiscoveryBench ctxgraph-8b-synth synth repeat 2/3; strict 0.1250, vista job 932522. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 149/239 answered, mean HMS 0.1678 over answered / 0.1046 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24).
Dataset Info
Rows: 149
Columns: 10
Columns
Column
Type
Description
task_id
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2.noise_linguistic_smallbrowsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2
browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2
DiscoveryBench ctxgraph-8b-dpo DPO eval 2/2; strict 0.0762, answered 156, vista job 933235. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 156/239 answered, mean HMS 0.1168 over answered / 0.0762 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs 0.060-0.076 and on par with 30B fold 0.078-0.086); answer counts 172/177/178.
Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2.browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3
browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3
DiscoveryBench react-8b Qwen3-8B react baseline repeat 3/3; strict 0.0575, vista job 932458. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 132/239 answered, mean HMS 0.1041 over answered / 0.0575 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24).
Dataset Info
Rows: 132
Columns: 10
Columns
Column
Type
Description
task_id
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3.browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239
SFT-v3 ctxgraph-8B — DiscoveryBench real 239, 3 eval runs
Qwen3-8B + LoRA-SFT (v3 clean corpus, 138 cross-method trajectories, 2 epochs, r16, job vista:955512),
merged, evaluated 3x on the 239 real DiscoveryBench tasks. Judge: gpt-5-nano (Azure), HMS scoring.
run
vista job
answered
mean HMS (answered)
strict (no-answer=0)
run1
958275
157/239
0.1211
0.0796
run2
958276
163/239
0.0992
0.0676
run3
959769
151/239
0.1236
0.0781
Baselines (same config/judge):… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239.common_voice_tamil_english-labeled-Data-filtered-v4rollout_lingbot_test_20260702_141704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/maximellerbach/rollout_lingbot_test_20260702_141704.
