CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chendelong /linguistic-similaritytabularn<1K1 likes708 downloads2y agoHugging Face02Wjjjh /libero_lingbot_va LIBERO datasets pre-encoded for LingBot-VA This repository mirrors three community-preprocessed LIBERO suites for easier migration and training with the official Robbyant/LingBot-VA pipeline. Contents libero_spatial/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents. libero_goal/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents. libero_object/: 500 episodes, 10 tasks, LeRobot v2.1, dual-camera Wan2.2 latents. Each suite contains… See the full description on the dataset page: https://huggingface.co/datasets/Wjjjh/libero_lingbot_va.tabularrobotics100K<n<1M0 likes687 downloads2mo agoHugging Face03Coffeecoderss /humanego_serve_bread_lingbot_lerobot_with_latents HumanEgo Serve Bread LingBot LeRobot With Latents This dataset contains LeRobot-format robot demonstrations for the task: pick up the bread and place it on the plate The repository has two standalone LeRobot-style roots: humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos. humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos. Each split includes: data/: episode parquet files. videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.tabular10K<n<100K0 likes519 downloads3mo agoHugging Face04lingao123 /coyo-700m Dataset Card for COYO-700M Dataset Summary COYO-700M is a large-scale dataset that contains 747M image-text pairs as well as many other meta-attributes to increase the usability to train various models. Our dataset follows a similar strategy to previous vision-and-language datasets, collecting many informative pairs of alt-text and its associated image in HTML documents. We expect COYO to be used to train popular large-scale foundation models complementary to other… See the full description on the dataset page: https://huggingface.co/datasets/lingao123/coyo-700m.imagetext-to-image100M<n<1B0 likes446 downloads9mo agoHugging Face05LinguaLift /IndicMMLU-Pro IndicMMLU Dataset This dataset contains the following languages: punjabi hindi urdu telugu gujrati kannada tamil marathi bengali UPLOAD Cite our work. This dataset is also described in IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding. @dataset{kj2024indicmmlupro, author = {Kj, Sankalp and Kumar, Ashutosh and Balaji, Laxmaan and Kotecha, Nikunj and Jain, Vinija and Chadha, Aman and Bhaduri, Sreyoshi}, title =… See the full description on the dataset page: https://huggingface.co/datasets/LinguaLift/IndicMMLU-Pro.tabulartext-generation100K<n<1M4 likes444 downloads2y agoHugging Face06lingbow /tiktok-video-engagement-200k TikTok Creator and Video Engagement (200K) This release contains 209,543 TikTok videos from 1,872 creators with daily engagement and follower statistics, covering videos posted from 2024-06-24 to 2024-11-09. The release contains derived video-content fields. Audio transcripts, screenshots, textual data, and music metadata were used to generate short video summaries; topic labels and emotion scores were then derived from those summaries using machine learning models. These… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-200k.tabulartext-classification1M<n<10M11 likes375 downloads17d agoHugging Face07linguabase /linguabase Linguabase A large, sense-aware map of English word meaning — released free to the public domain by IDEA.org, a nonprofit. Linguabase is a set of flat tables describing how English words relate to one another: which words are associated, which are opposed, which share a root, what senses a word has, how to define or clue it, and how to filter it for an audience. It is built for people who want rich semantic data they can build on — game makers first, but also educators… See the full description on the dataset page: https://huggingface.co/datasets/linguabase/linguabase.tabular10M<n<100M2 likes337 downloads3mo agoHugging Face08lingbow /tiktok-video-engagement-1m TikTok Creator and Video Engagement (1M) This release contains 1,035,817 TikTok videos from 4,926 creators with daily engagement and follower statistics, covering videos posted from 2024-06-09 to 2025-03-20. Github: https://github.com/lingbowzd/tiktok-creator-video-trend-data Cite this dataset: When does Trend-following Pay off? Evidence from Trending Content and Hashtag use Uses This dataset supports research on TikTok creator behavior, content strategy, trend… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-1m.tabularfeature-extraction10M<n<100M4 likes263 downloads17d agoHugging Face09bowang0911 /lingcomp-qa-es License & Attribution MTEB-format derivative of somosnlp/LingComp_QA (Spanish computational-linguistics QA). Query = question; corpus = answer. Licensed under Apache-2.0 (same as source). tabulartext-retrieval1K<n<10K0 likes220 downloads3mo agoHugging Face10Ishaank18 /screenplay-features-linguistic Screenplay Features - Linguistic Categories This dataset reorganizes the features from screenplay-features into theoretically-motivated linguistic categories. Dataset Structure The dataset contains 837 features organized into 10 linguistic categories: 1. SURPRISAL (57 features) Language model predictability features measuring cognitive processing difficulty. bert_surprisal (15) surprisal (5) - GPT-2 surprisal gpt2_char_surprisal (6) ngram_surprisal (5)… See the full description on the dataset page: https://huggingface.co/datasets/Ishaank18/screenplay-features-linguistic.tabulartext-classification1M<n<10M3 likes146 downloads9mo agoHugging Face11xzx34 /cross-lingual-pitfalls Cross-Lingual Pitfalls Cross-Lingual Pitfalls is a fixed, failure-focused dataset from the ACL 2025 paper "Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models." It contains 6,713 bilingual English-to-target-language question pairs across 16 target languages. The paper's search-based multilingual LLM evaluation method uses beam search and LLM-based simulation to discover cases where a model answers correctly in English but fails… See the full description on the dataset page: https://huggingface.co/datasets/xzx34/cross-lingual-pitfalls.tabularquestion-answering1K<n<10K0 likes130 downloads23h agoHugging Face12lingbow /tiktok-trending-hashtags-music TikTok Trending Hashtags and Music (2024 - 2025) This release contains the top 100 daily trending hashtags and music records from TikTok Creative Center, covering the period from 2024-05-23 to 2025-07-09. It includes 13,399 unique hashtags and 11,157 unique songs. The data comes from TikTok Creative Center: https://ads.tiktok.com/business/creativecenter/inspiration/popular/hashtag/pc/en The trending hashtag and music rankings are not personalized and are updated daily by TikTok.… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-trending-hashtags-music.tabularfeature-extraction100K<n<1M3 likes120 downloads17d agoHugging Face13LinguaLift /IndicMMLU IndicMMLU Dataset This dataset contains the following languages: bengali gujarati hindi kannada marathi punjabi tamil telugu urdu tabular100K<n<1M0 likes115 downloads2y agoHugging Face14LingoIITGN /HinGEAbstract Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.tabulartranslation1K<n<10K1 likes99 downloads2y agoHugging Face15juiceb0xc0de /ling-3.0-tiny-atlas ling-3.0-tiny-atlas image1M<n<10M0 likes83 downloads25d agoHugging Face16lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 DiscoveryBench ctxgraph-8b-clean Qwen3-8B ctxgraph fair config repeat 2/3; strict 0.0646, vista job 932507. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 157/239 answered, mean HMS 0.0983 over answered / 0.0646 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 157 Columns: 10 Columns Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2.tabularn<1K0 likes55 downloads1mo agoHugging Face17AI-Ling00 /ChiEngMixBench-Dataset ChiEngMixBench v0.2.0 Paper: arXiv:2601.16217Code and frozen release: GitHubGitHub release: v0.2.0 ChiEngMixBench evaluates terminology-form choice in Chinese AI/CS discourse. It contains controlled Chinese-English minimal pairs, item-level model outputs, anonymized human ratings, and auditable analysis code. This release deliberately separates two views: Paired terminology choice: whether an open-weight model assigns higher length-normalized sequence likelihood to an English… See the full description on the dataset page: https://huggingface.co/datasets/AI-Ling00/ChiEngMixBench-Dataset.tabulartext-classification1K<n<10K0 likes54 downloads2mo agoHugging Face18Alga2025 /multi-lingual-greetings-test Multi-Lingual-Greetings-Test This dataset was generated using NeMo Data Designer, a comprehensive framework for creating high-quality synthetic datasets from scratch or using seed data. Custom Description This dataset is a test dataset for multi-lingual greetings. About NeMo Data Designer NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides: Diverse data generation using… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/multi-lingual-greetings-test.textn<1K0 likes53 downloads8mo agoHugging Face19lingah /grading-question-triage-datasettabular1K<n<10K0 likes49 downloads2mo agoHugging Face20lingyi0101 /MedCalc-Bench-Verified Updates Updates to MedCalc-Bench Verified will be made on this page going forward. Here is the github link for our repository: https://github.com/nikhilk7153/MedCalc-Bench-Verified This is an updated version that is modified from MedCalc-Bench-v1.2. While we have audited MedCalc-Bench Verified on mulitple occasions, should there by any corrections or enhancements, we will update with a new release and specify any changes. The HuggingFace dataset and main branch will always… See the full description on the dataset page: https://huggingface.co/datasets/lingyi0101/MedCalc-Bench-Verified.tabularquestion-answering10K<n<100K0 likes49 downloads27d agoHugging Face21lingchensanwen /browsecomp-ctxgraph-30b-rl-sft-v3-corpus SFT-v3 training corpus (clean) — 138 trajectories Best-of-pool trajectory per synth DiscoveryBench task, from 21 runs across 4 sources (8b-fold 58 / 8b-ctxgraph 44 / 8b-react 30 / 30b-ctxgraph 6), all rendered under the ctxgraph code_graph prompt. 62/200 synth tasks were EXCLUDED by a data audit (gold hypothesis references constant/missing columns in the task's own data — see audit_synth_tasks.py); corpus covers all 138 clean tasks. Selection (threshold-light): binary gates… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-sft-v3-corpus.tabularn<1K0 likes48 downloads21d agoHugging Face22XXXXXXDU /LingzuV2_pick_box LingzuV2 Pick Box Dataset A robot manipulation dataset for picking up objects, created using LeRobot v3.0 format. Dataset Statistics Property Value Total Episodes 7 Total Frames 5,628 FPS 30 Codebase Version v3.0 Robot Type single_arm_7dof Arm Side right Data Features Observations 2 Camera Streams (480×640×3 RGB video @ 30 FPS, AV1 codec): observation.images.desk - Desk camera view observation.images.wrist - Wrist… See the full description on the dataset page: https://huggingface.co/datasets/XXXXXXDU/LingzuV2_pick_box.tabularrobotics1K<n<10K0 likes46 downloads6mo agoHugging Face23lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3 browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3 DiscoveryBench fold-8b Qwen3-8B fold baseline repeat 3/3; strict 0.0728, vista job 932461. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 150/239 answered, mean HMS 0.1160 over answered / 0.0728 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 150 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3.tabularn<1K0 likes46 downloads1mo agoHugging Face24lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2 DiscoveryBench ctxgraph-8b-synth synth repeat 2/3; strict 0.1250, vista job 932522. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 149/239 answered, mean HMS 0.1678 over answered / 0.1046 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 149 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2.tabularn<1K0 likes45 downloads1mo agoHugging Face25sunilregmi /noise_linguistic_smalltabular10K<n<100K0 likes45 downloads20d agoHugging Face26lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2 DiscoveryBench ctxgraph-8b-dpo DPO eval 2/2; strict 0.0762, answered 156, vista job 933235. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 156/239 answered, mean HMS 0.1168 over answered / 0.0762 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs 0.060-0.076 and on par with 30B fold 0.078-0.086); answer counts 172/177/178. Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2.tabularn<1K0 likes44 downloads1mo agoHugging Face27lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3 browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3 DiscoveryBench react-8b Qwen3-8B react baseline repeat 3/3; strict 0.0575, vista job 932458. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 132/239 answered, mean HMS 0.1041 over answered / 0.0575 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 132 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3.tabularn<1K0 likes43 downloads1mo agoHugging Face28lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239 SFT-v3 ctxgraph-8B — DiscoveryBench real 239, 3 eval runs Qwen3-8B + LoRA-SFT (v3 clean corpus, 138 cross-method trajectories, 2 epochs, r16, job vista:955512), merged, evaluated 3x on the 239 real DiscoveryBench tasks. Judge: gpt-5-nano (Azure), HMS scoring. run vista job answered mean HMS (answered) strict (no-answer=0) run1 958275 157/239 0.1211 0.0796 run2 958276 163/239 0.0992 0.0676 run3 959769 151/239 0.1236 0.0781 Baselines (same config/judge):… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239.tabularn<1K0 likes43 downloads21d agoHugging Face29Lingalingeswaran /common_voice_tamil_english-labeled-Data-filtered-v4tabularaudio-classification1K<n<10K0 likes42 downloads2y agoHugging Face30maximellerbach /rollout_lingbot_test_20260702_141704This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/maximellerbach/rollout_lingbot_test_20260702_141704.tabularroboticsn<1K0 likes42 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.