datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixamo-Animations-Characters
Mixamo Animations and Characters
A complete snapshot of the Mixamo library: 2,317 motion clips and
114 rigged characters, exported as binary FBX (FBX 7.7 / fbx7_2019) with per-file metadata.
All animations share one uniform 65-joint mixamorig skeleton, so any clip can drive any
compatible character without remapping.
Use animation_motion/ and character_refined/. The full export contains 2,446 animation
files, but 129 are single-pose assets that carry no motion (Mixamo's *_Pose*… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Mixamo-Animations-Characters.UniML3D
UniML3D
UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.PHINCAbstract
Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.Truebones-ZOO-Annotations
Truebones ZOO Annotations
Text prompts, per-clip metadata, rest-pose renders and the exact build pipeline for
Truebones ZOO — 1,097 animal motion clips across 74 skeletons: mammals, birds,
reptiles, insects, marine and prehistoric creatures. 1.02 hours, 111,158 frames, uniformly
30 fps. Rigs range from 9 to 143 joints; clips from 0.3 to 18.5 seconds.
The motion files themselves are not in this repository. Truebones ZOO is a commercial
library by Truebones Motions Animation… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Truebones-ZOO-Annotations.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.linkedin-job-postingsDMS-Fold2-Benchmark-Dataset
DMS-Fold2 Benchmark Dataset
This repository contains the benchmark datasets used to evaluate DMS-Fold2. The files are organized into two benchmark sets:
CASP/CAMEO targets (casp_cameo_targets/)
Mega-scale targets (megascale_targets/)
Each target includes the sequence and experimental data required to reproduce the benchmark inputs.
Directory Structure
paper_benchmarks/
├── casp_cameo_targets/
│ ├── alignments.tar.gz
│ ├── fastas/
│ ├── pdbs/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/LindertLab/DMS-Fold2-Benchmark-Dataset.Follow-Line-Combine-Datasetlinux-man-pages-tldr-summarized
Dataset Card for linux-man-pages-tldr-summarized
Dataset Summary
This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages.
Supported Tasks
This dataset should be used to fine-tune language models for summarization tasks.
linkedin_job_listingsbilu-linial-extensive-spectral-failure-v4
Extensive Spectral Failure of Bilu–Linial Signings at Arbitrary Girth
Subtitle: Positive-Density Outliers, Exterior-Power Obstructions, and Finite Moment CertificatesAuthor: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific release: v4.0.0 · Date: 2026-09-16Repository: PureOne/bilu-linial-extensive-spectral-failure-v4Status: public expert-review research release; not peer reviewed or proof-assistant formalized.
This Hugging Face repository is an AI-friendly… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/bilu-linial-extensive-spectral-failure-v4.linkedin-company-profilelingnli-multi
Dataset Card for Dataset Name
Dataset Summary
This repository contains a collection of machine translations of LingNLI dataset
into 9 different languages (Bulgarian, Finnish, French, Greek, Italian, Korean, Lithuanian, Portuguese, Spanish). The goal is to predict textual entailment (does sentence A
imply/contradict/neither sentence B), which is a classification task (given two sentences,
predict one of three labels). It is here formatted in the same manner as the… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/lingnli-multi.paper_result
Local-MIP 2.0: Experimental Data
This repository contains the instance lists and key experimental results accompanying the Local-MIP 2.0 paper.
Datasets
MIPLIB-2017-instances.csv: Inclusion status, reference objective values, and exclusion reasons for all 1,065 instances in the MIPLIB 2017 Collection, with Easy/Hard/Open groups for the included instances. The instance_url column links to each instance's official MIPLIB webpage. See Section 7.1.… See the full description on the dataset page: https://huggingface.co/datasets/linpeng0105/paper_result.AirCa
Contents
1. About Dataset
2. Download
3. Description
3.1 AirCa-W
3.2 AirCa-N
3.3 Constraints description
4. The AirCa APIs
5. References
Dataset Download: https://huggingface.co/datasets/LINC-BIT/AirCaDataset Website: https://huggingface.co/datasets/LINC-BIT/AirCaCode Link: https://github.com/LINC-BIT/AirCaPaper Link:
1. About Dataset
AirCa is a publicly available aircraft cargo loading dataset with millions of instances from industry. It has three unique… See the full description on the dataset page: https://huggingface.co/datasets/LINC-BIT/AirCa.LinuxCommandslinguistic_calibrationThis Datasets repo contains training and evaluation datasets for the paper "Linguistic Calibration of Long-Form Generations".
Please refer to our GitHub repo at https://github.com/tatsu-lab/linguistic_calibration for more information, and check out our paper for our research findings: https://arxiv.org/abs/2404.00474
LinkedInJobPostings
📊 LinkedIn Job Posting Engagement Analysis
Which LinkedIn job posting characteristics predict candidate engagement (views) — and how well can engagement be predicted or classified using only posting-level features?
Personal motivation: As someone in entrepreneurship, understanding which job posting features attract candidates is directly relevant to future hiring decisions.
📹 Presentation Video
<video… See the full description on the dataset page: https://huggingface.co/datasets/MichaelYitzchak/LinkedInJobPostings.linalg-bench-llm
LinAlg-Bench: Where LLMs Stop Computing and Start Hallucinating
Ten frontier LLMs drop from near-perfect to near-zero on 5×5 eigenvalue problems. Complete computational collapse is dimension-gated: rare at 3×3, dominant at 4×4 and 5×5. Failures dissociate cleanly by task — eigenvalues fail by constraint-aware fabrication (invented eigenvalues that still match the matrix trace), determinants by sign-accumulation drift. Nearly a third of irrational-spectrum eigenvalue failures are… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/linalg-bench-llm.line-msg-fact-check-tw
Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies
The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community.
Attribution to the Community
This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers.
To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.HinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.lingvanex_test_references
LTR
LTR -- Lingvanex Test References for MT Evaluation from English into a total of 30 target languages for a big variety of cases.
TEST CASES
Parameter
Description
Length
Sentences from 1 to 100 words.
Domain
Medicine (12%), Automobile (11%), Finance (8%)
Tokenizer
Jupiter is 1.000.000 km far. Ask Mr. Johnson for training
Tags
I want to eat and swim
Capitalisation (Case)
HELLO my Dear frIEND
Different languages in one text (Up to 3 languages)
I see… See the full description on the dataset page: https://huggingface.co/datasets/lingvanex/lingvanex_test_references.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.link_predictionneuml-linkedin-202501
NeuML LinkedIn Company Posts
This dataset is 12 months of NeuML's LinkedIn Company Posts as of January 2025. It contains the post text along with engagement metrics.
It was created as follows:
Export the company posts from the analytics page, see this link for instructions.
Run the following code to create a dataset
import pandas as pd
from datasets import load_dataset
df = pd.read_excel("export_data.xls", sheet_name=1, header=1)
df = df.dropna(axis="columns")… See the full description on the dataset page: https://huggingface.co/datasets/NeuML/neuml-linkedin-202501.COMI-LINGUA
Dataset Details
COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks.
COMI-LINGUA provides annotations for several key NLP tasks:
Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units.
Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/COMI-LINGUA.maternal-health-pregnancy
Synthetic Maternal Health & Pregnancy Complications Dataset
The complete bundle — including the full dataset (35,000 rows), trained xgboost model (AUC-ROC: 0.990), fully-executed notebook, and full Paper— is available on Gumroad:
👉 Get the full bundle on Gumroad for $30
Abstract
This dataset provides 30,000 synthetic records (10,000 per scenario) of pregnant women attending antenatal care (ANC) in LMIC facility settings. Each record contains 16 clinically relevant… See the full description on the dataset page: https://huggingface.co/datasets/LinderStacy-1/maternal-health-pregnancy.linux_commands
