datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HumanoidRobotSoccer
Fall Prediction Dataset for Humanoid Robots
Dataset Summary
This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.olivers-mtor-atlas
Oliver's mTOR Atlas
The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 414 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.turkish-olive-production
Turkish Olive Production Data
Olive and olive oil production in Türkiye at province, region and variety
level. Türkiye is the world's largest producer of table olives, yet its
production figures have not been available as a single machine-readable set
below the national level. This dataset collects them.
Published by zeytin.net ·
Source repository: github.com/yudumnet/zeytinnet
Configurations
Config
Rows
Contents
provinces
45
2024-25 season by province:… See the full description on the dataset page: https://huggingface.co/datasets/yudumnet/turkish-olive-production.PersonaMem🚨 We invite everyone to checkout our PersonaMem-v2 on 🤗HuggingFace, focusing on realistic and implicit user preferences in long conversations!
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized… See the full description on the dataset page: https://huggingface.co/datasets/OliverCMU/PersonaMem.grocery_sales_NOjob_listingscarpark_availabilitySingapore-fake-news-clarification-llama2reddit_user_o1SCMJOBS_KEThis dataset contains 10,000 synthetic entries representing job postings in supply chain management across industries such as procurement, logistics, and operations. It is designed for benchmarking NLP tasks (e.g., named entity recognition, salary prediction, skill extraction) and analyzing trends in job markets. Fields include job titles, companies, locations, salaries, required skills, and more.
Dataset Structure
Each entry includes the following fields:
Job Title (string): Role-specific… See the full description on the dataset page: https://huggingface.co/datasets/Olive254/SCMJOBS_KE.oliveHack_train_df. AIHub에서 수집한 형사 사건 판결문을 전처리 한 데이터
math6243-cook-county-tax-shift-data
