datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ffp
FoundationPose paired synthetic renders
Each scene row contains two synchronized views and per-object pose, mask ID,
bounding box, visibility, and occlusion annotations. Depth is stored as the
original float32 NPY bytes; RGB and uint32 instance masks are stored as PNG
bytes. Matrices use row-major flattened arrays and canonical column-vector
names (camera_from_world and world_from_object). Raw source transforms are
retained alongside them.
The assets tables contain stable source… See the full description on the dataset page: https://huggingface.co/datasets/Krispin/ffp.KrishokChat
KrishokChat Dataset
KrishokChat is a provenance-traceable, multi-task Bengali agricultural dataset for safety-critical chemical advisory and domain-specific natural language understanding. Every instance in the dataset retains citation-level provenance (publisher, document title, page range, section path) linked directly to official agricultural extension handbooks and research manuals issued by government and NGO agricultural institutions in Bangladesh.
Figure 1:… See the full description on the dataset page: https://huggingface.co/datasets/RaiyanKhaan/KrishokChat.finemath-4plus-tokenizedso101-pick-placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 1194,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kristaqp/so101-pick-place.parquet_for_narprojfama_french_datalm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private
Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b
Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private.so101_lift_figureThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 16916,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kristaqp/so101_lift_figure.racf-ivit-finalq-trust-datasets
Q-Trust Datasets
Measurement and corpus package for the Q-Trust post-quantum migration project
(humoge7502/q-trust on GitHub). Every directory carries a MANIFEST.json
with status = REAL or SYNTHETIC_DEMO — never publish demo data as
expert/enterprise results (see docs/TRUTH_AUDIT.md in the repo).
Contents
Path
What
Status
qtrust_ai/artifacts/real_datasets/code_corpus.json
12,462 real code files (6,636 crypto-labeled), 34 repos, seed-42 scanner labels… See the full description on the dataset page: https://huggingface.co/datasets/KRISHNAPURI/q-trust-datasets.fineweb-edu-1B
FineWeb-Edu 1B
Dataset Description
FineWeb-Edu 1B is a high-quality, stratified subset of the HuggingFaceFW/fineweb-edu dataset. It contains approximately 1 billion tokens of educational web text, carefully sampled to preserve the original distribution of source data (CommonCrawl dumps).
This dataset provides an accessible, lightweight alternative to the larger FineWeb-Edu subsets (like sample-10BT or sample-100BT) while maintaining the same data diversity and quality… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/fineweb-edu-1B.StudyChat
StudyChat (Public Mirror)
Public ungated mirror of wmcnicho/StudyChat.
parquet_storagepoh_cube_corner_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/kris0/poh_cube_corner_v1.Anime_UserRatingslm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-v3-private
Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b-v3
Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b-v3
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-v3-private.Unified_Dataset_with_Emotionseval_so101_ACTThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 3,
"total_frames": 1704,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kristaqp/eval_so101_ACT.pickplace-blue-krill-demoThis dataset was created using LeRobot for RoboSim at Home.
100 teleop episodes of a SO-ARM100 in MuJoCo. Task: pick up the blue krill oil bottle and place it inside the blue tray.
Companion policy: PruhaNLP/SmolVLA-pickplace-blue-krill-demo
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"total_episodes": 100,
"total_frames": 30636,
"total_tasks": 1,
"robot_type": "so100",
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/PruhaNLP/pickplace-blue-krill-demo.clust_pathway_scoresBundesliga_Stats_2018-2024This dataset contains football stats from the german Bundesliga scraped from https://fbref.com/en/
Used in my (Kristoffer Sommer Troelsen) mini-project for my 7th semester Medialogy machine learning course
poh_eth_lisbon_rocksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/kris0/poh_eth_lisbon_rocks.kristjan123This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/kris0/kristjan123.marker_test_20260920_213650This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/krish5831/marker_test_20260920_213650.dance_sameThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 1495,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/KristinWei/dance_same.edge-llm-bench
Edge LLM Bench — GGUF Quantization Benchmarks on Edge Devices
Controlled inference benchmark dataset for 7 GGUF K-quant quantization variants
(Q2_K through Q8_0) of Llama 3.2 3B Instruct and Qwen 2.5 1.5B Instruct across three hardware platforms:
Device
SoC / CPU
RAM
Backend
Google Pixel 6a
Google Tensor G1 (ARM Cortex-X1)
6 GB LPDDR5
llama.cpp CPU
Apple M4 Mac
Apple M4 (ARM, 10-core)
16 GB unified
llama.cpp Metal
HP Pavilion x86
Intel Core i5-1235U (12th gen)
16 GB… See the full description on the dataset page: https://huggingface.co/datasets/krisdcosta/edge-llm-bench.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.marker_test_20260920_201308This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/krish5831/marker_test_20260920_201308.so101_pickplace_wall_v1_20260722_174720This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/kris0/so101_pickplace_wall_v1_20260722_174720.
