datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-07-29-msm-philosophy-spec-fabrication-probes
Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence?
experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes.MLLM-Fabric
🧵 MLLM-Fabric: Multimodal LLM-Driven Robotic Framework for Fabric Sorting and Selection
📄 Overview
This is the official repository for the paper:
MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
Accepted to IEEE Robotics and Automation Letters (RA-L)
🏫 About This Work
This work is from the Robot-Assisted Living LAboratory (RALLA) at the University of York, UK.
🧵 Fabric Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/EuniceF/MLLM-Fabric.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Fabric AI.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.discursos-senado-legislatura-56
Discursos da 56ª Legislatura do Senado Federal
Visão geral
Corpus de pronunciamentos do Plenário do Senado Federal relativos à 56ª Legislatura (2019–2023), coletados da API pública do Senado e consolidados em Parquet e CSV. Cada registro corresponde a um pronunciamento, com metadados e texto integral quando disponibilizado pela fonte.
Versão atual
Esta documentação corresponde à versão v1.1.1. Os arquivos de dados são os mesmos da versão v1.1.0; a… See the full description on the dataset page: https://huggingface.co/datasets/fabriciosantana/discursos-senado-legislatura-56.fabric
FABRIC: Financial AI Benchmark for Reliability in Indian Context
FABRIC is a benchmark for evaluating how reliably large language models provide financial advice for Indian markets across six languages.
Usage
from datasets import load_dataset
dataset = load_dataset("agenticclass/fabric")
# Access a question
q = dataset["train"][0]
print(q["question_en"]) # English version
print(q["question_hi"]) # Hindi version
print(q["answer"]) # Ground truth answer… See the full description on the dataset page: https://huggingface.co/datasets/agenticclass/fabric.tulu-3-sft-personas-instruction-following-es
Tulu 3 SFT Personas Instruction Following (Spanish Translation)
This dataset is a Spanish translation of allenai/tulu-3-sft-personas-instruction-following, a 30k-example supervised fine-tuning dataset designed to improve instruction following and constraint satisfaction in chat models.
The translation was created to make this style of instruction-following data more useful for Spanish-language model development while keeping the original task structure intact.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fabriciocarraro/tulu-3-sft-personas-instruction-following-es.FabricGen-CVPR26
Dataset of FabricGen: Microstructure-Aware Woven Fabric Generation (CVPR 26)
The dataset has two components:
Macro-scale texture dataset (img_gen): Dataset of microstructure-free fabric textures along with captions. The dataset is used to fine-tune an image diffusion model.
Weaving pattern dataset (pattern_gen): Dataset of weaving drafts along with tags (sourced from Handweaving.net). The dataset is used to fine-tune a large language model.
Detailed fine-tuning strategy and… See the full description on the dataset page: https://huggingface.co/datasets/llxtyin/FabricGen-CVPR26.vibeapps-chat-fabric
vibeapps-chat-fabric — agentic ChatML
1329 full agentic coding trajectories. A fake user persona drives a coding agent
to build a self-contained web app, critiquing functionality + aesthetics over 3 turns.
Browse the live apps: https://huggingface.co/spaces/AlexWortega/vibeapps-chat-fabric
coder/assistant: minimax/minimax-m3 (pi agent) - user-sim: qwen/qwen3.7-max
personas: dotoshny (nitpicky), mamochka (non-techie mom), startuper - 443 ideas x 3 personas x 3 turns… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/vibeapps-chat-fabric.
