datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.IndEgo
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
Vivek Chavan¹²*, Yasmina Imgrund²†, Tung Dao²†, Sanwantri Bai³†, Bosong Wang⁴†, Ze Lu⁵†, Oliver Heimann¹, Jörg Krüger¹²
¹Fraunhofer IPK, Berlin ²Technical University of Berlin ³University of Tübingen
⁴RWTH Aachen University ⁵Leibniz University Hannover
*Project Lead †Work done during student theses/projects at Fraunhofer IPK… See the full description on the dataset page: https://huggingface.co/datasets/FraunhoferIPK/IndEgo.Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.quantum-like-attention-framework-1.3b-untuned-validation
Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration.
Key Specifications & Architecture
Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture)
Parameters: 1.3B parameters total configuration (327M active parameter student subset)
Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.fractal_rawseamless-align-enA-frA.speaker-embedding.hubert-xlChatGPT-Jailbreak-PromptsMIC21
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18029/
MIC21
Original description
One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is constantly expanding, still there is a huge demand for more datasets in terms of variety of domains and object classes covered. The goal of the… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/MIC21.InternData-fractal20220817_dataseamless-align-enA-frA.speaker-embedding.xlsr-2bframes-benchmark
FRAMES: Factuality, Retrieval, And reasoning MEasurement Set
FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning.
Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941.
Dataset Overview
824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles
Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.France_Government_Conversationsseamless-align-enA-frA.speaker-embedding.w2vbert-600mego-all-framesfractalfragmented_test_streamsoc-rawq-multi-0921-controlthinking_fractal20220817_data_lerobot_output_qwen3vl2WikiMultihopQA
2WikiMultihopQA
This repository only repackages the original 2WikiMultihopQA data so that every example follows the field layout used by HotpotQA. The content of the underlying questions, answers and contexts is unaltered.
All intellectual credit for creating 2WikiMultihopQA belongs to the authors of the paper Constructing a Multi‑hop QA Dataset for Comprehensive Evaluation of Reasoning Steps (COLING 2020) and the accompanying code/data in their GitHub project… See the full description on the dataset page: https://huggingface.co/datasets/framolfese/2WikiMultihopQA.SACSoNFragFake
FragFake: VLM-Based Edited-Image Detection Dataset
This repository contains four groups of examples—Gemini-IG, GoT, MagicBrush, and UltraEdit—each with two difficulty levels: easy and hard. The YAML front matter above tells the HF Dataset Viewer to expose eight configurations in the “Configurations” dropdown. Once you select a configuration, you’ll see its single instruction split.
Sampling Policy for Edited Images
To prevent potential privacy or content leakage, only… See the full description on the dataset page: https://huggingface.co/datasets/Vincent-HKUSTGZ/FragFake.Unified_Agent_Framework
A Unified Framework for the Evaluation of LLM Agentic Capabilities
This repository contains the dataset (Benchmark, Toolkit, and Environment assets) for the paper A Unified Framework for the Evaluation of LLM Agentic Capabilities.
The official code and agent execution sandbox can be found on GitHub: whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities.
Dataset Description
The dataset integrates diverse agent benchmarks into a standardized… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.self-improve-fragilityTartanDrive_2basilretro-games-gameplay-frames-30k-512pGEO_satellite_maneuvers
Simulated GEO Satellite Maneuver Dataset
Dataset Summary
This dataset contains simulated GEO satellite trajectories for station-keeping scenarios.Each sample is exported as a pair of files:
State vectors: state_vectors/*_state_vectors.parquet
Metadata: metadata/*_metadata.parquet
Files share the same <timestamp>_<pid> prefix and belong together. The dataset is organized into two subfolders: state_vectors/ and metadata/.
Simulation Tool
Dataset generated… See the full description on the dataset page: https://huggingface.co/datasets/FraDra/GEO_satellite_maneuvers.fragbench
FragBench (Public Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Dataset Summary
FragBench is a benchmark for evaluating cross-session, fragmented attacks on
LLM agents that use tools via the Model Context Protocol (MCP). Each campaign
is decomposed into many small fragments distributed across sessions; a
defender must reconstruct the compositional intent. The public tier in this
repository… See the full description on the dataset page: https://huggingface.co/datasets/LidaSafety/fragbench.fractal20220817_data_rawLA_dataset_human_made_franka_eef_Lerobotv21_260511
