datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmGQA
mmGQA
Full GQA dataset in mm-eval format (id, media, messages), covering all 10 splits: train_balanced, val_balanced, test_balanced, testdev_balanced, train, val, test, testdev, challenge, submission. See metadata.json for prompt template and upstream field mapping.
gorilla-openfunctions-v1Salesforce-xlam-function-calling-60kflywire-fafb-connectome
FlyWire FAFB v783 Connectome — GNN-ready package
The complete proofread wiring diagram of an adult female Drosophila melanogaster brain — 139,255 neurons and their synaptic connections — repackaged as a ready-to-train graph dataset. Companion to the flywire-gnn Python package.
This is a dataset packaging of two public, no-auth sources:
File
Contents
Source
connections.parquet (474 MB)
15,091,983 unique directed neuron→neuron pairs at ≥1 synapse: pre, post, syn_count… See the full description on the dataset page: https://huggingface.co/datasets/SLOP011/flywire-fafb-connectome.fly-sud-simulation
FlyWire-informed odor-reward simulation: individual-behavior V4b
The full predeclared validation FAILED. This is synthetic simulation data,
not measured fly behavior or a quantitative reproduction of Kaun et al. (2011).
Detailed results ·
Code and protocols
Findings and limitations
256 independently seeded validation flies, four conditions (paired, unpaired,
untrained, retrieval-DAN-silenced), two delays (30 min, 24 h), 32 flies per cell:
8 reciprocal replicate… See the full description on the dataset page: https://huggingface.co/datasets/Histochemichael/fly-sud-simulation.AutoIF-instruct-61kecommerce_last_exam
E-Commerce Last Exam
A benchmark for evaluating LLM agents on 120 real-world travel planning and e-commerce tool-use tasks. Each task runs in an isolated Docker container with domain-specific CLI tools and SQLite databases. Agents must search, analyze, and produce structured recommendations.
Repository: alibaba-flyai/ecommerce_last_exam
Evaluation CLI: flyai-bench (pip install flyai-bench)
Leaderboard: FlyaiLab/ecommerce_last_exam_leaderboard
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/FlyaiLab/ecommerce_last_exam.AutoIF-instruct-61k-with-funcsiiif_snorkel_labelsOpenR1-Math-220k-pruned-keep-0.9-end-start-0.5-correctnessso101_pickplace_block_1_lerobot_v2.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 19499,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/fl-ymd/so101_pickplace_block_1_lerobot_v2.1.OpenOrcaOpenR1-Math-220k-pruned-head-random-perturbationOpenR1-Math-220k-pruned-think_midOpenR1-Math-220k-pruned-keep-0.75-end-start-0.5Policy-on-the-Fly-Benchmark
⚠️ Content Warning: This dataset contains harmful content for AI filter evaluation, including biased expressions, crime-related scenarios, and jailbreak attempts.
PoFBench: Policy-on-the-Fly Benchmark
Overview
PoFBench (Policy-on-the-Fly Benchmark) is a test-only benchmark designed to measure the performance of policy-based custom filters in LLM-powered systems.
Existing AI safety benchmarks evaluate against fixed risk taxonomies predefined by experts. However, in… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSDS-Research/Policy-on-the-Fly-Benchmark.flyte-slack-data
Dataset Card for "flyte-slack-data"
More Information needed
NousResearch-hermes-function-calling-v1OpenR1-Math-220k-pruned-midOpenR1-Math-220k-pruned-middle-random-perturbationgorilla-apibenchpython-codevec-flytech_python-codes-25kOpenR1-Math-220k-pruned-keep-0.5-end-start-0.5-acc-increamentalOpenR1-Math-220k-pruned-keep-0.75-end-start-1.0pi-llm
PI-LLM Bench: The Core Retrieval Challenge Behind MRCR
Update: Accepted to COLM 2026 (San Francisco).
Moonshot AI (Kimi) PI-LLM is being observed internally for agent state tracking and robustness to context interference
ICML 2025 Long-Context Foundation Models Workshop Accepted.
AAAI 2026 Worshop Oral: LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems)
A simple context interference evaluation.
Update: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/pi-llm.OpenR1-Math-220k-pruned-keep-0.2-end-start-0.5-accOpenR1-Math-220k-pruned-keep-0.5-end-start-0.5-add-aimeM-FLYT-input-scoresThis repository contains the input scores dataset used for training M-FLYT as described in the paper Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining. The scores are formatted as a parquet dataset, and can be used to reproduce our results or to improve them by adding more or better scoring methods.
For code to use these scores and more information visit our GitHub repository.
flybrain-nfl-polymarket
FlyBrain NFL x Polymarket dataset
NFL regular-season games 2023-2025 matched to resolved Polymarket moneyline markets, with kickoff-eve market odds.
pm_games.parquet: one row per game (schedule + scores + Polymarket market metadata + odds)
p_eve: market probability of outcome0 winning ~24h before kickoff (CLOB prices-history)
p_close: market probability at last trade before kickoff
outcome0 = first slug token's team (outcome0_is_home marks its role); slug order is alphabetical… See the full description on the dataset page: https://huggingface.co/datasets/sach0312/flybrain-nfl-polymarket.stingning-ultrachat
