datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
internship-warehouse
FlyRank Internship — Pseudonymized Warehouse Release (v20260703)
The open-ended, warehouse-shaped dataset (~81.8M rows; daily fact
78,835,655 rows) for advanced capstone work. Star schema with salted, namespaced,
fingerprinted hash keys. Built from warehouse v2 full history (frozen snapshot,
export date 2026-07-03): an unbalanced panel — per-client history depth differs;
see dim_clients.gsc_data_start / ga4_data_start.
Table
Rows
Grain
dim_clients
104
one row per… See the full description on the dataset page: https://huggingface.co/datasets/FlyRank/internship-warehouse.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508
Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.mmGQA
mmGQA
Full GQA dataset in mm-eval format (id, media, messages), covering all 10 splits: train_balanced, val_balanced, test_balanced, testdev_balanced, train, val, test, testdev, challenge, submission. See metadata.json for prompt template and upstream field mapping.
riddleIt contains 585 English riddles. The top 173 was adjusted by GPT4.
gorilla-openfunctions-v1Salesforce-xlam-function-calling-60kstackexchange_20260331
Stack Exchange data dump - community release
This data dump is sourced from the various sites in the Stack Exchange network of Q&A sites.
This dump contains data up to and including 2026-03-31.
This community release is an unofficial replacement for the now-killed official archive.org release. As a reminder, windowsphone.stackexchange.com is still excluded from this data dump, as it isn't possible to get the archives anymore due to the site shutting down a few years ago. See the… See the full description on the dataset page: https://huggingface.co/datasets/flymin/stackexchange_20260331.Ev-Flyingfly-sud-simulation
FlyWire-informed odor-reward simulation: individual-behavior V4b
The full predeclared validation FAILED. This is synthetic simulation data,
not measured fly behavior or a quantitative reproduction of Kaun et al. (2011).
Detailed results ·
Code and protocols
Findings and limitations
256 independently seeded validation flies, four conditions (paired, unpaired,
untrained, retrieval-DAN-silenced), two delays (30 min, 24 h), 32 flies per cell:
8 reciprocal replicate… See the full description on the dataset page: https://huggingface.co/datasets/Histochemichael/fly-sud-simulation.AutoIF-instruct-61kAnimeDL-2MHere is the repo for our paper: AnimeDL-2M: Million-Scale AI-Generated Anime Image Detection and Localization in Diffusion Era
Our paper is accepted by ACM MM 25 DFF Workshop (Oral)
Disclaimer
This dataset includes publicly available images on the internet and is intended solely for academic research and non-commercial use. All copyrights and intellectual property rights of the original images remain with their respective owners. We do not claim any ownership over the content contained within… See the full description on the dataset page: https://huggingface.co/datasets/FlyTweety/AnimeDL-2M.ecommerce_last_exam
E-Commerce Last Exam
A benchmark for evaluating LLM agents on 120 real-world travel planning and e-commerce tool-use tasks. Each task runs in an isolated Docker container with domain-specific CLI tools and SQLite databases. Agents must search, analyze, and produce structured recommendations.
Repository: alibaba-flyai/ecommerce_last_exam
Evaluation CLI: flyai-bench (pip install flyai-bench)
Leaderboard: FlyaiLab/ecommerce_last_exam_leaderboard
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/FlyaiLab/ecommerce_last_exam.Ev-FlyingAutoIF-instruct-61k-with-funcsEv-Flyinggrocery-flyers-rawFly_cylinders_teddy_rawiiif_snorkel_labelsinternship-starter
FlyRank Internship — Starter Dataset (Anonymized)
The public, safe starting point for the FlyRank Applied Search Intelligence ML internship.
30,000 anonymized content-performance rows across 32 pseudonymized clients (53 columns).
Public-safe: hashed content_id / client_id + numeric/categorical metrics only — no titles, URLs, keywords, domains, or client names.
What it's for
Week 1–2 quick wins and the ready-now capstone lanes (ranking-signal analysis, lifecycle /… See the full description on the dataset page: https://huggingface.co/datasets/FlyRank/internship-starter.flyte-slack-data
Dataset Card for "flyte-slack-data"
More Information needed
OpenR1-Math-220k-pruned-keep-0.9-end-start-0.5-correctnessOpenOrcallama-python-codes-30k
Python Codes - 30k examples, Llama1&2 tokenized dataset
Author
FlyTech
For general guide on how to create, quantize, merge or inference the model and more, visit:
hackmd.io/my_first_ai
Overview
This dataset serves as a rich resource for various Natural Language Processing tasks such as:
Question Answering
Text Generation
Text-to-Text Generation
It primarily focuses on instructional tasks in Python, tokenized specifically for the Llama architecture.… See the full description on the dataset page: https://huggingface.co/datasets/flytech/llama-python-codes-30k.OpenR1-Math-220k-pruned-think_midultrafeedback_clean
Dataset Card for UltraFeedback Cleaned
Dataset Description
This is a cleaned version of the HuggingFaceH4/ultrafeedback_binarized
and was turned into jsonl format for DPO or PPO training.
I did the following clean steps:
Remove all lines with 'translation' or 'translate'. I believe few translation tasks are not good for fine-tuning.
Remove all answers starts with 'User: As an AI assistan'. It's a mistake that assistant answers have prompt.
Remove all lines with 'As an… See the full description on the dataset page: https://huggingface.co/datasets/flyingfishinwater/ultrafeedback_clean.python-codevec-flytech_python-codes-25kgorilla-apibenchOpenR1-Math-220k-pruned-head-random-perturbationOpenR1-Math-220k-pruned-midNousResearch-hermes-function-calling-v1
