datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
malaysian_journal_of_analytical_sciencefew-nerdFineWeb-Edu-Analytic
FineWeb-Edu-Analytic (v1)
FineWeb-Edu-Analytic (v1) is an English-language dataset containing 9908 documents, intended as a resource for training language models.
The dataset was generated by taking text sequences from the FineWeb-Edu dataset (CC-MAIN-2025-26 subset) to serve as a source. Each source sequence was then processed by a 48-billion parameter language model to generate a corresponding structured, analytical document.
Disclaimer: This dataset is not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/MultivexAI/FineWeb-Edu-Analytic.verified-analytics-tasks
Verified Analytics Tasks
150+ mainly small analytics and data-engineering tasks. The point of the set is
the answer key: every task ships its own automated checker, and every gold
answer was run through that checker and scored a clean 1.0 before the task was
allowed in. So the labels are more like "here's the checker, score it
yourself" instead of "just trust me bro."
I wanted to create a synthetic dataset inspired by this paper:
Autodata: An agentic data scientist to create… See the full description on the dataset page: https://huggingface.co/datasets/Eve39570/verified-analytics-tasks.mit-movie-triviamit-restaurantpro-onboarding-analyticsconll2003universal-nerwnut2017bionlp2004ontonotes5finenglish-introduction-to-data-analytics-30tweebankpidgin0.2ai-sports-analytics-2026
