datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-layoffs-by-metro-area-msa-warn-act
US layoffs by metro area: 54,166 WARN notices mapped to 765 metro and micro areas
Rebuilt 2026-09-22. 765 of the 935 US core-based statistical areas carry at least one
layoff notice on record — 361 metropolitan and 404 micropolitan.
Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the
Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish
the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
QnA_Descriptiveedisum_dataset
Dataset Card for Edisum
Dataset Description
For more details:
Github repository: https://github.com/epfl-dlab/edisum
Paper: https://arxiv.org/pdf/2404.03428.pdf
Languages
Edisum only contains Wikipedia data collected from English Wikipedia. Consequently, synthetic data is also only generated in English.
Dataset Structure
The Edisum meta-dataset actually comprises 5 datasets:
wikiepdia_processed_data (Filtered existing Wikipedia data)… See the full description on the dataset page: https://huggingface.co/datasets/msakota/edisum_dataset.handwritten_multihop_reasoning_data
Dataset used to better understand how to:
Correct Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
This is a handwritten dataset created to aid in better understanding the multi-hop reasoning capabilities of LLMs.
To learn how the dataset was constructed please check out the project page, paper, and demo linked below.
This is the link to the Project Page.
This repo contains the code that was used to conduct the experiments in this paper.… See the full description on the dataset page: https://huggingface.co/datasets/msakarvadia/handwritten_multihop_reasoning_data.Arabic_dialects_to_MSAawesome-chatgpt-prompts
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/msantiiisocial/awesome-chatgpt-prompts.msa-darja-pairs-completeSiamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc)
and text scraped from books, news articles, reviews etc
license: apache-2.0
owls_trait_bias_SFTff-challenge-res
Tested Model Information
Model Name: SmolVLM-Base (2B Parameters)
Model Link: https://huggingface.co/HuggingFaceTB/SmolVLM-Base
Model Type: Multimodal Vision-Language Base Model
Loading Methodology & Python Code
I loaded the model using a Google Colab T4 GPU. To accommodate the 2B parameters within a 16GB VRAM limit,
the model was loaded in half precision like torch.float16 and mapped to the GPU using device_map="auto".
Inference was conducted using greedy decoding… See the full description on the dataset page: https://huggingface.co/datasets/msaleem-aisci/ff-challenge-res.25_blind_spotsarabic-msa-sample
Arabic — Modern Standard Arabic (MSA) Sample
Native-written, human-verified Modern Standard Arabic. No scraping. No machine
translation. No synthetic generation. Every sentence written from scratch by a
first-language speaker in formal news / official-statement register, then reviewed
line by line against a written checklist and measured for structural diversity across
the whole set.
A public demonstration sample (50 items). Larger MSA datasets and other varieties
(Levantine… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-msa-sample.arabic-dialect-to-msaforc
Dataset Card for Forc
Dataset Description
For more details:
Github repository: https://github.com/epfl-dlab/forc
Paper: https://arxiv.org/pdf/2308.06077
The dataset was built from the data released by HELM (https://arxiv.org/pdf/2211.09110)
Citation Information
@inproceedings{šakota2024flyswat,
title={Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling},
author={Marija Šakota and Maxime Peyrard and Robert West}… See the full description on the dataset page: https://huggingface.co/datasets/msakota/forc.poetrymsa-darja-pairsQA_finetuning_simple_KIfill_in_the_blanks_simple_KIKI_simple_QA_dataset_for_finetuning_LLMsKI_simple_cloze_for_finetuning_LLMsbenstokes_love_haiasapbook-review
