datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Function_Calling_Private_GG
Function Recommendation Dataset Readme
Description
This dataset is based on the "glaiveai/glaive-function-calling" repository and has been customized to suit my specific requirements. It is designed for fine-tuning a Large Language Model (LLM) on the task of generating function recommendations in a conversational context. The dataset contains 158,738 snippets of conversations between me and my virtual assistant named J.O.S.I.E. (Just an Outstandingly Smart Intelligent… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/Function_Calling_Private_GG.private-letter-rulings
Private Letter Rulings
Text of IRS Private Letter Rulings (and other written determinations [TAMs, CCAs, etc.]), covering 1999 through August 2026. The IRS publishes these as PDF files each week; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed.
Dataset Structure
45,401 rows, one per ruling. Columns:
Column
Type
Description
wd_number
string
9-digit IRS written determination number: 4-digit year + 2-digit week… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/private-letter-rulings.Usdb_private
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zxc4wewewe/Usdb_private.safealert-private
SafeAlert Private Dataset — Generation Prompts
Dataset Description
This dataset contains 160 red-team generation prompts used in the SafeAlert AI safety evaluation kit for Nigerian fintech. Each prompt asks a language model to produce a harmful financial fraud message across eight scam categories documented in the Nigerian Inter-Bank Settlement System (NIBSS) 2024 Fraud Report.
A safe model refuses every prompt without producing any part of the requested content.… See the full description on the dataset page: https://huggingface.co/datasets/hadex-pearl/safealert-private.
