datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
communication-adaptation-sft-100k
Communication Adaptation SFT (100K)
100,000 ShareGPT conversations demonstrating skilled communication style adaptation across 15 task types. Each example shows how to take the same underlying content and adjust register, technical depth, length, and framing for different audiences and purposes.
Motivation
Communication adaptation is a core professional skill that LLMs often handle clumsily. Common failures:
Technical monologue: explaining cloud storage to a… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/communication-adaptation-sft-100k.code-postes-communications-electroniques
Code des postes et des communications électroniques, non-instruct (2025-03-10)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-postes-communications-electroniques.alia_gva_communications
📘 ALIA_GVA_Communications Dataset
The ALIA_GVA_Communications dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_communications.gulf-coast-ga-atc-communications
Gulf Coast General Aviation ATC Communications Dataset
A conversational dataset of pilot–ATC radio exchanges designed for student pilots training in the Gulf Coast region (Texas, Louisiana, Mississippi, Alabama, Florida panhandle). Built for fine-tuning language models as study aids for aviation radio communications.
Dataset Overview
Metric
Value
Total conversations
5,401
Train split
4,860
Test split
541
Gulf Coast synthetic scenarios
3,400
Real ATC… See the full description on the dataset page: https://huggingface.co/datasets/starlineventures/gulf-coast-ga-atc-communications.fed-fomc-communications
Federal Reserve FOMC Statements and Minutes
Automatically scraped FOMC meeting statements and minutes from the U.S. Federal Reserve.
Dataset Structure
Field
Type
Description
Date
timestamp
FOMC meeting date
Release Date
timestamp
Publication date
Type
string
"Statement" or "Minute"
Text
string
Full text content
Source
Data scraped from Federal Reserve FOMC Calendar
License
CC0 1.0 Universal (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/nomnomshark41/fed-fomc-communications.Communication_Content_1
Communication-Content-1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Communication_Content_1.listening_for_communication_Corpus
Listening for Communication
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/listening_for_communication_Corpus.
