annon-neurips-2026/beyond-n-grams
Beyond N-Grams (BNG) Dataset Summary This dataset is designed to study language trends over time. It combines multi-source metadata, aggregated review signals, and LLM-generated features to enable hypothesis-driven and exploratory research on how narrative and reception characteristics evolve over time. The dataset does not contain original book text. Instead, it uses LLM-generated proxy content and derived features to approximate semantic and evaluative… See the full description on the dataset page: https://huggingface.co/datasets/annon-neurips-2026/beyond-n-grams.
Beyond N-Grams (BNG)
Dataset Summary
This dataset is designed to study language trends over time. It combines multi-source metadata, aggregated review signals, and LLM-generated features to enable hypothesis-driven and exploratory research on how narrative and reception characteristics evolve over time.
The dataset does not contain original book text. Instead, it uses LLM-generated proxy content and derived features to approximate semantic and evaluative properties.
Dataset Description
Construction Process
The dataset is built using a three-stage pipeline:
- Data Collection
- Book metadata and summaries collected from:
- Google Books
- Wikipedia
- Goodreads
- Project Gutenberg
- Includes publication details, author attributes, ratings, and review statistics
- Proxy Content Generation
- LLM-generated content used to approximate book semantics
- Based on research-backed prompting methods
- Feature Engineering
- ~100 structured features per book, including:
- Metadata features
- Review-based signals
- LLM-derived evaluative scores
Dataset Structure
- Each row represents a single book
- Includes:
- Book metadata
- Author metadata
- Publication year
- Aggregated review statistics
- Derived feature set
- Some fields may contain nested structures (e.g., reviews)
Data Fields
The dataset consists of structured features grouped into three main categories: metadata, review-based signals, and LLM-generated evaluative attributes.
Metadata Features
book_title: Book title/namebook_author: Book author namegutenberg_book_link: Publically accessible book linkpreview_text: Publically available book textreconstructed_text: Proxy text of the book generated by the modelauthor_degree_level: Highest known academic qualification of the authorauthor_gender: Gender of the authorauthor_nationality: Nationality of the authorpub_year: Year of publicationgenre_count: Number of genres associated with the bookgenres: List of genres
Review and Engagement Features
review_count: Number of reviews includedreviews: Nested list containing:reviewer_review_count: Total reviews by reviewer (if available)reviewer_follower_count: Follower count of reviewerrating_out_of_5: Rating given in the reviewreview_year: Year of the reviewlikes: Number of likes on the reviewreview_length: Length of the review
total_reviews: Total number of reviewstotal_ratings: Total number of ratingsaverage_rating: Average rating scoreearly_avg_rating: Early-stage average rating (if available)star_1_counttostar_5_count: Distribution of ratings across 1–5 stars
LLM-Generated Features
The dataset includes ~100 LLM-generated features grouped into the following categories:
1. Sentiment and Polarity
Includes features measuring overall sentiment, intensity, shifts, and implicit emotional signals. Examples: polarity_score, sentiment_intensity, aspect_based_sentiment, sentiment_shift, implicit_sentiment, mixed_sentiment_ratio
2. Emotion and Affect
Captures emotional states based on Ekman categories and affective dynamics. Examples: ekman_anger_intensity, ekman_joy_intensity, ekman_sadness_intensity, ekman_fear_intensity, ekman_surprise_intensity, affective_arousal, emotional_variance
3. Sarcasm, Irony, and Figurative Language
Measures figurative expression, ambiguity, and creative language use. Examples: sarcasm_probability, situational_irony, hyperbole_intensity, linguistic_creativity, idiomatic_density, interpretation_flexibility, ambiguity_score
4. Political Ideology and Populism
Represents ideological stance and group dynamics. Examples: left_right_leaning, authoritarian_libertarian, populist_rhetoric_intensity, anti_elitism_score, in_group_favoritism, out_group_derogation
5. Social Media and Cultural Signals
Captures informal communication and cultural context markers. Examples: emoji_sentiment_alignment, hashtag_relevance, slang_density, code_mixing_ratio, meme_reference_likelihood, cultural_resonance, acronym_frequency
6. Readability, Cohesion, and Cognitive Load
Measures complexity, structure, and interpretability of content. Examples: cognitive_load_score, information_density, semantic_entropy, surprisal_score, novelty_score, conceptual_depth, abstraction_level, context_dependence, self_containment, attention_retention_curve
7. Argumentation and Epistemology
Captures logical structure, claims, and evidence. Examples: argument_polarity, causal_density, claim_verifiability, evidence_implicitness, claim_explicitness, evidence_strength, logical_coherence
8. Subjectivity and Stance
Represents opinion strength and perspective diversity. Examples: opinionatedness, subjectivity_score, stance_towards_topic, direct_opinion_presence, implicit_stance_cues, perspective_multiplicity
9. Style, Formality, and Tone
Captures writing style and tonal consistency. Examples: temporal_drift, tone_stability, hook_strength, memory_salience, formality_score, academic_tone_intensity, conversational_fluidity, authoritative_tone
10. Toxicity, Controversy, and Discourse
Measures conversational impact and conflict potential. Examples: discursive_engagement, quote_worthiness, reader_alignment, controversy_potential, toxicity_score, insult_intensity
11. Factuality, Misinformation, and Credibility
Evaluates truthfulness and information reliability. Examples: truth_alignment_score, sensationalism_index, rumor_stance, source_attribution_clarity, clickbait_potential, fabrication_likelihood
12. Persuasion and Propaganda
Captures rhetorical influence and manipulation strategies. Examples: pathos_appeal, ethos_appeal, loaded_language_density, fear_mongering_index, glittering_generality_score, logical_fallacy_presence, call_to_action_strength
13. Dialogue Acts and Conversational Intent
Represents communicative intent and interaction signals. Examples: information_seeking_intent, agreement_score, disagreement_intensity, directive_intent_score, commissive_intent, acknowledgment_presence
14. Empathy and Psychological Signals
Captures emotional understanding and mental state indicators. Examples: reflection_trigger_score, cognitive_empathy_score, emotional_support_intensity, distress_indicator, validation_score, self_disclosure_depth
15. Humor and Wordplay
Represents humor styles and creative expression. Examples: humor_intensity, pun_presence_score, absurdity_index, self_deprecation_level, dark_humor_probability
Data Sources
This dataset is derived from publicly accessible sources:
- Google Books, Gutenberg (metadata and summaries)
- Wikipedia (author and contextual information)
- Goodreads (aggregated review statistics)
- LLMs (proxy content and synthetic annotation)
No raw proprietary text is redistributed. Only aggregated and derived representations are included.
Annotation Process
LLM-based annotation was used to generate proxy content and evaluative features:
- Prompt-based generation strategy
- Model-generated scores for 100+ features.
Human Validation
- 48 participants evaluated a subset of samples
- Fleiss' kappa: 0.59
Intended Use
This dataset is suitable for:
- Research on persuasion trends across time
- Hypothesis testing in computational social science
- Analysis of relationships between content and reception
- Studying behavior of LLM-generated features
Out-of-Scope Use
This dataset should not be used for:
- Training production models
- Ground-truth benchmarking
- Author profiling or individual evaluation
- High-stakes decision-making systems
- Literary or stylistic analysis of original text
Bias and Limitations
- Popularity bias due to reliance on review platforms
- Demographic attributes may be inferred
- LLM-generated features may reflect model bias
Personal and Sensitive Information
- No personally identifiable information (PII)
- Data is aggregated and non-traceable to individuals
- Contains limited demographic attributes (e.g., gender, nationality)
License
This dataset is released under CC-BY-NC-4.0.
Rationale:
- Derived from mixed sources including platform-restricted data
- Includes synthetic and aggregated features
- Restricted to non-commercial research use
Ethical Considerations
- Dataset is intended for research use only
- Not suitable for decision-making involving individuals
- Potential biases and misuse risks are documented
- No raw user-generated content is directly included
