CoolFace
Datasetpublic

annon-neurips-2026/beyond-n-grams

Beyond N-Grams (BNG) Dataset Summary This dataset is designed to study language trends over time. It combines multi-source metadata, aggregated review signals, and LLM-generated features to enable hypothesis-driven and exploratory research on how narrative and reception characteristics evolve over time. The dataset does not contain original book text. Instead, it uses LLM-generated proxy content and derived features to approximate semantic and evaluative… See the full description on the dataset page: https://huggingface.co/datasets/annon-neurips-2026/beyond-n-grams.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes11downloads
Dataset Card

Beyond N-Grams (BNG)

Dataset Summary

This dataset is designed to study language trends over time. It combines multi-source metadata, aggregated review signals, and LLM-generated features to enable hypothesis-driven and exploratory research on how narrative and reception characteristics evolve over time.

The dataset does not contain original book text. Instead, it uses LLM-generated proxy content and derived features to approximate semantic and evaluative properties.


Dataset Description

Construction Process

The dataset is built using a three-stage pipeline:

  1. 1.Data Collection
  2. 2.Book metadata and summaries collected from:
  3. 3.Google Books
  4. 4.Wikipedia
  5. 5.Goodreads
  6. 6.Project Gutenberg
  7. 7.Includes publication details, author attributes, ratings, and review statistics
  1. 1.Proxy Content Generation
  2. 2.LLM-generated content used to approximate book semantics
  3. 3.Based on research-backed prompting methods
  1. 1.Feature Engineering
  2. 2.~100 structured features per book, including:
  3. 3.Metadata features
  4. 4.Review-based signals
  5. 5.LLM-derived evaluative scores

Dataset Structure

  • —Each row represents a single book
  • —Includes:
  • —Book metadata
  • —Author metadata
  • —Publication year
  • —Aggregated review statistics
  • —Derived feature set
  • —Some fields may contain nested structures (e.g., reviews)

Data Fields

The dataset consists of structured features grouped into three main categories: metadata, review-based signals, and LLM-generated evaluative attributes.


Metadata Features

  • —book_title: Book title/name
  • —book_author: Book author name
  • —gutenberg_book_link: Publically accessible book link
  • —preview_text: Publically available book text
  • —reconstructed_text: Proxy text of the book generated by the model
  • —author_degree_level: Highest known academic qualification of the author
  • —author_gender: Gender of the author
  • —author_nationality: Nationality of the author
  • —pub_year: Year of publication
  • —genre_count: Number of genres associated with the book
  • —genres: List of genres

Review and Engagement Features

  • —review_count: Number of reviews included
  • —reviews: Nested list containing:
  • —reviewer_review_count: Total reviews by reviewer (if available)
  • —reviewer_follower_count: Follower count of reviewer
  • —rating_out_of_5: Rating given in the review
  • —review_year: Year of the review
  • —likes: Number of likes on the review
  • —review_length: Length of the review
  • —total_reviews: Total number of reviews
  • —total_ratings: Total number of ratings
  • —average_rating: Average rating score
  • —early_avg_rating: Early-stage average rating (if available)
  • —star_1_count to star_5_count: Distribution of ratings across 1–5 stars

LLM-Generated Features

The dataset includes ~100 LLM-generated features grouped into the following categories:


1. Sentiment and Polarity

Includes features measuring overall sentiment, intensity, shifts, and implicit emotional signals. Examples: polarity_score, sentiment_intensity, aspect_based_sentiment, sentiment_shift, implicit_sentiment, mixed_sentiment_ratio


2. Emotion and Affect

Captures emotional states based on Ekman categories and affective dynamics. Examples: ekman_anger_intensity, ekman_joy_intensity, ekman_sadness_intensity, ekman_fear_intensity, ekman_surprise_intensity, affective_arousal, emotional_variance


3. Sarcasm, Irony, and Figurative Language

Measures figurative expression, ambiguity, and creative language use. Examples: sarcasm_probability, situational_irony, hyperbole_intensity, linguistic_creativity, idiomatic_density, interpretation_flexibility, ambiguity_score


4. Political Ideology and Populism

Represents ideological stance and group dynamics. Examples: left_right_leaning, authoritarian_libertarian, populist_rhetoric_intensity, anti_elitism_score, in_group_favoritism, out_group_derogation


5. Social Media and Cultural Signals

Captures informal communication and cultural context markers. Examples: emoji_sentiment_alignment, hashtag_relevance, slang_density, code_mixing_ratio, meme_reference_likelihood, cultural_resonance, acronym_frequency


6. Readability, Cohesion, and Cognitive Load

Measures complexity, structure, and interpretability of content. Examples: cognitive_load_score, information_density, semantic_entropy, surprisal_score, novelty_score, conceptual_depth, abstraction_level, context_dependence, self_containment, attention_retention_curve


7. Argumentation and Epistemology

Captures logical structure, claims, and evidence. Examples: argument_polarity, causal_density, claim_verifiability, evidence_implicitness, claim_explicitness, evidence_strength, logical_coherence


8. Subjectivity and Stance

Represents opinion strength and perspective diversity. Examples: opinionatedness, subjectivity_score, stance_towards_topic, direct_opinion_presence, implicit_stance_cues, perspective_multiplicity


9. Style, Formality, and Tone

Captures writing style and tonal consistency. Examples: temporal_drift, tone_stability, hook_strength, memory_salience, formality_score, academic_tone_intensity, conversational_fluidity, authoritative_tone


10. Toxicity, Controversy, and Discourse

Measures conversational impact and conflict potential. Examples: discursive_engagement, quote_worthiness, reader_alignment, controversy_potential, toxicity_score, insult_intensity


11. Factuality, Misinformation, and Credibility

Evaluates truthfulness and information reliability. Examples: truth_alignment_score, sensationalism_index, rumor_stance, source_attribution_clarity, clickbait_potential, fabrication_likelihood


12. Persuasion and Propaganda

Captures rhetorical influence and manipulation strategies. Examples: pathos_appeal, ethos_appeal, loaded_language_density, fear_mongering_index, glittering_generality_score, logical_fallacy_presence, call_to_action_strength


13. Dialogue Acts and Conversational Intent

Represents communicative intent and interaction signals. Examples: information_seeking_intent, agreement_score, disagreement_intensity, directive_intent_score, commissive_intent, acknowledgment_presence


14. Empathy and Psychological Signals

Captures emotional understanding and mental state indicators. Examples: reflection_trigger_score, cognitive_empathy_score, emotional_support_intensity, distress_indicator, validation_score, self_disclosure_depth


15. Humor and Wordplay

Represents humor styles and creative expression. Examples: humor_intensity, pun_presence_score, absurdity_index, self_deprecation_level, dark_humor_probability


Data Sources

This dataset is derived from publicly accessible sources:

  • —Google Books, Gutenberg (metadata and summaries)
  • —Wikipedia (author and contextual information)
  • —Goodreads (aggregated review statistics)
  • —LLMs (proxy content and synthetic annotation)

No raw proprietary text is redistributed. Only aggregated and derived representations are included.


Annotation Process

LLM-based annotation was used to generate proxy content and evaluative features:

  • —Prompt-based generation strategy
  • —Model-generated scores for 100+ features.

Human Validation

  • —48 participants evaluated a subset of samples
  • —Fleiss' kappa: 0.59

Intended Use

This dataset is suitable for:

  • —Research on persuasion trends across time
  • —Hypothesis testing in computational social science
  • —Analysis of relationships between content and reception
  • —Studying behavior of LLM-generated features

Out-of-Scope Use

This dataset should not be used for:

  • —Training production models
  • —Ground-truth benchmarking
  • —Author profiling or individual evaluation
  • —High-stakes decision-making systems
  • —Literary or stylistic analysis of original text

Bias and Limitations

  • —Popularity bias due to reliance on review platforms
  • —Demographic attributes may be inferred
  • —LLM-generated features may reflect model bias

Personal and Sensitive Information

  • —No personally identifiable information (PII)
  • —Data is aggregated and non-traceable to individuals
  • —Contains limited demographic attributes (e.g., gender, nationality)

License

This dataset is released under CC-BY-NC-4.0.

Rationale:

  • —Derived from mixed sources including platform-restricted data
  • —Includes synthetic and aggregated features
  • —Restricted to non-commercial research use

Ethical Considerations

  • —Dataset is intended for research use only
  • —Not suitable for decision-making involving individuals
  • —Potential biases and misuse risks are documented
  • —No raw user-generated content is directly included