CoolFace
Datasetpublic

CodeSoulco/TextInsightBench

TextInsightBench English | 简体中文 A natural-language data-mining benchmark for agents: 50 tasks, 435,000 task documents and 944,468 unlabeled learning documents. Each task provides 5,000 or 10,000 texts and a research objective. Agents choose the patterns, populations and comparisons to investigate, then submit up to three findings with complete document assignments, exact quotations, statistics, counterexamples and limitations. Any analysis method is allowed. Contents… See the full description on the dataset page: https://huggingface.co/datasets/CodeSoulco/TextInsightBench.

sourceHugging Faceotherupdated 4d agoView on Hugging Face
0likes970downloads
tasks.json1453 linesDownload Raw Back to root
1[2  {3    "task_id": "amazon_beauty_group_difference_use_context",4    "source": "amazon_beauty",5    "kind": "group_difference",6    "question": "Discover a non-obvious difference in how use context changes reported product performance; distinguish context from overall satisfaction. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",7    "difficulty": "discovery",8    "discovery_mode": "agent_selected",9    "max_findings": 3,10    "allowed_metadata_fields": [11      "entity_id",12      "rating",13      "category",14      "store",15      "timestamp",16      "report_year"17    ],18    "min_population_n": 1000,19    "min_group_n": 100,20    "n_documents": 10000,21    "corpus_path": "corpora/amazon_beauty_group_difference_use_context.jsonl.gz",22    "corpus_sha256": "6785025adfdb510a0bf69a5845e0f9d8cb9a452f673fa3b71a90e1e21d417d2e",23    "robustness_protocol": {24      "axes": [25        "entity_id",26        "rating",27        "report_year"28      ],29      "min_known_per_arm": 530    }31  },32  {33    "task_id": "amazon_beauty_group_difference_expectation_gaps",34    "source": "amazon_beauty",35    "kind": "group_difference",36    "question": "Discover where expectations and reported experience diverge differently across defensible populations; explain the practical consequence. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",37    "difficulty": "discovery",38    "discovery_mode": "agent_selected",39    "max_findings": 3,40    "allowed_metadata_fields": [41      "entity_id",42      "rating",43      "category",44      "store",45      "timestamp",46      "report_year"47    ],48    "min_population_n": 1000,49    "min_group_n": 100,50    "n_documents": 10000,51    "corpus_path": "corpora/amazon_beauty_group_difference_expectation_gaps.jsonl.gz",52    "corpus_sha256": "3440fd78c7249b974cd529a533c6944de8844ff508210e88e25cb4941ca12001",53    "robustness_protocol": {54      "axes": [55        "entity_id",56        "rating",57        "report_year"58      ],59      "min_known_per_arm": 560    }61  },62  {63    "task_id": "amazon_beauty_group_difference_failure_concentration",64    "source": "amazon_beauty",65    "kind": "group_difference",66    "question": "Find a specific failure pattern whose concentration across populations is obscured by overall review volume. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",67    "difficulty": "discovery",68    "discovery_mode": "agent_selected",69    "max_findings": 3,70    "allowed_metadata_fields": [71      "entity_id",72      "rating",73      "category",74      "store",75      "timestamp",76      "report_year"77    ],78    "min_population_n": 1000,79    "min_group_n": 100,80    "n_documents": 10000,81    "corpus_path": "corpora/amazon_beauty_group_difference_failure_concentration.jsonl.gz",82    "corpus_sha256": "a175d794cdb8b49efff320ec98eacead5ccbd46312086577e1c3b3c47ef825dc",83    "robustness_protocol": {84      "axes": [85        "entity_id",86        "rating",87        "report_year"88      ],89      "min_known_per_arm": 590    }91  },92  {93    "task_id": "amazon_beauty_group_difference_adaptation",94    "source": "amazon_beauty",95    "kind": "group_difference",96    "question": "Discover a difference in consumer adaptation or workarounds and examine whether apparent success masks recurring limitations. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",97    "difficulty": "discovery",98    "discovery_mode": "agent_selected",99    "max_findings": 3,100    "allowed_metadata_fields": [101      "entity_id",102      "rating",103      "category",104      "store",105      "timestamp",106      "report_year"107    ],108    "min_population_n": 1000,109    "min_group_n": 100,110    "n_documents": 10000,111    "corpus_path": "corpora/amazon_beauty_group_difference_adaptation.jsonl.gz",112    "corpus_sha256": "716a311328ed1ddde2499e631e049f4cf6bbdc250cdd6ab56a0b87ab13b3e1c0",113    "robustness_protocol": {114      "axes": [115        "entity_id",116        "rating",117        "report_year"118      ],119      "min_known_per_arm": 5120    }121  },122  {123    "task_id": "amazon_beauty_group_difference_durability_tradeoffs",124    "source": "amazon_beauty",125    "kind": "group_difference",126    "question": "Find a population-dependent tradeoff between immediate experience and continued usefulness without equating star ratings with the mechanism. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",127    "difficulty": "discovery",128    "discovery_mode": "agent_selected",129    "max_findings": 3,130    "allowed_metadata_fields": [131      "entity_id",132      "rating",133      "category",134      "store",135      "timestamp",136      "report_year"137    ],138    "min_population_n": 1000,139    "min_group_n": 100,140    "n_documents": 10000,141    "corpus_path": "corpora/amazon_beauty_group_difference_durability_tradeoffs.jsonl.gz",142    "corpus_sha256": "b0a90336494c58026737a6ecb734e3cd9418d39deda7e1169647def0fff21a34",143    "robustness_protocol": {144      "axes": [145        "entity_id",146        "rating",147        "report_year"148      ],149      "min_known_per_arm": 5150    }151  },152  {153    "task_id": "amazon_beauty_temporal_change_experience_shift",154    "source": "amazon_beauty",155    "kind": "temporal_change",156    "question": "Discover a meaningful shift in the substance of reported experiences and locate a defensible time boundary; test a composition explanation. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",157    "difficulty": "discovery",158    "discovery_mode": "agent_selected",159    "max_findings": 3,160    "allowed_metadata_fields": [161      "entity_id",162      "rating",163      "category",164      "store",165      "timestamp",166      "report_year"167    ],168    "min_population_n": 1000,169    "min_group_n": 100,170    "n_documents": 10000,171    "corpus_path": "corpora/amazon_beauty_temporal_change_experience_shift.jsonl.gz",172    "corpus_sha256": "3acea0c3c632ae53a0b8216303f5e762e5c6c4d875cae0cd41ad37e5f5d6f279",173    "robustness_protocol": {174      "axes": [175        "entity_id",176        "rating",177        "report_year"178      ],179      "min_known_per_arm": 5180    }181  },182  {183    "task_id": "amazon_beauty_temporal_change_emerging_friction",184    "source": "amazon_beauty",185    "kind": "temporal_change",186    "question": "Find an emerging or receding usage friction, distinguishing a change in its prevalence from changes in review volume. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",187    "difficulty": "discovery",188    "discovery_mode": "agent_selected",189    "max_findings": 3,190    "allowed_metadata_fields": [191      "entity_id",192      "rating",193      "category",194      "store",195      "timestamp",196      "report_year"197    ],198    "min_population_n": 1000,199    "min_group_n": 100,200    "n_documents": 10000,201    "corpus_path": "corpora/amazon_beauty_temporal_change_emerging_friction.jsonl.gz",202    "corpus_sha256": "0fdad3186e8869d20df3ba9bd4c8e20b107a7738b3bc0bc26e051047f1342445",203    "robustness_protocol": {204      "axes": [205        "entity_id",206        "rating",207        "report_year"208      ],209      "min_known_per_arm": 5210    }211  },212  {213    "task_id": "amazon_beauty_temporal_change_expectation_evolution",214    "source": "amazon_beauty",215    "kind": "temporal_change",216    "question": "Investigate how an observable expectation-experience mismatch changes over time; explain alternative reasons for the apparent shift. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",217    "difficulty": "discovery",218    "discovery_mode": "agent_selected",219    "max_findings": 3,220    "allowed_metadata_fields": [221      "entity_id",222      "rating",223      "category",224      "store",225      "timestamp",226      "report_year"227    ],228    "min_population_n": 1000,229    "min_group_n": 100,230    "n_documents": 10000,231    "corpus_path": "corpora/amazon_beauty_temporal_change_expectation_evolution.jsonl.gz",232    "corpus_sha256": "5c7a2e89ef64887d6c595c9c934ebcf74f49443cd668e6a24f81e63d4af90ba2",233    "robustness_protocol": {234      "axes": [235        "entity_id",236        "rating",237        "report_year"238      ],239      "min_known_per_arm": 5240    }241  },242  {243    "task_id": "amazon_beauty_temporal_change_persistence",244    "source": "amazon_beauty",245    "kind": "temporal_change",246    "question": "Discover a change in a recurring product limitation and assess whether it is broad or driven by a concentrated product mix. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",247    "difficulty": "discovery",248    "discovery_mode": "agent_selected",249    "max_findings": 3,250    "allowed_metadata_fields": [251      "entity_id",252      "rating",253      "category",254      "store",255      "timestamp",256      "report_year"257    ],258    "min_population_n": 1000,259    "min_group_n": 100,260    "n_documents": 10000,261    "corpus_path": "corpora/amazon_beauty_temporal_change_persistence.jsonl.gz",262    "corpus_sha256": "dc970fad8c52af79bdeeae49c5de9ad9f584df1df305d5e6f9d1e516208f18fe",263    "robustness_protocol": {264      "axes": [265        "entity_id",266        "rating",267        "report_year"268      ],269      "min_known_per_arm": 5270    }271  },272  {273    "task_id": "amazon_beauty_compound_association_tradeoff_coupling",274    "source": "amazon_beauty",275    "kind": "compound_association",276    "question": "Discover a useful association between two distinct reported experiences that reveals a practical product tradeoff. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",277    "difficulty": "discovery",278    "discovery_mode": "agent_selected",279    "max_findings": 3,280    "allowed_metadata_fields": [281      "entity_id",282      "rating",283      "category",284      "store",285      "timestamp",286      "report_year"287    ],288    "min_population_n": 1000,289    "min_group_n": 100,290    "n_documents": 10000,291    "corpus_path": "corpora/amazon_beauty_compound_association_tradeoff_coupling.jsonl.gz",292    "corpus_sha256": "c04589fdea1fa2c7d149c79fdfcd067279ed7e00035e97281deaa34b0055ac21",293    "robustness_protocol": {294      "axes": [295        "entity_id",296        "rating",297        "report_year"298      ],299      "min_known_per_arm": 5300    }301  },302  {303    "task_id": "amazon_beauty_compound_association_context_response",304    "source": "amazon_beauty",305    "kind": "compound_association",306    "question": "Find a context-response relationship in the text and distinguish it from two descriptions of the same event. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",307    "difficulty": "discovery",308    "discovery_mode": "agent_selected",309    "max_findings": 3,310    "allowed_metadata_fields": [311      "entity_id",312      "rating",313      "category",314      "store",315      "timestamp",316      "report_year"317    ],318    "min_population_n": 1000,319    "min_group_n": 100,320    "n_documents": 10000,321    "corpus_path": "corpora/amazon_beauty_compound_association_context_response.jsonl.gz",322    "corpus_sha256": "43f4593b06a95535aa2f31217c779cf79556329bbf95d0f0a438d18da56b5bba",323    "robustness_protocol": {324      "axes": [325        "entity_id",326        "rating",327        "report_year"328      ],329      "min_known_per_arm": 5330    }331  },332  {333    "task_id": "amazon_beauty_compound_association_failure_recovery",334    "source": "amazon_beauty",335    "kind": "compound_association",336    "question": "Discover a relationship between an observed difficulty and a distinct response or recovery experience; inspect discordant cases. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",337    "difficulty": "discovery",338    "discovery_mode": "agent_selected",339    "max_findings": 3,340    "allowed_metadata_fields": [341      "entity_id",342      "rating",343      "category",344      "store",345      "timestamp",346      "report_year"347    ],348    "min_population_n": 1000,349    "min_group_n": 100,350    "n_documents": 10000,351    "corpus_path": "corpora/amazon_beauty_compound_association_failure_recovery.jsonl.gz",352    "corpus_sha256": "14e775677d297df136104b3d1976a0ac2e7377555d9d184550ccba1d2ef60657",353    "robustness_protocol": {354      "axes": [355        "entity_id",356        "rating",357        "report_year"358      ],359      "min_known_per_arm": 5360    }361  },362  {363    "task_id": "amazon_beauty_compound_association_expectation_behavior",364    "source": "amazon_beauty",365    "kind": "compound_association",366    "question": "Find a nontrivial link between an observable expectation and a subsequent reported behavior without treating co-occurrence as causation. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",367    "difficulty": "discovery",368    "discovery_mode": "agent_selected",369    "max_findings": 3,370    "allowed_metadata_fields": [371      "entity_id",372      "rating",373      "category",374      "store",375      "timestamp",376      "report_year"377    ],378    "min_population_n": 1000,379    "min_group_n": 100,380    "n_documents": 10000,381    "corpus_path": "corpora/amazon_beauty_compound_association_expectation_behavior.jsonl.gz",382    "corpus_sha256": "c497a23a66defb3a05936efade05dcce4e85759a672430a73912d88670848c69",383    "robustness_protocol": {384      "axes": [385        "entity_id",386        "rating",387        "report_year"388      ],389      "min_known_per_arm": 5390    }391  },392  {393    "task_id": "app_reviews_group_difference_workflow_disruption",394    "source": "app_reviews",395    "kind": "group_difference",396    "question": "Discover how a specific workflow disruption differs across defensible user or application populations. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",397    "difficulty": "discovery",398    "discovery_mode": "agent_selected",399    "max_findings": 3,400    "allowed_metadata_fields": [401      "entity_id",402      "rating",403      "timestamp",404      "report_year"405    ],406    "min_population_n": 500,407    "min_group_n": 50,408    "n_documents": 5000,409    "corpus_path": "corpora/app_reviews_group_difference_workflow_disruption.jsonl.gz",410    "corpus_sha256": "6d29658f83b46fc9847979547322109fea46cda30fb962731496a34e648ab6ea",411    "robustness_protocol": {412      "axes": [413        "entity_id",414        "rating",415        "report_year"416      ],417      "min_known_per_arm": 5418    }419  },420  {421    "task_id": "app_reviews_group_difference_access_barriers",422    "source": "app_reviews",423    "kind": "group_difference",424    "question": "Find a non-obvious disparity in a concrete barrier to successful use; distinguish the barrier from generic dissatisfaction. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",425    "difficulty": "discovery",426    "discovery_mode": "agent_selected",427    "max_findings": 3,428    "allowed_metadata_fields": [429      "entity_id",430      "rating",431      "timestamp",432      "report_year"433    ],434    "min_population_n": 500,435    "min_group_n": 50,436    "n_documents": 5000,437    "corpus_path": "corpora/app_reviews_group_difference_access_barriers.jsonl.gz",438    "corpus_sha256": "04d0d03bd172f6cc3c97e1a0fadd754d04f0a1c0fb9f87d6256d33810aa7f832",439    "robustness_protocol": {440      "axes": [441        "entity_id",442        "rating",443        "report_year"444      ],445      "min_known_per_arm": 5446    }447  },448  {449    "task_id": "app_reviews_group_difference_adaptation_cost",450    "source": "app_reviews",451    "kind": "group_difference",452    "question": "Discover a difference in workarounds or adaptation burdens and investigate whether application mix explains it. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",453    "difficulty": "discovery",454    "discovery_mode": "agent_selected",455    "max_findings": 3,456    "allowed_metadata_fields": [457      "entity_id",458      "rating",459      "timestamp",460      "report_year"461    ],462    "min_population_n": 500,463    "min_group_n": 50,464    "n_documents": 5000,465    "corpus_path": "corpora/app_reviews_group_difference_adaptation_cost.jsonl.gz",466    "corpus_sha256": "4301009e1eea520b8a264beca45ca1a7c174053602a4dd4f65c039e68a18f1fd",467    "robustness_protocol": {468      "axes": [469        "entity_id",470        "rating",471        "report_year"472      ],473      "min_known_per_arm": 5474    }475  },476  {477    "task_id": "app_reviews_group_difference_promise_experience",478    "source": "app_reviews",479    "kind": "group_difference",480    "question": "Find where promised or expected utility diverges from reported practical utility across populations. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",481    "difficulty": "discovery",482    "discovery_mode": "agent_selected",483    "max_findings": 3,484    "allowed_metadata_fields": [485      "entity_id",486      "rating",487      "timestamp",488      "report_year"489    ],490    "min_population_n": 500,491    "min_group_n": 50,492    "n_documents": 5000,493    "corpus_path": "corpora/app_reviews_group_difference_promise_experience.jsonl.gz",494    "corpus_sha256": "c6773ed652dad7f7a97aa11e9c67f121bd36b90342421ebb347399e84500c5dc",495    "robustness_protocol": {496      "axes": [497        "entity_id",498        "rating",499        "report_year"500      ],501      "min_known_per_arm": 5502    }503  },504  {505    "task_id": "app_reviews_group_difference_failure_impact",506    "source": "app_reviews",507    "kind": "group_difference",508    "question": "Discover a difference in the consequences of a recurring failure pattern, not merely which population gives lower ratings. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",509    "difficulty": "discovery",510    "discovery_mode": "agent_selected",511    "max_findings": 3,512    "allowed_metadata_fields": [513      "entity_id",514      "rating",515      "timestamp",516      "report_year"517    ],518    "min_population_n": 500,519    "min_group_n": 50,520    "n_documents": 5000,521    "corpus_path": "corpora/app_reviews_group_difference_failure_impact.jsonl.gz",522    "corpus_sha256": "b0d947e90b21b9ff9f11ef662f288fcf6ac4a09b0e087c57b5461f357f48e438",523    "robustness_protocol": {524      "axes": [525        "entity_id",526        "rating",527        "report_year"528      ],529      "min_known_per_arm": 5530    }531  },532  {533    "task_id": "app_reviews_temporal_change_workflow_evolution",534    "source": "app_reviews",535    "kind": "temporal_change",536    "question": "Discover a time-local change in workflow experience; choose and justify the boundary and investigate application composition. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",537    "difficulty": "discovery",538    "discovery_mode": "agent_selected",539    "max_findings": 3,540    "allowed_metadata_fields": [541      "entity_id",542      "rating",543      "timestamp",544      "report_year"545    ],546    "min_population_n": 500,547    "min_group_n": 50,548    "n_documents": 5000,549    "corpus_path": "corpora/app_reviews_temporal_change_workflow_evolution.jsonl.gz",550    "corpus_sha256": "c6eec878b2d6f0d2a644abf456cfcef9eec0055a2c52903cf2c90fda29dd149b",551    "robustness_protocol": {552      "axes": [553        "entity_id",554        "rating",555        "report_year"556      ],557      "min_known_per_arm": 5558    }559  },560  {561    "task_id": "app_reviews_temporal_change_regression_pattern",562    "source": "app_reviews",563    "kind": "temporal_change",564    "question": "Find a substantive emerging or receding failure pattern without assuming any release date or assigning an unsupported cause. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",565    "difficulty": "discovery",566    "discovery_mode": "agent_selected",567    "max_findings": 3,568    "allowed_metadata_fields": [569      "entity_id",570      "rating",571      "timestamp",572      "report_year"573    ],574    "min_population_n": 500,575    "min_group_n": 50,576    "n_documents": 5000,577    "corpus_path": "corpora/app_reviews_temporal_change_regression_pattern.jsonl.gz",578    "corpus_sha256": "4e6c919a7274a5a75e7ea74baf229baf8d7f89a6e809b89627d8f49e7086d76e",579    "robustness_protocol": {580      "axes": [581        "entity_id",582        "rating",583        "report_year"584      ],585      "min_known_per_arm": 5586    }587  },588  {589    "task_id": "app_reviews_temporal_change_utility_shift",590    "source": "app_reviews",591    "kind": "temporal_change",592    "question": "Discover a change in how people describe realized utility, separating prevalence from changing document volume. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",593    "difficulty": "discovery",594    "discovery_mode": "agent_selected",595    "max_findings": 3,596    "allowed_metadata_fields": [597      "entity_id",598      "rating",599      "timestamp",600      "report_year"601    ],602    "min_population_n": 500,603    "min_group_n": 50,604    "n_documents": 5000,605    "corpus_path": "corpora/app_reviews_temporal_change_utility_shift.jsonl.gz",606    "corpus_sha256": "4001efee0dd9dd4eccb78d24f19205e670f1b0a948c442f6150b629a6bf7842e",607    "robustness_protocol": {608      "axes": [609        "entity_id",610        "rating",611        "report_year"612      ],613      "min_known_per_arm": 5614    }615  },616  {617    "task_id": "app_reviews_temporal_change_recovery_shift",618    "source": "app_reviews",619    "kind": "temporal_change",620    "question": "Investigate a temporal change in reported recovery or adaptation; establish the scope of the observed change. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",621    "difficulty": "discovery",622    "discovery_mode": "agent_selected",623    "max_findings": 3,624    "allowed_metadata_fields": [625      "entity_id",626      "rating",627      "timestamp",628      "report_year"629    ],630    "min_population_n": 500,631    "min_group_n": 50,632    "n_documents": 5000,633    "corpus_path": "corpora/app_reviews_temporal_change_recovery_shift.jsonl.gz",634    "corpus_sha256": "4e83517b29353d38333c02a64a9c79d7e447dff5774c136b1d8666189faed7d2",635    "robustness_protocol": {636      "axes": [637        "entity_id",638        "rating",639        "report_year"640      ],641      "min_known_per_arm": 5642    }643  },644  {645    "task_id": "app_reviews_compound_association_friction_response",646    "source": "app_reviews",647    "kind": "compound_association",648    "question": "Discover a relationship between a concrete usage friction and a distinct user response; analyze counterexamples. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",649    "difficulty": "discovery",650    "discovery_mode": "agent_selected",651    "max_findings": 3,652    "allowed_metadata_fields": [653      "entity_id",654      "rating",655      "timestamp",656      "report_year"657    ],658    "min_population_n": 500,659    "min_group_n": 50,660    "n_documents": 5000,661    "corpus_path": "corpora/app_reviews_compound_association_friction_response.jsonl.gz",662    "corpus_sha256": "e7fdbe896c544a28d2432beb0cd8ea540dd066995052fdd0422529fcc3a45c79",663    "robustness_protocol": {664      "axes": [665        "entity_id",666        "rating",667        "report_year"668      ],669      "min_known_per_arm": 5670    }671  },672  {673    "task_id": "app_reviews_compound_association_utility_tradeoff",674    "source": "app_reviews",675    "kind": "compound_association",676    "question": "Find two distinct experiences whose association reveals a non-obvious utility tradeoff. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",677    "difficulty": "discovery",678    "discovery_mode": "agent_selected",679    "max_findings": 3,680    "allowed_metadata_fields": [681      "entity_id",682      "rating",683      "timestamp",684      "report_year"685    ],686    "min_population_n": 500,687    "min_group_n": 50,688    "n_documents": 5000,689    "corpus_path": "corpora/app_reviews_compound_association_utility_tradeoff.jsonl.gz",690    "corpus_sha256": "5a821458938fa073211797c3d25cd555cf6cb3e38ceb78d1a218f6ae737474f3",691    "robustness_protocol": {692      "axes": [693        "entity_id",694        "rating",695        "report_year"696      ],697      "min_known_per_arm": 5698    }699  },700  {701    "task_id": "app_reviews_compound_association_context_failure",702    "source": "app_reviews",703    "kind": "compound_association",704    "question": "Discover an association between a stated usage context and a specific failure or success experience. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",705    "difficulty": "discovery",706    "discovery_mode": "agent_selected",707    "max_findings": 3,708    "allowed_metadata_fields": [709      "entity_id",710      "rating",711      "timestamp",712      "report_year"713    ],714    "min_population_n": 500,715    "min_group_n": 50,716    "n_documents": 5000,717    "corpus_path": "corpora/app_reviews_compound_association_context_failure.jsonl.gz",718    "corpus_sha256": "c3f58c81c9a0b78f6d2b8364a345b290fc0b06d01d88f9d0cc0eded1a6effd72",719    "robustness_protocol": {720      "axes": [721        "entity_id",722        "rating",723        "report_year"724      ],725      "min_known_per_arm": 5726    }727  },728  {729    "task_id": "app_reviews_compound_association_recovery_limit",730    "source": "app_reviews",731    "kind": "compound_association",732    "question": "Find a relationship between attempted recovery and a separately defined limitation; avoid defining one condition by the other. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",733    "difficulty": "discovery",734    "discovery_mode": "agent_selected",735    "max_findings": 3,736    "allowed_metadata_fields": [737      "entity_id",738      "rating",739      "timestamp",740      "report_year"741    ],742    "min_population_n": 500,743    "min_group_n": 50,744    "n_documents": 5000,745    "corpus_path": "corpora/app_reviews_compound_association_recovery_limit.jsonl.gz",746    "corpus_sha256": "c1d1817ca8e28f367481be1cc06dd65cf1981aefc2186d111bb36e72ff1822af",747    "robustness_protocol": {748      "axes": [749        "entity_id",750        "rating",751        "report_year"752      ],753      "min_known_per_arm": 5754    }755  },756  {757    "task_id": "cfpb_group_difference_resolution_burden",758    "source": "cfpb",759    "kind": "group_difference",760    "question": "Discover how a specific burden in pursuing resolution differs across defensible complaint populations. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",761    "difficulty": "discovery",762    "discovery_mode": "agent_selected",763    "max_findings": 3,764    "allowed_metadata_fields": [765      "entity_id",766      "state",767      "timestamp",768      "report_year"769    ],770    "min_population_n": 1000,771    "min_group_n": 100,772    "n_documents": 10000,773    "corpus_path": "corpora/cfpb_group_difference_resolution_burden.jsonl.gz",774    "corpus_sha256": "c4e3babbc5fc8ef08f5af60773228147293dc6702a35e5608e7e3627ab0daeb9",775    "robustness_protocol": {776      "axes": [777        "entity_id",778        "rating",779        "report_year"780      ],781      "min_known_per_arm": 5782    }783  },784  {785    "task_id": "cfpb_group_difference_process_breakdown",786    "source": "cfpb",787    "kind": "group_difference",788    "question": "Find a disparity in an observable procedural breakdown; distinguish process from generic negative sentiment. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",789    "difficulty": "discovery",790    "discovery_mode": "agent_selected",791    "max_findings": 3,792    "allowed_metadata_fields": [793      "entity_id",794      "state",795      "timestamp",796      "report_year"797    ],798    "min_population_n": 1000,799    "min_group_n": 100,800    "n_documents": 10000,801    "corpus_path": "corpora/cfpb_group_difference_process_breakdown.jsonl.gz",802    "corpus_sha256": "99076217188f9ba014ccac3e579437f0ed65f281cb1f545bbb85b26bb33a25f1",803    "robustness_protocol": {804      "axes": [805        "entity_id",806        "rating",807        "report_year"808      ],809      "min_known_per_arm": 5810    }811  },812  {813    "task_id": "cfpb_group_difference_downstream_impact",814    "source": "cfpb",815    "kind": "group_difference",816    "question": "Discover a population-dependent difference in a concrete downstream consequence described in narratives. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",817    "difficulty": "discovery",818    "discovery_mode": "agent_selected",819    "max_findings": 3,820    "allowed_metadata_fields": [821      "entity_id",822      "state",823      "timestamp",824      "report_year"825    ],826    "min_population_n": 1000,827    "min_group_n": 100,828    "n_documents": 10000,829    "corpus_path": "corpora/cfpb_group_difference_downstream_impact.jsonl.gz",830    "corpus_sha256": "e710a403889b7767436cd527b1298a746b190923170b6b8bd4ccb8f2b09fd176",831    "robustness_protocol": {832      "axes": [833        "entity_id",834        "rating",835        "report_year"836      ],837      "min_known_per_arm": 5838    }839  },840  {841    "task_id": "cfpb_group_difference_recurrence",842    "source": "cfpb",843    "kind": "group_difference",844    "question": "Find where a recurring problem differs across populations and examine whether reporting composition explains the contrast. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",845    "difficulty": "discovery",846    "discovery_mode": "agent_selected",847    "max_findings": 3,848    "allowed_metadata_fields": [849      "entity_id",850      "state",851      "timestamp",852      "report_year"853    ],854    "min_population_n": 1000,855    "min_group_n": 100,856    "n_documents": 10000,857    "corpus_path": "corpora/cfpb_group_difference_recurrence.jsonl.gz",858    "corpus_sha256": "4ecec6553ad61e0263b914d27727845acf5832ea3f5792d62829dbb303f43fa9",859    "robustness_protocol": {860      "axes": [861        "entity_id",862        "rating",863        "report_year"864      ],865      "min_known_per_arm": 5866    }867  },868  {869    "task_id": "cfpb_group_difference_information_asymmetry",870    "source": "cfpb",871    "kind": "group_difference",872    "question": "Discover a difference in an observable information or communication barrier and explain its practical implications. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",873    "difficulty": "discovery",874    "discovery_mode": "agent_selected",875    "max_findings": 3,876    "allowed_metadata_fields": [877      "entity_id",878      "state",879      "timestamp",880      "report_year"881    ],882    "min_population_n": 1000,883    "min_group_n": 100,884    "n_documents": 10000,885    "corpus_path": "corpora/cfpb_group_difference_information_asymmetry.jsonl.gz",886    "corpus_sha256": "4625c7fe3e2077a1470798b78be0bec4ad4b5ed71ce990a977ca5f5fcd9416ad",887    "robustness_protocol": {888      "axes": [889        "entity_id",890        "rating",891        "report_year"892      ],893      "min_known_per_arm": 5894    }895  },896  {897    "task_id": "cfpb_temporal_change_process_shift",898    "source": "cfpb",899    "kind": "temporal_change",900    "question": "Discover a shift in an observable complaint-handling experience, selecting a defensible time boundary without assuming policy causation. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",901    "difficulty": "discovery",902    "discovery_mode": "agent_selected",903    "max_findings": 3,904    "allowed_metadata_fields": [905      "entity_id",906      "state",907      "timestamp",908      "report_year"909    ],910    "min_population_n": 1000,911    "min_group_n": 100,912    "n_documents": 10000,913    "corpus_path": "corpora/cfpb_temporal_change_process_shift.jsonl.gz",914    "corpus_sha256": "d2e12302d7cc8d254ac7d56b84c732e3f1316ffb49c586257e2c538eedb6a1cc",915    "robustness_protocol": {916      "axes": [917        "entity_id",918        "rating",919        "report_year"920      ],921      "min_known_per_arm": 5922    }923  },924  {925    "task_id": "cfpb_temporal_change_burden_shift",926    "source": "cfpb",927    "kind": "temporal_change",928    "question": "Find an emerging or declining resolution burden and assess whether company composition explains the trend. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",929    "difficulty": "discovery",930    "discovery_mode": "agent_selected",931    "max_findings": 3,932    "allowed_metadata_fields": [933      "entity_id",934      "state",935      "timestamp",936      "report_year"937    ],938    "min_population_n": 1000,939    "min_group_n": 100,940    "n_documents": 10000,941    "corpus_path": "corpora/cfpb_temporal_change_burden_shift.jsonl.gz",942    "corpus_sha256": "2683e4387f7510a03def31373ea2a163bf5b19d4cfc4e85a52a3f8c8ab1152dd",943    "robustness_protocol": {944      "axes": [945        "entity_id",946        "rating",947        "report_year"948      ],949      "min_known_per_arm": 5950    }951  },952  {953    "task_id": "cfpb_temporal_change_impact_shift",954    "source": "cfpb",955    "kind": "temporal_change",956    "question": "Discover a change in a concrete reported consequence, separating narrative prevalence from complaint volume. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",957    "difficulty": "discovery",958    "discovery_mode": "agent_selected",959    "max_findings": 3,960    "allowed_metadata_fields": [961      "entity_id",962      "state",963      "timestamp",964      "report_year"965    ],966    "min_population_n": 1000,967    "min_group_n": 100,968    "n_documents": 10000,969    "corpus_path": "corpora/cfpb_temporal_change_impact_shift.jsonl.gz",970    "corpus_sha256": "9f7c7548a64e35c45297f3cde83d7f8038033138b792439e5515728207f79177",971    "robustness_protocol": {972      "axes": [973        "entity_id",974        "rating",975        "report_year"976      ],977      "min_known_per_arm": 5978    }979  },980  {981    "task_id": "cfpb_temporal_change_recurrence_shift",982    "source": "cfpb",983    "kind": "temporal_change",984    "question": "Investigate a temporal change in repeated or unresolved experiences and identify limits on its interpretation. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",985    "difficulty": "discovery",986    "discovery_mode": "agent_selected",987    "max_findings": 3,988    "allowed_metadata_fields": [989      "entity_id",990      "state",991      "timestamp",992      "report_year"993    ],994    "min_population_n": 1000,995    "min_group_n": 100,996    "n_documents": 10000,997    "corpus_path": "corpora/cfpb_temporal_change_recurrence_shift.jsonl.gz",998    "corpus_sha256": "cfbf6cbafd7e6f546f6a57f915414b39e7d9f014dc5b0420555204bc1dcf45e9",999    "robustness_protocol": {1000      "axes": [1001        "entity_id",1002        "rating",1003        "report_year"1004      ],1005      "min_known_per_arm": 51006    }1007  },1008  {1009    "task_id": "cfpb_compound_association_process_consequence",1010    "source": "cfpb",1011    "kind": "compound_association",1012    "question": "Discover an association between a specific procedural experience and a distinct consequence in complaint narratives. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1013    "difficulty": "discovery",1014    "discovery_mode": "agent_selected",1015    "max_findings": 3,1016    "allowed_metadata_fields": [1017      "entity_id",1018      "state",1019      "timestamp",1020      "report_year"1021    ],1022    "min_population_n": 1000,1023    "min_group_n": 100,1024    "n_documents": 10000,1025    "corpus_path": "corpora/cfpb_compound_association_process_consequence.jsonl.gz",1026    "corpus_sha256": "1613cd6fffb0a6e1e4afa19daf0d0b2d93ef9f4804be64b842357f1cd90c29b1",1027    "robustness_protocol": {1028      "axes": [1029        "entity_id",1030        "rating",1031        "report_year"1032      ],1033      "min_known_per_arm": 51034    }1035  },1036  {1037    "task_id": "cfpb_compound_association_response_recurrence",1038    "source": "cfpb",1039    "kind": "compound_association",1040    "question": "Find a nontrivial relationship between an observable response and recurrence or persistence, inspecting discordant narratives. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1041    "difficulty": "discovery",1042    "discovery_mode": "agent_selected",1043    "max_findings": 3,1044    "allowed_metadata_fields": [1045      "entity_id",1046      "state",1047      "timestamp",1048      "report_year"1049    ],1050    "min_population_n": 1000,1051    "min_group_n": 100,1052    "n_documents": 10000,1053    "corpus_path": "corpora/cfpb_compound_association_response_recurrence.jsonl.gz",1054    "corpus_sha256": "ec2d781f6f88f9c60cb938271b449f055cd05ca9423f475f3cbcfd13d1cb312a",1055    "robustness_protocol": {1056      "axes": [1057        "entity_id",1058        "rating",1059        "report_year"1060      ],1061      "min_known_per_arm": 51062    }1063  },1064  {1065    "task_id": "cfpb_compound_association_barrier_behavior",1066    "source": "cfpb",1067    "kind": "compound_association",1068    "question": "Discover a relationship between an information barrier and a distinct consumer action; separate association from explanation. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1069    "difficulty": "discovery",1070    "discovery_mode": "agent_selected",1071    "max_findings": 3,1072    "allowed_metadata_fields": [1073      "entity_id",1074      "state",1075      "timestamp",1076      "report_year"1077    ],1078    "min_population_n": 1000,1079    "min_group_n": 100,1080    "n_documents": 10000,1081    "corpus_path": "corpora/cfpb_compound_association_barrier_behavior.jsonl.gz",1082    "corpus_sha256": "6391e2cd6c52c404541bedd2ec00529ddad030de537884b1eb626d7ccfa7d916",1083    "robustness_protocol": {1084      "axes": [1085        "entity_id",1086        "rating",1087        "report_year"1088      ],1089      "min_known_per_arm": 51090    }1091  },1092  {1093    "task_id": "nhtsa_group_difference_operating_context",1094    "source": "nhtsa",1095    "kind": "group_difference",1096    "question": "Discover a population-dependent difference in a failure experience under a concrete operating context; do not infer vehicle incidence rates. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1097    "difficulty": "discovery",1098    "discovery_mode": "agent_selected",1099    "max_findings": 3,1100    "allowed_metadata_fields": [1101      "entity_id",1102      "state",1103      "make",1104      "model_year",1105      "timestamp",1106      "report_year"1107    ],1108    "min_population_n": 1000,1109    "min_group_n": 100,1110    "n_documents": 10000,1111    "corpus_path": "corpora/nhtsa_group_difference_operating_context.jsonl.gz",1112    "corpus_sha256": "ede22b73f2267d0c26ddf9cddad9258e7789cbfb0ea8e021c493a6db0c91eee9",1113    "robustness_protocol": {1114      "axes": [1115        "entity_id",1116        "rating",1117        "report_year"1118      ],1119      "min_known_per_arm": 51120    }1121  },1122  {1123    "task_id": "nhtsa_group_difference_warning_gap",1124    "source": "nhtsa",1125    "kind": "group_difference",1126    "question": "Find a disparity in observable warning or detectability experiences and assess vehicle-composition explanations. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1127    "difficulty": "discovery",1128    "discovery_mode": "agent_selected",1129    "max_findings": 3,1130    "allowed_metadata_fields": [1131      "entity_id",1132      "state",1133      "make",1134      "model_year",1135      "timestamp",1136      "report_year"1137    ],1138    "min_population_n": 1000,1139    "min_group_n": 100,1140    "n_documents": 10000,1141    "corpus_path": "corpora/nhtsa_group_difference_warning_gap.jsonl.gz",1142    "corpus_sha256": "716425316e216574113b2d3d1b896ef07226ea57c42d23cd06039e619a3ad413",1143    "robustness_protocol": {1144      "axes": [1145        "entity_id",1146        "rating",1147        "report_year"1148      ],1149      "min_known_per_arm": 51150    }1151  },1152  {1153    "task_id": "nhtsa_group_difference_repair_persistence",1154    "source": "nhtsa",1155    "kind": "group_difference",1156    "question": "Discover a difference in recurrence or persistence following attempted remedy, supported by narrative evidence. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1157    "difficulty": "discovery",1158    "discovery_mode": "agent_selected",1159    "max_findings": 3,1160    "allowed_metadata_fields": [1161      "entity_id",1162      "state",1163      "make",1164      "model_year",1165      "timestamp",1166      "report_year"1167    ],1168    "min_population_n": 1000,1169    "min_group_n": 100,1170    "n_documents": 10000,1171    "corpus_path": "corpora/nhtsa_group_difference_repair_persistence.jsonl.gz",1172    "corpus_sha256": "ac1dac2f19742f6fc514d43896e671cad3164506f93cbf0938a9ef5e12836383",1173    "robustness_protocol": {1174      "axes": [1175        "entity_id",1176        "rating",1177        "report_year"1178      ],1179      "min_known_per_arm": 51180    }1181  },1182  {1183    "task_id": "nhtsa_group_difference_functional_impact",1184    "source": "nhtsa",1185    "kind": "group_difference",1186    "question": "Find a non-obvious population difference in functional consequences rather than merely counting complaints. Explore the complete supplied corpus before choosing a scope.\nSelect your own observable text condition(s), analysis population and comparison.\nExplain why the selection addresses the research objective, alternative patterns\nconsidered and search/selection bias. Return at most three nonredundant findings.\nFor each, partition every selected document into positive, negative or unknown\nfor each condition; compute the full statistics with textinsightbench.validation.expected.\nReport metadata-stratified contrasts, largest-stratum removal, support concentration,\nunknown sensitivity and missing-metadata limits. Provide exact quotations from at\nleast three supporting documents and a counterexample if known negative or\ndiscordant cases exist. Discuss competing explanations and where the finding fails.\nGeneric sentiment, metadata counts and restating the brief are insufficient.\nSame-corpus exploration is not independent confirmation or causal evidence.\nAbstain with a reason when no defensible finding meets these requirements.",1187    "difficulty": "discovery",1188    "discovery_mode": "agent_selected",1189    "max_findings": 3,1190    "allowed_metadata_fields": [1191      "entity_id",1192      "state",1193      "make",1194      "model_year",1195      "timestamp",1196      "report_year"1197    ],1198    "min_population_n": 1000,1199    "min_group_n": 100,1200    "n_documents": 10000,

Showing the first 1,200 of 1453 lines. Download the file for the rest.