datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
BEAM
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
This huggingface page contains data for the paper: Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
Abstract
Evaluating the abilities of large language models (LLMs) for tasks that require long-term memory and thus long-context reasoning, for example in conversational settings, is hampered by the existing benchmarks, which often lack narrative coherence, cover narrow… See the full description on the dataset page: https://huggingface.co/datasets/Mohammadta/BEAM.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.BEAM-10M
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
This huggingface page contains data for the paper: Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
Abstract
Evaluating the abilities of large language models (LLMs) for tasks that require long-term memory and thus long-context reasoning, for example in conversational settings, is hampered by the existing benchmarks, which often lack narrative coherence, cover narrow… See the full description on the dataset page: https://huggingface.co/datasets/Mohammadta/BEAM-10M.hakidame_lorabeamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
beam-100k-qa-with-idsgym-aqa-beam-finegym
FineGym-AQA Balance Beam Scoring Setup
Prepared data for fine-tuning an AQA (Action Quality Assessment) model on balance-beam
routines, using the FineGym-AQA annotations (official competition D/E/ND/total scores).
Commercial-intent note (2026-09-21): the end product is commercial. This
FineGym-derived dataset and prototype are research-only (CC BY-NC 4.0); the
commercial model will be trained exclusively on owned, consent-gated data — see
data-collection/README.md for the full… See the full description on the dataset page: https://huggingface.co/datasets/dschultz0404/gym-aqa-beam-finegym.Llama-3.2-1B-Instruct-beam-search-completionsrag_datasetLlama-3.2-3B-Instruct-beam-search-completionsbeam-datasetRC_Beams_Dataset_V1
PN Engineering Datasets
DATA DICTIONARY – Professional Specification
This document defines the structure and attributes of all datasets released
under PN Engineering Datasets.
Dataset Metadata
dataset_name: RC_Beams_Dataset_V1
dataset_version: V1
element_type: Reinforced concrete beams
pdf_count: 25
png_count: 25
File Structure
PDF files:
• Flattened
• Anonymous
• Metadata removed
• OCR-ready
PNG files:
• 1200 DPI resolution
• Clean and uniform background
• High contrast for… See the full description on the dataset page: https://huggingface.co/datasets/PNEngineeringDatasets/RC_Beams_Dataset_V1.BEAMSTER
BEAMSTER
Brain mEtAstases segMentation for STEreotactic Radiotherapy — a
retrospective MRI dataset with expert segmentations. Re-host of the author deposit
figshare 10.6084/m9.figshare.29365844 v1
(CC BY 4.0), from University Hospital Ostrava, Czech Republic, Oct 2019 - Sep 2024.
140 patients, one contrast-enhanced T1w 3D MPRAGE volume each, with a binary
brain-metastasis mask drawn by a board-certified radiation oncologist (13 years'
experience) for stereotactic radiotherapy… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/BEAMSTER.BeamRL-TrainData
BeamRL-TrainData
BeamRL-TrainData is a synthetic dataset of beam mechanics question-answer pairs used to train the BeamPERL model via Group Relative Policy Optimization (GRPO) with verifiable reward signals. Each row corresponds to a unique simply supported beam configuration solved symbolically, paired with natural-language questions and ground-truth reaction force answers.
Dataset Details
Property
Value
Rows
180
Beam type
Simply supported (pin at x=0… See the full description on the dataset page: https://huggingface.co/datasets/tphage/BeamRL-TrainData.BeamTreeBeamRL-TrainData
BeamRL-TrainData
BeamRL-TrainData is a synthetic dataset of beam mechanics question-answer pairs used to train the BeamPERL model via Group Relative Policy Optimization (GRPO) with verifiable reward signals. Each row corresponds to a unique simply supported beam configuration solved symbolically, paired with natural-language questions and ground-truth reaction force answers.
Dataset Details
Property
Value
Rows
180
Beam type
Simply supported (pin at x=0… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/BeamRL-TrainData.BEAM_50K
BEAM 50K Synthetic Compact Dataset
This is a wholly synthetic, machine-generated dataset inspired by the column
schema of Mohammadta/BEAM. It is not an official BEAM release and has not
received BEAM's human validation.
Contents
HF conversation records: 2,500
Probing questions nested in those records: 50,000
Probing questions per conversation: 20
Data shards: 25
Generation route: OpenRouter
Requested model: google/gemini-3.8-flash
Each row contains the eight… See the full description on the dataset page: https://huggingface.co/datasets/vm2825/BEAM_50K.new_rag_datasetbeamit-annotated_full_texts_dataset
Dataset Card for "beamit-annotated_full_texts_dataset"
More Information needed
Llama-3.2-1B-Instruct-beam_search_16-completions-a100test_beam
BEAM Compact Gemini Scaling Sample
This is a wholly synthetic, machine-generated scaling sample inspired by the
schema of Mohammadta/BEAM. It is not an official BEAM release and has not
received BEAM's human validation.
Contents
Conversations: 50
Probing questions: 1000
Questions per conversation: 20
Model: gemini-3.8-flash
Each dataset row contains the eight published BEAM columns:
conversation_id, conversation_seed, narratives, user_profile,
conversation_plan… See the full description on the dataset page: https://huggingface.co/datasets/vm2825/test_beam.BeamRL-EvalData
BeamRL-EvalData
BeamRL-EvalData is a synthetic dataset of beam mechanics question-answer pairs used to evaluate the BeamPERL model. It is the companion evaluation set to tphage/BeamRL-TrainData, and is deliberately designed with harder, more varied configurations to test out-of-distribution generalization: a fixed beam length (9*L) and load magnitude (-13*P) are used, but configurations span 1–3 simultaneous point loads and variable support positions (not just pin at x=0 and roller… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/BeamRL-EvalData.slurp-babble-Qwen2.5-Omni-3B-beam-v3restaurant-verified-email-access-in-columbus-ohio-us-173432
Restaurant Verified Email Access in Columbus, Ohio, US
Free sample dataset from BeamStation
Restaurant Verified Email Access in Columbus, Ohio, US
This dataset provides weekly‑verified email addresses for 665 established, independent restaurants (or micro‑chains) located in Columbus, Ohio. Large chain enterprises are excluded, ensuring the list focuses on independent operators. Each record includes a validated email address that has undergone our proprietary verification… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/restaurant-verified-email-access-in-columbus-ohio-us-173432.BeamRL-EvalData
BeamRL-EvalData
BeamRL-EvalData is a synthetic dataset of beam mechanics question-answer pairs used to evaluate the BeamPERL model. It is the companion evaluation set to tphage/BeamRL-TrainData, and is deliberately designed with harder, more varied configurations to test out-of-distribution generalization: a fixed beam length (9*L) and load magnitude (-13*P) are used, but configurations span 1–3 simultaneous point loads and variable support positions (not just pin at x=0 and roller… See the full description on the dataset page: https://huggingface.co/datasets/tphage/BeamRL-EvalData.BeamRL-EvalData-v2
BeamRL-EvalData-v2
Expanded evaluation dataset for BeamPERL (paper, code): 123 symbolic beam-mechanics problems asking for the reaction forces at the supports, with ground-truth answers as signed coefficients of the load symbol P (e.g. 0.222222P).
Categories
category
samples
description
id
30
single point load, supports at beam ends
ood_loads
30
2–4 point loads, supports at beam ends
ood_supports
18
varying support positions, 1–3 point loads… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/BeamRL-EvalData-v2.all-restaurants-in-lee-s-summit-missouri-us-103773
All Restaurants in Lee's Summit, Missouri, US
Free sample dataset from BeamStation
This dataset provides a complete export of every restaurant operating in Lee's Summit, Missouri, United States. It contains 146 records, each representing a distinct establishment with all available profile columns included. The data is refreshed on a weekly basis to keep information current for users who need up-to-date listings. Researchers, developers, and local business analysts can use this… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-lee-s-summit-missouri-us-103773.white-space-finder-in-milwaukee-waukesha-metro-wisconsin-us-111239
White Space Finder in Milwaukee-Waukesha (Metro), Wisconsin, US
Free sample dataset from BeamStation
The "White Space Finder in Milwaukee-Waukesha (Metro), Wisconsin, US" dataset highlights specific markets within the Milwaukee‑Waukesha metropolitan area where consumer demand for a given category is strong while direct competition remains scarce. Each of the 21 records represents a location that satisfies a strict tri‑factor test: high population density indicating ample demand… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/white-space-finder-in-milwaukee-waukesha-metro-wisconsin-us-111239.slurp-babble-Qwen2.5-Omni-3B-beam-v1
