balance
Datasets
All datasets matching “balance”dcvlm-balanced-200b
DCVLM-Balanced (200B tokens)
DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.balanced-copa
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.crop-disease-balanced-5022cantabile-runs
cantabile-runs
Work queue and checkpoint store for the Cantabile dynamics study. The directory tree
is the plan — there is no plan file and no database.
main/<song>/<method>/.gitkeep queued, unclaimed
main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat)
main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints
main/<song>/<method>/<seed>/FAILED crashed, needs a human
A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.Stocks-Quarterly-BalanceSheet
Stocks Quarterly Balance Sheet
This dataset includes quarterly balance sheet data for various stocks.
342,800 rows over 6,893 symbols, 39 columns, covering 1983-06-30 to 2026-06-30. Refreshed monthly.
Strategies Built on This Data
1,067 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,012 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.27, and 37% clear a t-statistic of 1.96 on… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Quarterly-BalanceSheet.ai2thor-perspective-qa-20k-balanced-splits-with-obj
