CoolFace
20 results

bart

jbduran /bartholomew-dataset-v3 BART Dataset v3 The final version of the BART pretraining corpus, focused on removing anything that betrays a post-1930 origin. This is our best vintage dataset yet. Documents 146,031 (97.52% of v2) Characters 102,798,688,961 (96.73% of v2) Tokens ~23B (estimated) Shards 473 (one per v2 shard, same basename) Source BART Dataset v2 Cutoff 1930 Schema single string column text Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v3.100K<n<1M1 likes2.8k downloads1mo agoHugging Facejbduran /bartholomew-dataset-v1 BART Dataset v1 The first version of the BART pretraining corpus: pre-1930 English books drawn from Institutional Books 1.0 and filtered hard on OCR quality, language, date, and tokenizability. Documents 160,263 Characters 118,745,375,871 Tokens ~27B (estimated) Shards 473 (472 train + 1 val) Source Institutional Books 1.0 (242B tokens, ~983K documents) Schema single string column text Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.text100K<n<1M1 likes2.5k downloads1mo agoHugging FaceBarthel /variouscryptodata variouscryptodata Crypto market datasets collected as a by-product of our own research and published so they are not lost. One sub-folder per dataset; each appended nightly where collection is still running. folder what coverage cadence polymarket_updown_orderbook/ Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference 2026-05-24 → present appended nightly (previous UTC day) hyperliquid_trades/ Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.tabular1B<n<10B0 likes2k downloads22h agoHugging Facejbduran /bartholomew-dataset-v2 BART Dataset v2 The second version of the BART pretraining corpus, focused on stripping low-quality text — boilerplate and OCR corruption — out of v1. Documents 149,745 (93.44% of v1) Characters 106,274,384,672 (89.50% of v1) Tokens ~24B (estimated) Shards 473 (one per v1 shard, same basename) Source BART Dataset v1 Schema single string column text Lineage — three cumulative filtering stages over the same corpus: Institutional Books 1.0 → v1 → v2 → v3… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.100K<n<1M0 likes1.8k downloads1mo agoHugging Facebartn8 /EventHubDatasetimage1 likes1.4k downloads3mo agoHugging Facebarto17 /speech_commands Dataset Card for "speech_commands" More Information needed timeseries10K<n<100K0 likes853 downloads3y agoHugging Face