datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wrapped-asset-parity
Wrapped-Asset Parity Ledger
Council of AI · CSOAI Ltd (UK #16939677) · live source https://councilof.ai/interop/wrapped-asset-parity-latest.json · door GET https://councilof.ai/api/wrapper?id=<pair> (free &preview=1, x402 402 challenge for the signed card)
One row per bridged or custodial wrapper pair, read from public RPC with no key at provider-reported finalized blocks: the wrapped token's totalSupply() on its chain and, where an escrow exists, the canonical token's… See the full description on the dataset page: https://huggingface.co/datasets/csoai/wrapped-asset-parity.osworld-native-parity-runs
OSWorld Native Parity Runs
This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset.
Archives
fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst
Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1
Source upstream: xlang-ai/OSWorld at fe8c78e
Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.parity-juridico-dataset
parity-juridico-dataset
Dataset de fine-tuning para retrieval semântico em domínio jurídico
brasileiro (licitações públicas, jurisprudência TCU, Lei 14.133/21).
Curado a partir do banco de conhecimento da plataforma Parity
(parity.doublethree.com.br).
Estrutura
Arquivo
Conteúdo
Linhas
parity-triplets.jsonl
(anchor, positive, negative) — train split
117
parity-eval.jsonl
mesmas chaves, eval split 10%
14
parity-pairs.jsonl
(texto_a, texto_b, label∈{0,1}) p/… See the full description on the dataset page: https://huggingface.co/datasets/SamuelMauli/parity-juridico-dataset.parity-juridico-dataset-v2
parity-juridico-dataset-v2
Dataset agregado para retrieval jurídico brasileiro. v2 = união de fontes
HF públicas + corpus curado Parity + (opcional) scrapings TCU/PNCP.
Volume
Triplets: 2366 (anchor, positive, negative)
Pairs: 7317 (texto_a, texto_b, label∈{0,1})
Corpus: 26079 textos jurídicos brutos (PT-BR)
Fontes incluídas
parity-v1: triplets=131, pairs=313
assin2: corpus=13000, pairs_pos=3250, pairs_neg=3250
pierreguillou-lener-lm: corpus=12415… See the full description on the dataset page: https://huggingface.co/datasets/SamuelMauli/parity-juridico-dataset-v2.parity-juridico-dataset-v3
parity-juridico-dataset-v3
Dataset jurídico brasileiro PURO — 100% texto jurídico de domínio,
sem ruído de NLI genérico ou frases não-jurídicas.
Volume
Triplets: 195
Pairs: 445
Corpus: 957 textos jurídicos brutos
Diferenças do v2
REMOVIDO: ASSIN2 (NLI genérico — "jóqueis montando cavalos" não é jurídico)
ADICIONADO: filtro regex jurídico estrito em todas as fontes
ADICIONADO: PNCP editais reais via API pública
ADICIONADO: TCU acórdãos via dados.tcu.gov.br… See the full description on the dataset page: https://huggingface.co/datasets/SamuelMauli/parity-juridico-dataset-v3.parity-juridico-dataset-v4
parity-juridico-dataset-v4
Dataset DOMÍNIO PARITY PURO — apenas triplets gerados a partir de:
131 acórdãos/súmulas TCU curados manualmente (v1)
Mining agressivo dos 832 editais Parity em prod
(top-K cosine + match de setor/modalidade/tipoAnalise + hard negatives)
Por que v4?
v3 tentou usar datasets HF jurídicos (LeNER-Br, legal-bench-br) e
REGREDIU em retrieval real (40% → 20%) porque trouxe domínios fora
do core (direito civil/criminal/tributário). v4 corrige usando só… See the full description on the dataset page: https://huggingface.co/datasets/SamuelMauli/parity-juridico-dataset-v4.tpu-kernel-parity-lab-v5e
JAX/Pallas attention measured on one TPU v5e device
I measured a Pallas masked-softmax kernel against the JAX/XLA reference inside grouped-query
attention. The Kaggle worker exposed eight TPU devices. The unsharded arrays ran on JAX's default
single device, so these numbers make no multi-device scaling claim.
All seven correctness checks passed. The benchmark contains twenty cases, with five warmups and
twenty synchronized timing samples in each case. XLA was faster in all ten… See the full description on the dataset page: https://huggingface.co/datasets/LaelaZorana/tpu-kernel-parity-lab-v5e.wbc-khan-parity-datafinetune_parity-aware-bpe
