datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hyperliquid-ohlcv-1m
Tessera Analytics Hyperliquid Order-Flow OHLCV (1-minute) — free sample
A growing monthly sample of gold_ohlcv_1m, the
order-flow-enriched 1-minute OHLCV dataset served by Tessera Analytics:
open / high / low / close and volume for every minute, plus the order-flow context raw candles can't
show — the aggressor buy/sell volume split, cumulative volume delta (CVD), distinct taker counts,
taker fees and realized PnL, and the hourly forward-filled funding rate.
Coverage: BTC, ETH… See the full description on the dataset page: https://huggingface.co/datasets/tessera-analytics/hyperliquid-ohlcv-1m.syntheticDocQA_government_reports_test_tesseracttesseract-zhen-s2sttesseract-testCIVQA-TesseractOCR-LayoutLM
CIVQA TesseractOCR LayoutLM Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM.
The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices.
Invoice number
Variable… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR-LayoutLM.tessera-experimentos
tessera-experimentos
Resultados dos experimentos sobre
tessera-extraidollm-gpt-5-mini-corrigido:
a mesma questão da OAB apresentada ao modelo com e sem a norma aplicável no contexto.
As condições
condição
o que vai no contexto
sem_lei
nada — a linha de base
com_lei_certa
só a norma da alternativa correta
com_lei_total
todas as normas transcritas na questão
com_lei_errada
a norma de outra questão — o controle
O controle é o que faz o resto… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-experimentos.tessera-8a92b237
tessera — corpus epoch 13, full sweep
Teacher-anchored SFT data harvested from every published Affine (Bittensor
SN120) duel scored against corpus epoch 13 — 230 duel records, chal-00760
through chal-01102, covering 2026-08-16 to 2026-08-24.
42,006 rows over 42,006 distinct turns (one row per turn), drawn from
4,981 trajectories and 3,751 strata. That is 70% of the 59,745-turn epoch-13
corpus, and 2.3× the 18,138 rows of
iamPi/tessera-77d11909,
which sampled a subset of the same… See the full description on the dataset page: https://huggingface.co/datasets/iamPi/tessera-8a92b237.tessera-resultados-tabelas
tessera — tabelas de resultado (com e sem RAG)
As tabelas do relatório em formato tabular, uma por arquivo, .parquet e .csv.
São os resultados de apresentar a mesma questão da OAB com e sem a norma aplicável
no contexto.
Gerações cruas, métricas por modelo e o relatório completo estão no dataset irmão:
juliadollis/tessera-experimentos.
Dataset de origem das questões:
juliadollis/tessera-extraidollm-gpt-5-mini-corrigido.
O desenho
366 questões × 9 modelos × 4… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-resultados-tabelas.CIVQA-TesseractOCR
CIVQA TesseractOCR Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR, and it is suitable for adding labels for the chosen model.
The encoded dataset for LayoutLM model can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR-LayoutLM
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing… See the full description on the dataset page: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR.CIVQA-TesseractOCR-LayoutLM
CIVQA TesseractOCR LayoutLM Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM.
The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices.
Invoice number
Variable… See the full description on the dataset page: https://huggingface.co/datasets/SpringRollMonster/CIVQA-TesseractOCR-LayoutLM.docvqa_test_subsampled_tesseracttessera_sample
Tessera Sample Dataset
This dataset is a small sample of the Tessera dataset developed by the University of Cambridge Earth Observation Group.This sample is intended for testing workflows, experimentation, and demonstration purposes.
Original Tessera dataset: GitHub Repository
Dataset Overview
TESSERA (Temporal Embeddings of Surface Spectra for Earth Representation and Analysis) encodes spatio-temporal information from Earth observation data into compact embeddings.This… See the full description on the dataset page: https://huggingface.co/datasets/torchgeo/tessera_sample.arxivqa_test_subsampled_tesseracttessera-mixed-pages
Tessera mixed-page corpus
Synthetic multilingual web pages with exact span boundaries and multi-label
language annotations, for training and evaluating span-level language
identification.
Why this is synthetic, stated up front
No public corpus annotates span boundaries and languages inside real multilingual
web pages at scale. Annotating one by hand across 479
languages, most of them low-resource, is not feasible and would itself be error-prone
in exactly the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/tessera-mixed-pages.syntheticDocQA_energy_test_tesseracttatdqa_test_tesseractinfovqa_test_subsampled_tesseracttabfquad_test_subsampled_tesseracttessera-77d11909tessera-quantization-research-evidence
Tessera Quantization Research Evidence
This dataset is the primary-source measurement evidence from an ongoing research
program studying calibrated low-bit quantization (ternary, int4, vector-quantized
codebooks) for LLM inference on heterogeneous AMD hardware (RDNA3 iGPU, XDNA1/2
NPU, Zen 4/5 CPU). The work is done in a fork of llama.cpp (project name
"Tessera") that adds calibrated per-tensor ternary/payload4/VQ quantization,
NPU offload, and RDNA3-native GPU kernels.
This is… See the full description on the dataset page: https://huggingface.co/datasets/Tribunus-dev/tessera-quantization-research-evidence.syntheticDocQA_artificial_intelligence_test_tesseractshiftproject_test_tesseractTessera2025
📚 Tessera: Exposing the Challenges of LLM-based Test Generation for Low-Resource Programming Languages
Tessera is a validation benchmark designed to measure how well models can generate unit tests for Low-Resource Programming Languages (LRPLs) — specifically Rust, Go, and Julia.
📌 Purpose
Evaluate how well a model can generate test code, given a focal function's source code and additional context.
📂 Dataset Structure
Each sample contains:
function_name:… See the full description on the dataset page: https://huggingface.co/datasets/Tessera2025/Tessera2025.Audio_FilessyntheticDocQA_healthcare_industry_test_tesseractAIIT-Tessera24B-dataset
AIIT-Tessera24B-dataset
The ~24.5B-token pretraining corpus used to train Tessera 1B, from AIIT-THRESHOLD.
What makes it different: this is not a single lab dump. The bulk is web text (DCLM), but the long tail was hand-selected, source by source — a deliberate curation rather than a firehose.
Contents
Tokens
24,504,827,904 seen (~24.5B); curated set ≈24.66B
Shards
674, tokenized with the Tessera tokenizer (byte-level BPE, vocab 65,536, memory-organ… See the full description on the dataset page: https://huggingface.co/datasets/AIIT-Threshold/AIIT-Tessera24B-dataset.Tessera-WADT-Dilemmas
Tessera WADT Dilemmas
WADT — Wike Adversarial Dilemma Training. 658 structured ethical-dilemma pairs
built to train a model to commit to a decision under pressure instead of hedging,
flattering, or deferring — the opposite instinct of a sycophantic model, applied
to hard cases with no clean answer.
Why this exists
Most "AI ethics" training data teaches a model to discuss dilemmas. WADT trains
a model to decide — every example follows a fixed structure: name the… See the full description on the dataset page: https://huggingface.co/datasets/AIIT-Threshold/Tessera-WADT-Dilemmas.tessera-extraidollm
tessera-extraidollm
O dataset tessera com a norma
aplicável extraída do comentário por um LLM.
Questões de múltipla escolha de direito brasileiro (Exame de Ordem, edições 38 a 45). O
comentário oficial de cada questão transcreve o texto da lei ao justificar a resposta — a
extração isola esse texto.
O funil
questões no CSV original
640
− anuladas
−16
− sem as quatro letras A/B/C/D
−10
base
614
com norma extraída
522 (85.0%)
com a norma da… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-extraidollm.tessera-gotessera-extraidollm-gpt-5-mini-corrigido
tessera-extraidollm-gpt-5-mini-corrigido
Questões de múltipla escolha de direito brasileiro (Exame de Ordem, edições 38 a 45), com
o texto da norma aplicável extraído do comentário oficial e verificado caractere a
caractere contra ele.
É a versão corrigida de
tessera-extraidollm-gpt-5-mini, e é a que
deve ser usada. O dado bruto de origem está em
tessera.
Números
questões
614
com norma extraída
508 (82.7%)
com a norma da alternativa correta
372… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-extraidollm-gpt-5-mini-corrigido.
