datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
idt5-v4-results-final-fft-s2026-20260911T063507183162Z
final-fft-s2026-20260911T063507183162Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 88.01498127340824,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s2026-20260911T063507183162Z.idt5-v4-results-final-fft-s42-20260910T135823740810Z
final-fft-s42-20260910T135823740810Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 93.63295880149813,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s42-20260910T135823740810Z.idt5-v4-results-final-fft-s123-20260910T215641650583Z
final-fft-s123-20260910T215641650583Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 93.25842696629213,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s123-20260910T215641650583Z.fftc
Model Information
The model tested in this exercise is https://huggingface.co/Nanbeige/Nanbeige4.1-3B.
This is a 3-billion parameter base model developed by the Nanbeige LLM Lab. It is designed with high reasoning density in math and code but lacks "instruction-tuning," making it prone to task drift and logical inconsistencies when prompted directly.
Reproduction Code
The model was loaded using the transformers library with half-precision (float16) to optimize VRAM usage on a GPU (standard… See the full description on the dataset page: https://huggingface.co/datasets/ODanchi-1/fftc.
