Comprehensive
llama-3.2-3b-phaedo-ext-comprehensive160-ggufllama-3.2-3b-phaedo-ext-comprehensive335-ggufllama-3.2-3b-phaedo-ext-comprehensive81-ggufllama-3.2-3b-phaedo-ext-comprehensive335-q4f16-MLCLlama-3.2-1B-Instruct-ft-comprehensive-qafasalai-produce-comprehensiveqwen3-1.7b-ft-touch-rugby-comprehensive-qaPhi-4-mini-instruct-touch-rugby-comprehensive-qa
comprehensive-arithmetic-problemscomprehensive-arithmetic-problems-carriesindian-stocks-comprehensive-fundamentals-dataset
Indian Stocks Comprehensive Fundamentals Dataset
From screener.in | 5701 Stocks | 375.20 MB+ Data | Weekly Updates
Highlights :
Total Number of stocks : 5701
Dataset Size : 375.20 MB
Status :
last updated on Friday, 25 Sep 2026 00:39:40 +0000
Usage Notes :
They are stored in 5701 individual files.
[stockname].json means all data related to that stock.
For example, titan.json contains all available fundamental data… See the full description on the dataset page: https://huggingface.co/datasets/AYUSHKHAIRE/indian-stocks-comprehensive-fundamentals-dataset.llm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.Comprehensive-Antiquarian-and-Rare-Books-Archive
Comprehensive Antiquarian & Rare Books Archive
Dataset Description
This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.
