evidence-first
EvidenceFirst-Audio
EvA Open Data
This folder contains the public EvA caption, QA and audio data.
The audio clips are sourced from AudioSet Strong Labels and are stored as audio
bytes in the parquet shards.
Files
captions.jsonl: caption metadata with fields id, audio_path, caption, and caption_with_asr.
qa.jsonl: instruction QA data using the same relative audio paths as captions.jsonl.
parquet/: audio parquet shards. Each row contains id, audio_path, audio_filename, caption… See the full description on the dataset page: https://huggingface.co/datasets/SatsukiVie/EvidenceFirst-Audio.evidence-first-research-memory
Evidence-First Research Memory — Synthetic Demo Index
This is the small, fully synthetic artifact for the
Evidence-First Research Memory
portfolio project.
한국어 안내
이 Dataset은 Evidence-First Research Memory 포트폴리오의 작고 완전한 synthetic
데모 artifact입니다. 실제 연구 코퍼스가 아니라, LLM 에이전트가 검색한 뒤 원문 근거를
검증하는 retrieval 구조를 재현하기 위한 공개 예제입니다.
Files
events.jsonl — four canonical synthetic source records
index.sqlite — derived SQLite FTS5 and vector index
manifest.json —… See the full description on the dataset page: https://huggingface.co/datasets/mp-juuuns/evidence-first-research-memory.
