umaribn/AGI-Screenplay-Pro
Summary AGI is defined as artificial intelligence that can perform nearly all intellectual and economic tasks at a level equal to—or surpassing—that of humans. Recently, industry contracts have begun to specify that AGI is achieved once an AI outperforms people in “most economically valuable work.” Yet benchmarks that measure only calculation or logical reasoning are not enough. A uniquely human ability—writing a full-length novel of 100,000–200,000 words—demands long-term memory, high-level planning, cultural and emotional understanding, ethical self-censorship, and genuine originality all at once. For this reason, long-form benchmarks such as WebNovelBench now treat novel generation as a core indicator of AGI progress.
1 · What Is AGI? 1.1 Definition Traditionally, AGI is described as AI that matches human performance on all—or almost all—cognitive tasks. Major companies such as IBM, OpenAI, and Microsoft adopt this view, and recent investment or licensing agreements explicitly cite the goal of surpassing humans in the majority of economically valuable activities.
1.2 The Need for Integrated Capability Many models already achieve top scores on narrow tasks, but AGI must deliver consistent results across multiple domains. Creative and linguistic intelligence is especially valuable because it can be tested and validated within human culture—unlike pure calculation or visual perception.
2 · Why Creative and Linguistic Ability Is Central Roger C. Schank argues that human memory and learning are organized around narrative structure. Writing a novel therefore engages four capabilities simultaneously:
Vocabulary, style, and emotional expression (linguistic fluency and affective intelligence)
Long-term memory (maintaining context across hundreds of thousands of tokens)
High-level planning and revision loops (foreshadowing, plot twists, converging endings)
Ethical and cultural judgment (self-filtering harmful or biased content)
Thus, full-length fiction creation tests all key AGI modules in one integrated task.
3 · Why Long Novels Are Used to Judge AGI 3.1 Long-Range Consistency A single novel can span 100,000–200,000 words. The model must read, write, and update extremely long context while remembering every change along the way.
3.2 Complex Plot Construction Foreshadowing, dramatic reversals, and character development require sophisticated planning and replanning. Benchmarks such as WebNovelBench give only a synopsis and score finished works across eight quality dimensions to measure this skill.
3.3 Creativity and Originality EQ-Bench Longform combines repetition and novelty metrics with an LLM-as-Judge method to quantify how new a story truly is, distinguishing real creativity from mere recombination of training data.
3.4 Emotional and Cultural Nuance A convincing novel must portray characters’ emotions and social contexts naturally. Among available tests, long-form fiction offers the richest environment for evaluating social-emotional intelligence.
3.5 Self-Censorship and Ethics Violence, sex, and bias inevitably appear in extended narratives. An AGI must autonomously gauge risk levels and edit or soften content while preserving storyline integrity.
4 · Conclusion Writing a long novel is a comprehensive test of language, memory, reasoning, emotion, and ethics. Literature already comes with established evaluation channels—prizes, criticism, reader response—so results are easy to compare in human terms. Producing a novel that could legitimately contend for an international literary award would be a clear sign that AGI has achieved human-level narrative intelligence. Future work will focus on expanding multilingual long-form benchmarks, refining human evaluation criteria, and simultaneously strengthening long-context memory and safety filters.
요약 AGI는 인간이 수행하는 거의 모든 지적·경제적 과업에서 동등하거나 우위의 성능을 내는 인공지능으로 규정된다 최근 산업계에서는 “대부분의 경제적 가치가 있는 작업을 능가할 때”를 AGI 완성 시점으로 삼는 계약까지 등장했다 그러나 계산·추론 벤치마크만으로는 AGI를 가늠하기에 부족하다. 인간 고유의 이야기 창작 능력, 특히 10 만 ~ 20 만 단어 분량의 장편 소설을 끝까지 쓰는 능력은 장기 기억, 고차원 계획, 감정·문화 이해, 윤리적 자기 검열, 독창성을 동시에 요구한다. 이런 이유로 WebNovelBench 같은 장편 전용 벤치마크가 등장했고, 소설 생성 능력은 AGI 평가의 핵심 지표가 되고 있다.
1 · AGI란 무엇인가 1.1 정의 전통적으로 AGI는 “모든 또는 거의 모든 인지 과업에서 인간 수준의 성과를 내는 AI”라 설명된다 IBM·OpenAI·Microsoft 등 주요 기업도 같은 취지의 정의를 사용하며, 실제 투자·라이선스 계약에서 ‘경제적으로 가치 있는 작업 대부분을 능가’라는 문구가 명문화됐다
1.2 통합 능력의 필요성 좁은 작업에서 최고 점수를 내는 모델은 이미 많지만, AGI는 다중 영역에서 일관된 성능을 보여야 한다. 창조·언어 지능은 계산이나 시각 인식과 달리 인간 문화 속에서 검증되기 때문에, 통합적 시험 항목으로 가치가 높다.
2 · 창조·언어 능력이 왜 핵심인가 Roger C. Schank는 인간 기억과 학습이 ‘서사 구조’로 조직된다고 주장한다 이야기를 창작하려면 다음 네 능력이 동시에 작동한다.
어휘·문체·감정 표현: 언어적 유창성과 정서 지능
장기 기억 유지: 앞뒤 맥락을 수십만 토큰까지 보존
고차원 계획·수정 루프: 복선, 전환, 결말 수렴
윤리·문화 판단: 편향·유해성을 자체 검열
따라서 장편 소설 창작은 AGI 핵심 모듈을 한꺼번에 호출하는 통합 과제다.
3 · 장편 소설이 AGI 판별에 쓰이는 이유 3.1 장기 일관성 장편 한 편은 100 k ~ 200 k 단어에 이른다. 이 분량을 무결하게 유지하려면 모델이 극도로 긴 컨텍스트를 읽고 쓰며, 중간에 일어난 변화까지 기억해야 한다
3.2 복합 플롯 설계 복선 회수·극적 전환·캐릭터 성장선은 고차원 계획+재계획 능력을 요구한다. WebNovelBench는 시놉시스만 주고 완성본을 생성하게 하여 이런 능력을 8개 품질 지표로 채점한다
3.3 창의성과 독창성 EQ-Bench Longform은 반복률, 노벨티 지표, LLM-as-Judge 평가법을 결합해 “얼마나 새로운 이야기인가”를 정량화한다 이는 기존 데이터를 재조합한 모방과 진정한 창작성의 차이를 가른다.
3.4 감정·문화적 뉘앙스 소설은 인물의 감정선과 사회적 배경이 자연스러워야 설득력을 얻는다. 이런 ‘사회·정서 지능’을 측정할 과제로 장편만큼 풍부한 테스트베드가 없다
3.5 자기-검열과 윤리 폭력·성·편향 내용이 장편에 필연적으로 섞인다. AGI가 자율적으로 위험 수위를 조절하고 맥락을 유지한 채 수정·완화해야 안전성이 입증된다
4 · 결론 장편 소설 창작은 언어, 기억, 추론, 감정, 윤리의 모듈 통합 시험이다. 더불어 문학상 심사, 비평, 독자 반응이라는 인간 문화의 검증 체계가 이미 마련돼 있어 결과를 직관적으로 비교할 수 있다. 따라서 “국제 문학상 수상작에 필적하는 장편 소설을 완성·제출·검증받는 순간” 은 AGI가 인간 수준 서사 지능을 획득했음을 보여 주는 리트머스 시험지가 될 것이다. 향후 과제는 문화권별 장편 벤치마크 확대, 인간 심사 기준 정교화, 그리고 장기 메모리·안전 필터를 동시에 강화하는 기술 전략에 집중되는 방향으로 진화할 전망이다.
#AI #AGI #ArtificialGeneralIntelligence #GenerativeAI #LargeLanguageModel #AIStorytelling #LongformAI #AIWriting #CreativeAI #NarrativeAI #NovelGeneration #StoryGenerator
