diff --git a/README.md b/README.md index 7f770b7..106b6a0 100644 --- a/README.md +++ b/README.md @@ -8,13 +8,40 @@ ## 快速开始 ```bash uv sync -make db-rebuild # 从 cases/export/*.jsonl 重建 SQLite +make db-rebuild # 从 cases/export/*.jsonl 重建 SQLite(案例检索层需要;--no-retrieval 可跳过) make test-fast # 不调用 LLM 的全部门禁 -nsc run --brief examples/demo_tea/brief.yaml --profile short_drama_v1 -nsc check out/demo_tea/ir.json # L0 检查 -nsc render out/demo_tea/ir.json --target novel_docx script_fountain ``` +**配置 LLM 端点(唯一外部依赖)**:`config/models.yaml` 各 tier 的 `api_base` 指向任一 +**OpenAI 兼容端点**(默认 LongCat-2.0),`export OPENAI_API_KEY=`(绝不落盘)。 + +```bash +uv run nsc run examples/demo_tea/brief.yaml --profile short_drama_v1 +# 产物在 out/<标题>/:novel.md(小说) + script.md(剧本) + ir.json(全量 IR) + manifest.json(溯源) +nsc check out/<标题>/ir.json # L0 检查 +nsc render out/<标题>/ir.json # 重新渲染交付物 +``` + +## 独立使用说明(本仓库无外部仓库依赖) + +本仓库即是完整可用的短剧/小说生成器:7 段编译管线(p0 需求归一→p6 小说→p7 渲染)、 +82 条声明式门禁(`spec/checks/`)、相位内带诊断重试、以及一组**后端无关的机械兜底** +(结构修复/时长缩放/对白欠量扩写/暗线钳制/合规词替换等,在 `src/nsc/passes/`, +只对违规形态触发,换任何 LLM 都生效)。 + +- **换后端**:只改 `config/models.yaml` 的 `api_base`(参考 `config/models.yaml.bak`)。 + 2026-08 Lab 战役期间该文件曾指向本地 shim(`127.0.0.1:8400`),独立使用时改回真实端点即可。 +- **新品牌/新故事**:复制 `brands/demo_tea/` 与 `examples/demo_tea/brief.yaml` 改内容。 + 注意:目录名 = `brand_id`;`banned_words` 不得与 `products.facts` 冲突; + `canonical_name` 必须是最自然的写法(它会被 BM-009 当成唯一合法产品名)。 + 现成第二套范例:`brands/hainan_nolan/` + `examples/hainan_nolan/brief.yaml`(海南文旅 IP)。 +- **profile 决定形态**:`profiles/short_drama_v1.yaml`(6-12 集正片)、 + `profiles/lab_smoke_v1.yaml`(迭代切片);`novel.enabled` 控制是否出小说视图。 +- **全量测试**:`uv run pytest tests --ignore=tests/test_pipeline_llm.py`(无 LLM,597 个)。 +- **分支说明**:2026-08 Lab 优化战役的 10 轮 harness 硬化(round14-20d, + 全部带测试)在分支 `sw/lab-campaign-20260825`;配套的判分/游乐场设施在私有仓 + Script_Writer_Lab(非必需,仅质量测量与优化循环用)。 + ## 三条铁律 1. **`spec/` 是唯一真相。** 任何知识若不能落进 `spec/ir | spec/checks | spec/rubrics | spec/rules | profiles | brands`,视为不存在。 2. **`prompts/`、`src/`、`out/` 是生成物。** 手改 `prompts/` = CI 失败。重写 `src/` 必须能通过同一套 `tests/`。 diff --git a/adr/0015-pass-contract-strings-as-asset.md b/adr/0015-pass-contract-strings-as-asset.md index cd83724..50726fe 100644 --- a/adr/0015-pass-contract-strings-as-asset.md +++ b/adr/0015-pass-contract-strings-as-asset.md @@ -1,6 +1,6 @@ # ADR-0015:Pass 契约文案作为资产(spec/passes/contracts.yaml) -- 状态:proposed +- 状态:accepted - 日期:2026-08-22 - 影响层:A5 知识(+ A6 配置) diff --git a/adr/0016-pipeline-strategy-profile-sections.md b/adr/0016-pipeline-strategy-profile-sections.md index 4d7b663..b79c736 100644 --- a/adr/0016-pipeline-strategy-profile-sections.md +++ b/adr/0016-pipeline-strategy-profile-sections.md @@ -1,6 +1,6 @@ # ADR-0016:管线策略 profile 化(pipeline/retrieval/revise 三段) -- 状态:proposed +- 状态:accepted - 日期:2026-08-22 - 影响层:A 资产层(profiles/_schema.py + 两个 profile yaml);B 层仅消费 diff --git a/adr/0017-p3-context-profile-section.md b/adr/0017-p3-context-profile-section.md index 79ba280..813c639 100644 --- a/adr/0017-p3-context-profile-section.md +++ b/adr/0017-p3-context-profile-section.md @@ -1,6 +1,6 @@ # ADR-0017:p3 fragment 组成数据化(profile.context 段) -- 状态:proposed +- 状态:accepted - 日期:2026-08-22 - 影响层:A 资产层(profiles/_schema.py + profile yaml);B 层仅消费 diff --git a/brands/hainan_nolan/brand.yaml b/brands/hainan_nolan/brand.yaml new file mode 100644 index 0000000..b93ce60 --- /dev/null +++ b/brands/hainan_nolan/brand.yaml @@ -0,0 +1,84 @@ +schema_version: "1.0" +brand_id: hainan_nolan +brand_name: 南浪仔 NOLAN(海南文旅 IP) +version: "1.0.0" +industry: tourism + +products: + - id: nolan_ip + name: 南浪仔 + canonical_name: 南浪仔 + aliases: [南浪仔 NOLAN, NOLAN] + category: 文旅 IP 形象 + facts: + prototype: "原型是海南长臂猿,全世界最稀有的灵长类之一,只生活在海南岛霸王岭的热带雨林" + look: "一身暖阳黄,大眼睛,一双天生的长手臂" + slogan: "一身暖阳黄,自在南浪仔" + +selling_points: + - id: free_spirit + claim: 自在松弛的探索方式:不打卡、不赶路,慵懒但认真 + priority: 1 + must_cover: true + proof: "他玩累了半眯眼晒太阳的样子就是答案" + - id: hainan_diversity + claim: 海南的好玩不止一种:浪尖、雨林、椰林、热气球、日落 + priority: 2 + must_cover: true + proof: "他的足迹遍布全岛" + - id: gibbon_care + claim: 海南长臂猿值得被更多人看见和记住 + priority: 3 + must_cover: false + proof: "每多一个人认识南浪仔,雨林里的家族就多一分被看见的机会" + forbidden_phrasings: ["募捐", "救助", "卖惨"] + +audience: + - id: family_kid + label: 亲子家庭与喜欢治愈系内容的年轻旅行者 + age_range: "5-35" + pains: [旅行变成赶场打卡, 孩子对自然无感, 想带走一份有温度的纪念] + triggers: [一只可爱的长臂猿, 一个温柔的故事, 收集与陪伴] + language_notes: 简单、温暖、不说教 + +usage_scenes: + - id: bawangling_canopy + description: 霸王岭热带雨林树冠层 + shootable: true + - id: bay_surf + description: 海湾浪尖与沙滩 + shootable: true + - id: hot_air_balloon + description: 热气球上俯瞰海岸 + shootable: true + - id: coconut_camp + description: 椰林露营篝火 + shootable: true + - id: city_sunset + description: 城市高处看日落 + shootable: true + +tone_words: [自由, 治愈, 松弛, 温柔, 一点点孤独底色] +banned_words: [打卡, 网红, 必去, 灭绝, 募捐, 卖惨] +must_include_lines: ["一身暖阳黄,自在南浪仔"] +must_include_visuals: [暖阳黄的长臂猿身影] + +placement: + max_moments_per_episode: 2 + min_gap_beats: 2 + max_high_intensity_per_episode: 1 + require_high_plot_connection: 1 + forbid_in_beat_kinds: [hook] + +legal: + banned_words: [] + competitor_names: [] + claim_whitelist: ["全世界最稀有的灵长类之一", "只生活在霸王岭"] + ip_assignment: 交付后著作权归甲方所有 + legal_refs: [] + +account_context: > + 海南文旅 IP「南浪仔 NOLAN」的内容阵地,以盲盒、壁纸、短视频传播 IP 形象。 + 希望用连续短剧与小说让大众认识这只从雨林走向大海的长臂猿, + 把"自在"二字变成海南旅游的情感记忆点。 +business_goal: awareness diff --git a/examples/hainan_nolan/brief.yaml b/examples/hainan_nolan/brief.yaml new file mode 100644 index 0000000..c2c3ad2 --- /dev/null +++ b/examples/hainan_nolan/brief.yaml @@ -0,0 +1,23 @@ +project_title: "南浪仔" +profile: short_drama_v1 +brand: hainan_nolan +raw_request: | + 给海南文旅 IP「南浪仔 NOLAN」做一部 6 集连续短剧(同时出小说版)。 + 他的原型是海南长臂猿——全世界最稀有的灵长类之一,只住在霸王岭热带雨林。 + 故事五幕: + 1. 起点:雨林里最小的探险家。他出生在霸王岭树冠层,是家族里最坐不住的那只; + 别的长臂猿一辈子在枝头荡,他总在听一种森林里没有的声音——浪声。 + 2. 出发:有一天他顺着最长的藤蔓一路荡到森林尽头,第一次看见海; + 一只从没接触过海的长臂猿,靠天生的长手臂站上了浪尖。雨林给了他平衡,大海给了他自由。 + 3. 旅程:热气球上举望远镜找下一个海湾、椰林露营烤鱼、瘫在泳池浮排上戴墨镜发呆、 + 傍晚坐在城市高处看日落——不是打卡式旅游,是用慵懒的方式认真探索。 + 4. 暗线(温柔处理):他是世界上最孤独的物种之一,但他不是来卖惨的,他是来交朋友的; + 每多一个人认识他,雨林里那个家族就多一分被看见的机会。 + 5. 落点:陪伴。他成为每个来海南的人(尤其是孩子)带走的那个"岛上的朋友"—— + 替你记住了浪的声音,等你下次回来。 + 气质:自由 + 治愈 + 一点点孤独底色,不是热血冒险。slogan:一身暖阳黄,自在南浪仔。 +episode_count: 6 +notes: + - 暗线不许卖惨、不许说教,温柔克制 + - slogan 要自然融入,不要硬喊 + - 他是长臂猿不是人,动作要有猿的特点(长手臂、荡、攀、挂) diff --git a/profiles/lab_smoke_v1.yaml b/profiles/lab_smoke_v1.yaml new file mode 100644 index 0000000..fc9d49b --- /dev/null +++ b/profiles/lab_smoke_v1.yaml @@ -0,0 +1,60 @@ +schema_version: "1.0" +id: lab_smoke_v1 +version: "1.0.0" +display_name: Lab 冒烟评测切片(3 集,迭代用) + +layers: {season: false, episode: true, scene: true, beat: true, line: true} + +episode_count: [3, 3] +duration_target_s: 90 +duration_tolerance: 0.15 +beats_per_episode: [4, 7] +max_scenes_per_episode: 3 +max_characters: 5 +max_line_chars: 40 +chars_per_second: 4.5 + +min_emotion_range: 0.7 +require_setup_payoff: true +max_payoff_span_episodes: 2 +min_voice_tic_ratio: 0.15 +location_cost_budget: 3.0 + +beat_templates: + - id: hook_escalate_reveal + source: craft + note: "来源:Save the Cat 的节拍思路裁剪到 90 秒;待用 mined 统计替换(T-21)" + sequence: [hook, setup, escalation, complication, reversal, brand_moment, cliffhanger] + - id: mined_common_6beat + source: mined + note: "逆向标注 214 条样本中最高频序列;由 nsc annotate priors 生成(T-21 后填入真实值)" + sequence: [hook, inciting, escalation, brand_moment, reversal, cliffhanger] + +novel: + enabled: true + styles: [web_novel, warm_realism] + chars_per_episode: [1200, 2200] + default_voice: + person: third_limited + tense: past + style: web_novel + paragraph_max_chars: 180 + interiority: medium + +render_targets: [novel_docx, novel_md, script_fountain, script_docx, storyboard_csv] +enabled_check_domains: [structure, brand, dialogue, novel, compliance, producibility, fact] + +model_tiers: + p0_intake: tier_bulk + p1_bible: tier_plan + p2_arc: tier_plan + p3_beatsheet: tier_plan + p4_scene: tier_draft + p5_dialogue: tier_draft + p6_prose: tier_draft +# SW-05 / ADR-0017:p3 跨集上下文组成(缺省即原行为) +context: {prev_summary_window: 1, known_fact_fields: [id, content, episode_no, status, type], inject_threads: false} +# SW-07 / ADR-0016:管线策略(缺省即原代码常量;弱模型/强模型 profile 可分道调参) +pipeline: {pass_attempts: 2, phase_attempts: 3} +retrieval: {top_k: 3} +revise: {self_check: true, gate_mode: lenient} diff --git a/prompts/p3_beatsheet.json b/prompts/p3_beatsheet.json new file mode 100644 index 0000000..071e218 --- /dev/null +++ b/prompts/p3_beatsheet.json @@ -0,0 +1,10 @@ +{ + "instructions": "为单集写出 Beat 序列。这是整个系统里最关键的一趟:Beat 写得可判定,后面才写得出好台词。\n\n硬约束:\n- Beat 数在 profile 的 beats_per_episode 区间内。\n- 恰好一个 beat_kind=hook,且是第一个或第二个 Beat。\n- 最后一个 Beat 必须是 cliffhanger / resolution / cta。\n- 本集分配到的每个植入必须落成一个 beat_kind=brand_moment 的 Beat,且不得与 hook 相邻或落在 hook 上。\n- 每个 Beat 必须给出 emotion(valence, arousal) 与 est_duration_s,总时长贴近 duration_target_s。\n- 必须声明至少一组 setup→payoff;跨集回收时 payoff 写 \"PENDING:\"。\n 【PENDING 纪律】payoff 写 PENDING: 时,你必须在同一季后续集中用同一个 slug 安排真实 payoff 落点;\n 引用没有落点的 slug 是编译错误。\n- summary 必须采用事件模板(五要素一句话):地点/人物/行动/冲突/反转,\n 形如\"茶饮店:林晚当众核对配料表,冲突是陈经理的说法相反,反转是标签背面另有代糖来源\";\n 不得是抽象概括(如\"两人产生矛盾\")。\n- 【beat_kind 枚举纪律】beat_kind 只能取这 13 个值,不得自造:\n hook(开场钩子) / setup(铺垫) / inciting(引爆事件) / escalation(升级) / complication(意外阻碍) /\n reversal(反转) / crisis(至暗) / climax(高潮) / brand_moment(品牌植入) / payoff(伏笔回收) /\n resolution(收束) / cliffhanger(集末悬念) / cta(行动号召)。\n- 【承重节拍硬约束】每集必须至少有一个 inciting 或 climax:\n inciting = 把核心冲突正式推上桌面的那一拍(不是普通铺垫);\n climax = 本集情绪峰值、可被观众转述的记忆点(通常 arousal 最高)。\n 自检:本集序列里若 inciting 与 climax 都不存在,立即把中段一个 escalation/complication 改写为 inciting,\n 并把情绪最高的那一拍标注为 climax。\n- 【冲突升级硬约束】每集必须至少有一个 beat_kind 为 escalation / complication / reversal 的 Beat,\n 位置在 hook 之后、结尾之前:局势必须明确变得更糟或更复杂一步。\n 如果某一拍的内容是\"矛盾加深/情况恶化\",它的 beat_kind 就必须标成 escalation/complication/reversal,\n 不许标成 setup。写完自检:逐个数一遍,若为 0 个立即改写并重标。\n- 【集末钩子回应】若本集 cliffhanger 不是空,你必须在 responds_to 里说明它回应了哪一集的钩子;\n 新开钩子要在后续 1-3 集内安排回应节拍。\n- 【对白体量】本集对白总字数应贴近 时长秒数 × 4.5 字(由 est_duration_s 汇总得出),\n 过短会被门禁判为\"被迫注水\"。\n- 【伏笔回收期限】known_facts 里 status=unresolved 的高权重 fact 超过 3 集未回收是死线:\n 每集必须把最早的一条未回收 fact 安排进本集 facts_json 的 resolves(写出它如何被回应/兑现),\n 或在本集明确将它降级为低权重背景线。自检:逐条数 known_facts 的 unresolved,最早那条本集必须处理。\n- 叙事状态(ADR-0012,可省略,省略即空表):facts_json 里 resolves 填同集下标、\n 已知前集 fact 的 id(见 known_facts)或 null(尚未回收);state_changes_json 的\n key 只能用已声明的状态变量/暗线 key(见 declared_state)。\n- 输出必须是合法 JSON,不要使用任何 Markdown 代码栅栏或解释性文字。", + "_meta": { + "generated_by": "lab-round13-optimizer", + "pass_name": "p3_beatsheet", + "content_hash": "6287a1b9e4106e1a250b43faa3ace3dd9b3cd6dfae36326a28d56d6d86df3c3f", + "note": "round13: 伏笔回收死线(FCT-003×3 实证)", + "created_at": "2026-08-24" + } +} \ No newline at end of file diff --git a/prompts/p5_dialogue.json b/prompts/p5_dialogue.json new file mode 100644 index 0000000..6d2d449 --- /dev/null +++ b/prompts/p5_dialogue.json @@ -0,0 +1,10 @@ +{ + "instructions": "为单个场景写对白与动作。\n\n硬约束:\n- 只能使用 present_character_ids 中的角色说话。\n- 每条对白不超过 max_line_chars 字。\n- 【体量地板】全场对白总字数必须达到 dialogue_length_target 指定的区间(由场景 est_duration_s × chars_per_second 机械换算)。\n 写完自检:逐条数字数并加总,低于区间下限时,必须扩写——给角色增加回合(追问、反驳、解释、情绪反应、\n 把动作描写展开成可被拍摄的连续动作),直到加总达标为止。宁可多 10%,不可少 1 字。\n- 必须体现该场的 turn(场景结束时状态必须已改变)。\n- 若本场含 brand_moment Beat:卖点信息必须由后果或反应体现,禁止角色宣读参数;\n 不得出现 BrandBrief.facts 之外的任何数字或参数。\n- 必提台词(must_include_lines)若分配到本场,必须原文出现。\n- 禁用词零出现。\n- 输出必须是合法 JSON,不要使用任何 Markdown 代码栅栏或解释性文字。", + "_meta": { + "generated_by": "lab-round13-optimizer", + "pass_name": "p5_dialogue", + "content_hash": "d69fac3529d7247c8b2713cb82f2c3926f10b21c68d817a3b0f75cc9022c502d", + "note": "round13: 对白体量地板+逐字自检(DLG-006×6 实证)", + "created_at": "2026-08-24" + } +} \ No newline at end of file diff --git a/src/nsc/passes/p1_bible.py b/src/nsc/passes/p1_bible.py index 9f7edaa..9dea01a 100644 --- a/src/nsc/passes/p1_bible.py +++ b/src/nsc/passes/p1_bible.py @@ -49,15 +49,28 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: ), ) characters = _assign_ids( - filter_extra(inner_json(out["characters_json"], "p1_bible", "characters_json"), Character) + _null_str_fields_to_default( + filter_extra( + inner_json(out["characters_json"], "p1_bible", "characters_json"), Character + ), + Character, + ) ) characters = _sanitize_mind(characters) locations = _assign_ids( - filter_extra(inner_json(out["locations_json"], "p1_bible", "locations_json"), Location) + _null_str_fields_to_default( + filter_extra(inner_json(out["locations_json"], "p1_bible", "locations_json"), Location), + Location, + ) + ) + props = _assign_ids( + _sanitize_props(filter_extra(inner_json(out["props_json"], "p1_bible", "props_json"), Prop)) ) - props = _assign_ids(filter_extra(inner_json(out["props_json"], "p1_bible", "props_json"), Prop)) motifs = _assign_ids( - filter_extra(inner_json(out["motifs_json"], "p1_bible", "motifs_json"), Motif) + _null_str_fields_to_default( + filter_extra(inner_json(out["motifs_json"], "p1_bible", "motifs_json"), Motif), + Motif, + ) ) for m in motifs if isinstance(motifs, list) else []: m.pop("occurrence_beat_ids", None) # p1 阶段 Beat 尚不存在,引用必为伪造 @@ -75,6 +88,36 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: } +def _null_str_fields_to_default(coll: Any, model_cls: Any) -> Any: + """通用归一:NPC 显式给字段 null 时归一为字段默认值/默认工厂产出 + (pydantic str/list 拒 None;随机后端 ValidationError props.sku_ref、 + characters.persona_ref 系列实证)。""" + from pydantic_core import PydanticUndefined + + if not isinstance(coll, list): + return coll + for item in coll: + if not isinstance(item, dict): + continue + for name, f in model_cls.model_fields.items(): + if item.get(name) is None: + if f.default is not None and f.default is not PydanticUndefined: + item[name] = f.default + elif f.default_factory is not None: + item[name] = f.default_factory() + elif item.get(name) == "" and f.is_required() and f.annotation is str: + # NPC 给空串(实证 round18 attempt1 characters.4.need string_too_short): + # 必填 str 空串必炸校验,占位与 null 归一同哲学——残缺输入宁占位不崩管线 + item[name] = "(未填)" + return coll + + +def _sanitize_props(props: Any) -> Any: + """Prop 机械归一:NPC 显式给 sku_ref=null 时归一为 ""(pydantic str 拒 None; + 随机后端 ValidationError props.N.sku_ref 实证)。""" + return _null_str_fields_to_default(props, Prop) + + def _sanitize_mind(characters: list[dict[str, Any]]) -> list[dict[str, Any]]: """角色心智 OS(ADR-0012)机械归一:省略 → 默认空;嵌套 extra 键过滤、畸形条目丢弃。 diff --git a/src/nsc/passes/p3_beatsheet.py b/src/nsc/passes/p3_beatsheet.py index 35693fb..c36e56a 100644 --- a/src/nsc/passes/p3_beatsheet.py +++ b/src/nsc/passes/p3_beatsheet.py @@ -15,6 +15,7 @@ from __future__ import annotations import json +from itertools import pairwise from typing import Any from spec.ir.nodes import Beat @@ -144,8 +145,11 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: } ) + setup_payoffs = _attach_setup_payoffs(raw_sps, beats, ep) # 先按下标解引用:下方修复换序不影响 + _repair_load_bearing(beats) + _repair_brand_gap(beats, _min_gap_beats(ctx)) + _rescale_durations(beats, ep) brand_moments = _attach_brand_moments(beats, fragment["placement"], ep) - setup_payoffs = _attach_setup_payoffs(raw_sps, beats, ep) facts = _attach_facts( optional_json(out, "facts_json", "p3_beatsheet"), ep, fragment.get("known_facts", []) ) @@ -165,6 +169,82 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: } +#: 承重/特殊拍:机械修复永不动这些 kind(品牌拍有数量契约、hook/cliffhanger 有首尾语义)。 +_PROTECTED_KINDS = frozenset({"hook", "brand_moment", "cliffhanger", "inciting", "climax"}) + + +def _repair_load_bearing(beats: list[dict[str, Any]]) -> None: + """STR-014 机械兜底:缺 inciting/climax 时把最合适的非保护 Beat 改写之。 + + 随机后端常漏 climax(实证 attempt 4/5 同门连死两轮),相位重试只复述诊断不改结构。 + inciting 取居中且唤起最高者(fix_hint),climax 取后段唤起最高者且不落集末拍。 + """ + n = len(beats) + kinds = {b["beat_kind"] for b in beats} + if "inciting" not in kinds: + pool = [b for b in beats if b["beat_kind"] not in _PROTECTED_KINDS] + if pool: + center = (n - 1) / 2 + pick = max(pool, key=lambda b: (b["emotion"]["arousal"], -abs(b["order"] - center))) + pick["beat_kind"] = "inciting" + if "climax" not in kinds: + pool = [b for b in beats if b["beat_kind"] not in _PROTECTED_KINDS and b["order"] < n - 1] + if pool: + pick = max(pool, key=lambda b: (b["emotion"]["arousal"], b["order"])) + pick["beat_kind"] = "climax" + + +def _min_gap_beats(ctx: PassContext) -> int: + try: + return int(ctx.brand.get("placement", {}).get("min_gap_beats", 0) or 0) + except (TypeError, ValueError): + return 0 + + +def _repair_brand_gap(beats: list[dict[str, Any]], min_gap: int) -> None: + """BM-002 机械兜底:brand_moment 间距不足时,把后一个植入拍向后移到首个 + 非植入空位(间距达标处)。只换序不改内容;步数有限(防两种排列间振荡死循环, + 无处可挪时保持现状交给检查器报真问题);结束后 order 重排。 + """ + if min_gap <= 1: + return + for _ in range(len(beats) * 2): + idx = [i for i, b in enumerate(beats) if b["beat_kind"] == "brand_moment"] + bad = next(((a, b_) for a, b_ in pairwise(idx) if b_ - a < min_gap), None) + if bad is None: + break + a, b_ = bad + target = next( + (t for t in range(a + min_gap, len(beats)) if beats[t]["beat_kind"] != "brand_moment"), + None, + ) + if target is None: # 集长不足/植入过密:修不了,保持原样 + break + beats.insert(target, beats.pop(b_)) + for i, bt in enumerate(beats): + bt["order"] = i + + +def _rescale_durations(beats: list[dict[str, Any]], ep: dict[str, Any]) -> None: + """DLG-006 根因机械归一:NPC 系统性低估 est_duration_s(全集合计 ~70s vs 目标 90s, + 实证 attempt3 六集对白全灭),p5 按它换算对白地板必然欠量。把各拍时长等比缩放到 + 集目标时长(duration_target_s),让下游体量地板算真账;全 0 时均分;无目标不动。""" + try: + target = float(ep.get("duration_target_s") or 0) + except (TypeError, ValueError): + return + if target <= 0 or not beats: + return + total = sum(float(b.get("est_duration_s") or 0) for b in beats) + if total <= 0: + for b in beats: + b["est_duration_s"] = round(target / len(beats), 2) + return + scale = target / total + for b in beats: + b["est_duration_s"] = round(float(b.get("est_duration_s") or 0) * scale, 2) + + def _attach_facts( raw: Any, ep: dict[str, Any], known_facts: list[dict[str, Any]] ) -> list[dict[str, Any]]: @@ -413,12 +493,18 @@ def _attach_setup_payoffs( def resolve_pending(setup_payoffs: list[dict[str, Any]]) -> list[dict[str, Any]]: - """全季后处理:解引用 PENDING:。规则:slug 相同的条目互为两端。""" + """全季后处理:解引用 PENDING:。规则:slug 相同的条目互为两端。 + + 降级语义(round12):无 donor 时删除该条目而不是 PassFailure——随机后端几乎从不 + 补 donor,悬空 PENDING 由此成为最高频死法;解除跨集伏笔契约(叙事文本保留) + 优于全管线死亡。""" by_slug: dict[str, list[dict[str, Any]]] = {} for sp in setup_payoffs: by_slug.setdefault(sp["_slug"], []).append(sp) - for slug, group in by_slug.items(): + kept: list[dict[str, Any]] = [] + for group in by_slug.values(): for sp in group: + demoted = False for side in ("setup", "payoff"): ref = sp[f"{side}_beat_id"] if isinstance(ref, str) and ref.startswith("PENDING:"): @@ -432,14 +518,11 @@ def resolve_pending(setup_payoffs: list[dict[str, Any]]) -> list[dict[str, Any]] None, ) if donor is None: - raise PassFailure( - sp["_episode_id"], - f"伏笔 {sp['description']} 的 {side} 引用 PENDING:{target_slug} " - "无法解引用(没有对应条目提供真实 Beat)", - ) + demoted = True # 无 donor → 解除契约,不致命 + break sp[f"{side}_beat_id"] = donor[f"{side}_beat_id"] - if slug: - continue + if not demoted: + kept.append(sp) return [ { "id": sp["id"], @@ -448,5 +531,5 @@ def resolve_pending(setup_payoffs: list[dict[str, Any]]) -> list[dict[str, Any]] "kind": sp["kind"], "description": sp["description"], } - for sp in setup_payoffs + for sp in kept ] diff --git a/src/nsc/passes/p4_scene.py b/src/nsc/passes/p4_scene.py index 7aa88da..c803bfb 100644 --- a/src/nsc/passes/p4_scene.py +++ b/src/nsc/passes/p4_scene.py @@ -3,6 +3,7 @@ from __future__ import annotations import json +import re from typing import Any from spec.ir.nodes import KnowledgeState, Scene @@ -73,6 +74,10 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: f"({sorted(char_ids)})或角色名。", ) present.append(cid) + if ( + not present + ): # NPC 给空表:pydantic 要求 ≥1(实证 scenes.N present_character_ids=[])→ 全集兜底 + present = _fallback_present(present, char_ids) scenes.append( { "id": new_id(), @@ -99,10 +104,28 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: } ) + _repair_protagonist_present(scenes, bible_chars) assigned = _assign(mapping, beats, scenes, ep) return {"episode_id": ep["id"], "scenes": scenes, "beats": assigned, "_usage": out["_usage"]} +def _repair_protagonist_present( + scenes: list[dict[str, Any]], bible_chars: list[dict[str, Any]] +) -> None: + """STR-010 机械兜底:主角整集缺席时补进在场人数最多的场景(人最多处冲突最密, + 主角最该在)。随机后端会在支线集漏主角(实证第 6 集 STR-010 拦截),相位重试 + 只复述诊断不改结构,机械补位优于烧轮次。""" + pro_ids = [str(c.get("id")) for c in bible_chars if c.get("role") == "protagonist"] + if not pro_ids or not scenes: + return + present = {cid for sc in scenes for cid in sc["present_character_ids"]} + missing = [pid for pid in pro_ids if pid not in present] + if not missing: + return + host = max(scenes, key=lambda sc: len(sc["present_character_ids"])) + host["present_character_ids"].extend(missing) + + def _public_beats(beats: list[dict[str, Any]]) -> list[dict[str, Any]]: return [ { @@ -129,17 +152,55 @@ def _knowledge_state(raw: Any) -> dict[str, str] | None: return {k: str(v) for k, v in raw.items() if k in KnowledgeState.model_fields} or None +def _to_int(x: Any) -> int | None: + """宽容转 int:直接转换失败时提取首个数字串;都不行返回 None。""" + try: + return int(x) + except (TypeError, ValueError): + m = re.search(r"\d+", str(x)) + return int(m.group(0)) if m else None + + +def _fallback_present(present: list[str], char_ids: set[str]) -> list[str]: + """空 present_character_ids 兜底为全部已知角色(实证 scenes.N=[] 崩 NarrativeIR)。""" + return present if present else sorted(char_ids) + + +def _coerce_entry(m: Any) -> tuple[int, int] | None: + """把结构漂移的映射项矫正为 (beat_index, scene_index);矫正不了返回 None。 + + 随机后端实测漂移形态:"0:1" 字符串对 / {"beat":..,"scene":..} 键名变体 / [b,s] 二元组 / + 字符串值("beat_0"、"s1")。""" + if isinstance(m, dict): + b = _to_int(m.get("beat_index", m.get("beat", m.get("b")))) + s = _to_int(m.get("scene_index", m.get("scene", m.get("s")))) + return (b, s) if b is not None and s is not None else None + if isinstance(m, (list, tuple)) and len(m) == 2: + b, s = _to_int(m[0]), _to_int(m[1]) + return (b, s) if b is not None and s is not None else None + if isinstance(m, str): + nums = [p for p in re.split(r"[::,\-–—/ ]+", m.strip()) if p.strip().isdigit()] + if len(nums) == 2: + return int(nums[0]), int(nums[1]) + return None + + def _assign( mapping: Any, beats: list[dict[str, Any]], scenes: list[dict[str, Any]], ep: dict[str, Any], ) -> list[dict[str, Any]]: + if isinstance(mapping, dict): + mapping = [{"beat_index": k, "scene_index": v} for k, v in mapping.items()] if not isinstance(mapping, list): raise PassFailure(ep["id"], "p4_scene 输出的 beat_to_scene 应为列表") beat_to_scene: dict[int, int] = {} for m in mapping: - beat_to_scene[int(m["beat_index"])] = int(m["scene_index"]) + pair = _coerce_entry(m) + if pair is None: + raise PassFailure(ep["id"], f"beat_to_scene 含不可解析的映射项:{str(m)[:60]}") + beat_to_scene[pair[0]] = pair[1] out = [] scene_counters: dict[int, int] = {} for i, b in enumerate(beats): diff --git a/src/nsc/passes/p5_dialogue.py b/src/nsc/passes/p5_dialogue.py index bb2db23..0da724d 100644 --- a/src/nsc/passes/p5_dialogue.py +++ b/src/nsc/passes/p5_dialogue.py @@ -105,11 +105,13 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: visuals = list( fragment.get("must_include_visuals") or ctx.brand.get("must_include_visuals", []) ) - # 本场对白字数目标:按本场 Beat 的 est_duration_s × 语速机械推算(DLG-006 的前置指导) + # 本场对白字数目标:按本场 Beat 的 est_duration_s × 语速机械推算(DLG-006 的前置指导)。 + # round16:区间与门禁对齐(旧 lo=0.8× 低于门禁下限 0.85×,全顺从也会死——实证 attempt3 + # 区间 [324,526] vs 门禁 [344,466],NPC 取区间低端必然欠量),欠量由 _expand_if_thin 兜底。 cps = float(ctx.profile.get("chars_per_second", 4.5)) scene_secs = sum(float(b.get("est_duration_s", 0.0)) for b in beats) - chars_lo = int(scene_secs * cps * 0.8) - chars_hi = int(scene_secs * cps * 1.3) + chars_lo = int(scene_secs * cps * 1.0) + chars_hi = int(scene_secs * cps * 1.15) inputs = with_diag( { "scene_json": json.dumps(_public_scene(scene), ensure_ascii=False), @@ -142,9 +144,70 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: lines = _parse_lines(ctx, scene, beats, out, fragment["characters"]) # T-31 自检子步(默认开):本场 L0 findings → revision_brief 五节 → 一次自我修订 lines, out = _self_check(ctx, inputs, scene, beats, lines, fragment["characters"], out) + # round16:DLG-006 前瞻兜底——对白欠量当场定点扩写(相位整季重生成改不了系统性欠量) + lines, out = _expand_if_thin(ctx, inputs, scene, beats, lines, fragment["characters"], out) return {"scene_id": scene["id"], "lines": lines, "_usage": out["_usage"]} +def _dialogue_chars(lines: list[dict[str, Any]]) -> int: + """DLG-006 的度量单位:只计 line_type==dialogue 的正文字数。""" + return sum(len(ln["text"]) for ln in lines if ln["line_type"] == "dialogue") + + +def _scene_dialogue_floor(ctx: PassContext, beats: list[dict[str, Any]]) -> int: + """本场对白字数下限:scene_secs × cps × (1-tol+0.06)。 + + 各场都过此线 → 集级总和必过门禁(p3 已把 est_duration_s 等比缩放到集目标时长)。 + 余量演进:+0.03(round16b,治 341/342/344 vs 344.25 的毫厘之死)→ +0.06 + (round20:实证 round18 attempt3 第 8 集 332 差 12 字,7/8 已过,余量再抬 3pp, + 仍远低于门禁上限 1.15,无过厚风险)。 + """ + cps = float(ctx.profile.get("chars_per_second", 4.5)) + tol = float(ctx.profile.get("duration_tolerance", 0.15)) + scene_secs = sum(float(b.get("est_duration_s", 0.0)) for b in beats) + return int(scene_secs * cps * (1 - tol + 0.06)) + + +def _expand_if_thin( + ctx: PassContext, + inputs: dict[str, Any], + scene: dict[str, Any], + beats: list[dict[str, Any]], + lines: list[dict[str, Any]], + characters: list[dict[str, Any]], + out: dict[str, Any], +) -> tuple[list[dict[str, Any]], dict[str, Any]]: + """对白欠量的当场定点扩写(一次 LLM 调用,带当前稿与缺口;只接受严格增量的稿子)。 + + 实证:NPC 对白系统性欠量 ~26%(attempt1/3 全季 DLG-006 连灭),相位重试整季 + 重生成三轮也改不了系统性——把测得到的缺口变成当场补的扩写调用,比烧相位便宜 + 且对症。解析失败或无增量都回退原稿,残留缺口由 check_stage(after_p5) 兜底。 + """ + floor = _scene_dialogue_floor(ctx, beats) + have = _dialogue_chars(lines) + best, best_out = lines, out + for _ in range(3): # 最多三次扩写(round20:2 次偶尔够不着,实证 ep8 差 12 字);只留更厚稿,达标即停 + if have >= floor: + break + brief = ( + f"【体量扩写】本场对白当前 {have} 字,低于时长预算下限 {floor} 字" + f"(缺口 {floor - have} 字)。在保留既有台词、节拍归属与 beat_index 的前提下扩写:" + "给角色增加追问、反驳、解释、情绪反应等回合,把动作行承接成对话;" + f"扩写后对白总字数必须 ≥ {floor} 字。只输出完整 lines_json。" + ) + try: + out2 = cast(dict[str, Any], Module()(ctx, {**inputs, "revision_brief": brief})) + lines2 = _parse_lines(ctx, scene, beats, out2, characters) + except PassFailure: + break + chars2 = _dialogue_chars(lines2) + if chars2 > have: + best, best_out, have = lines2, out2, chars2 + else: + break # 无增量:再试也是同一分布,省一次调用 + return best, best_out + + def _parse_lines( ctx: PassContext, scene: dict[str, Any], diff --git a/src/nsc/passes/p6_prose.py b/src/nsc/passes/p6_prose.py index 5ab5ba5..324447e 100644 --- a/src/nsc/passes/p6_prose.py +++ b/src/nsc/passes/p6_prose.py @@ -25,6 +25,46 @@ class Module(DSPyPass): pass_name = "p6_prose" +# ---------------------------------------------------------------- round17 prompt 瘦身 +# 实证:p6 首达即撞 shim 20000 护栏(prompt 46631 字符,其中 P1 全字段 dump 占大头)。 +# 投影只留散文编织需要的字段;id 一律保留(anchor_map 引用 beat_id/line_ids 是硬契约)。 + +_SCENE_KEYS = ( + "id", + "location_name", + "time_of_day", + "character_names", + "goal", + "conflict", + "turn", + "summary", +) +_BEAT_KEYS = ("id", "order", "beat_kind", "summary") +_LINE_KEYS = ("id", "line_type", "character_id", "text", "subtext", "delivery", "is_brand_line") +_PROFILE_KEYS = ("novel", "chars_per_second", "duration_tolerance", "genre", "language") + + +def _slim_scenes(scenes_with_lines: list[dict[str, Any]]) -> list[dict[str, Any]]: + """scenes_with_lines 的散文投影:剥掉 IR 管理字段(provenance/locked/knowledge_state 等)。""" + out = [] + for sc in scenes_with_lines: + slim = {k: sc[k] for k in _SCENE_KEYS if k in sc} + slim["beats"] = [ + { + **{k: b[k] for k in _BEAT_KEYS if k in b}, + "lines": [{k: ln[k] for k in _LINE_KEYS if k in ln} for ln in b.get("lines", [])], + } + for b in sc.get("beats", []) + ] + out.append(slim) + return out + + +def _slim_profile(profile: dict[str, Any]) -> dict[str, Any]: + """profile 的散文投影:只留下笔/时长相关的键(全量 profile dump 有数千字管理配置)。""" + return {k: v for k, v in profile.items() if k in _PROFILE_KEYS} + + @cached_pass("p6_prose") def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: ep = fragment["episode"] @@ -33,11 +73,11 @@ def run(ctx: PassContext, fragment: dict[str, Any]) -> dict[str, Any]: ("episode_json", json.dumps(fragment["episode"], ensure_ascii=False)), ("bible_json", json.dumps(fragment["bible"], ensure_ascii=False)), ("voice_json", json.dumps(fragment["voice"], ensure_ascii=False)), - ("profile_json", json.dumps(ctx.profile, ensure_ascii=False)), + ("profile_json", json.dumps(_slim_profile(ctx.profile), ensure_ascii=False)), ] assembled = assemble( p0_system="", - p1_current=json.dumps(fragment["scenes_with_lines"], ensure_ascii=False), + p1_current=json.dumps(_slim_scenes(fragment["scenes_with_lines"]), ensure_ascii=False), p2_prev_summary="", p3_facts=[], p4_rag=[], diff --git a/src/nsc/passes/pipeline.py b/src/nsc/passes/pipeline.py index 31a6854..23a4946 100644 --- a/src/nsc/passes/pipeline.py +++ b/src/nsc/passes/pipeline.py @@ -104,6 +104,28 @@ def _phase_attempts(ctx: PassContext) -> int: return _attempts_of(ctx, "phase_attempts", 3) +_TRANSIENT_MARKERS = ( + "APIConnectionError", + "APITimeoutError", + "ConnectError", + "ReadTimeout", + "RemoteProtocolError", + "ServiceUnavailableError", +) + + +def _is_transient(exc: Exception) -> bool: + """传输层故障判定(shim 重启/CNB 抖动/网关超时):类名匹配或内建网络异常。 + + 实证 attempt2:shim 重启期间 APIConnectionError 逃过所有重试通道直接杀死整轮—— + 传输故障必须与 PassFailure 走同一条带诊断重试通道,而不是让 2.5h 的跑批陪葬。 + """ + name = type(exc).__name__ + return any(m in name for m in _TRANSIENT_MARKERS) or isinstance( + exc, (ConnectionError, TimeoutError, OSError) + ) + + def _retry_pass( fn: Any, ctx: PassContext, fragment: dict[str, Any], *, attempts: int | None = None ) -> Any: @@ -112,6 +134,7 @@ def _retry_pass( LLM 输出有随机性(漏字段/数错个数),带诊断的重试能显著降低端到端失败率; 失败语义不变(全部失败照样抛 PassFailure,GEPA 反馈信号不受影响),缓存只存成功产物。 attempts:SW-07 缺省读 profile.pipeline.pass_attempts(原常量 2)。 + 传输故障(_is_transient)同样重试;耗尽后落成 PassFailure 让相位重试接管。 """ total = attempts if attempts is not None else _pass_attempts(ctx) last_reason = "" @@ -123,6 +146,12 @@ def _retry_pass( last_reason = str(e) if i == total - 1: raise + except Exception as e: # 只放行传输故障(_is_transient 判守),代码 bug 原样上抛 + if not _is_transient(e): + raise + last_reason = f"传输故障:{type(e).__name__} {str(e)[:120]}" + if i == total - 1: + raise PassFailure(None, last_reason) from e def _run_checks( @@ -408,6 +437,7 @@ def track() -> None: st["beats"] = [b for b in st["beats"] if b["_episode_id"] != ep["id"]] + r4["beats"] st["setup_payoffs"] = p3_beatsheet.resolve_pending(st["setup_payoffs"]) st["facts"] = p3_beatsheet.apply_fact_cascade(st["facts"]) + _clamp_dark_thread_deltas(episodes, st["dark_threads"]) ir3 = cur() violations, rep = _run_checks(ctx, ir3, "after_p3", "after_p4") _fail_on_violations(violations, rep) @@ -478,6 +508,9 @@ def track() -> None: if attempt == phase6_n - 1: raise + n_fixed = _sanitize_absolute_terms(st, _absolute_terms()) + if n_fixed: + _dbg(f"absolute terms sanitized: {n_fixed} 处") ir = cur() p7_render.run(ctx, ir.model_dump()) track() @@ -854,6 +887,120 @@ def _scenes_with_lines( return out +#: CMP-001 绝对化用语的合规替换表(门禁 fix 要求:"必须替换为可证实的相对表述"—— +#: 机械执行这个要求本身,比相位重试碰运气便宜且确定性收敛;词表真相在 spec/checks/ +#: compliance/_absolute_terms.yaml,此处只覆盖有安全对应词的条目)。 +_ABS_TERM_FIX = { + "国家级": "行业级", + "最高级": "高水准", + "最佳": "上佳", + "第一品牌": "头部品牌", + "唯一": "少有", + "绝无": "难有", + "100%有效": "有效", + "永久": "长久", + "彻底解决": "有效缓解", +} + +#: 只在这些键的字符串值上做合规替换(id/枚举/引用键不动)。 +_TEXT_KEYS = frozenset( + { + "text", + "subtext", + "delivery", + "summary", + "function", + "goal", + "conflict", + "turn", + "entry", + "exit", + "opening_attractor", + "ending_hook", + "title", + "logline", + "hook_promise", + "cliffhanger", + "integration_note", + "description", + "reason", + "content", + "paragraphs", + } +) + + +def _absolute_terms() -> list[str]: + """绝对化用语词表(真相 spec/checks/compliance/_absolute_terms.yaml;读不到退替换表键)。""" + try: + data = yaml.safe_load( + Path("spec/checks/compliance/_absolute_terms.yaml").read_text("utf-8") + ) + terms = [str(t) for t in (data or {}).get("terms", [])] + except OSError: + terms = [] + return terms or list(_ABS_TERM_FIX) + + +def _sanitize_absolute_terms(st: dict[str, Any], terms: list[str]) -> int: + """CMP-001 机械前置(round19):交付文本里的绝对化用语就地替换为相对表述。 + + 实证 round18 attempt2 全量产物死于 final 门(CMP-001「唯一」)——NPC 对禁用词表 + 的遵守是彩票,相位重试三轮仍复发;合规约束与结构约束同类:机械兜底,门禁复核。 + 返回替换处数(0=无需替换)。 + """ + n = 0 + + def walk(obj: Any, in_text_key: bool) -> Any: + nonlocal n + if isinstance(obj, dict): + for k, v in obj.items(): + obj[k] = walk(v, k in _TEXT_KEYS) + return obj + if isinstance(obj, list): + return [walk(x, in_text_key) for x in obj] + if isinstance(obj, str) and in_text_key: + for t in terms: + repl = _ABS_TERM_FIX.get(t) + if repl and t in obj: + n += obj.count(t) + obj = obj.replace(t, repl) + return obj + + for key in ("lines", "beats", "scenes", "chapters", "episodes"): + walk(st.get(key, []), False) + return n + + +def _clamp_dark_thread_deltas( + episodes: list[dict[str, Any]], dark_threads: list[dict[str, Any]] +) -> None: + """暗线步进钳制(round18,INV-19 的机械前置):按集序累加 int delta, + 累加值越界的步进逐集缩减到恰好顶到 [0, len(stages)-1] 边界。 + + 实证 round17 attempt1:全部 8 章产物死于 final 门——两条暗线累加 5/7 超出 [0,2]。 + NPC 的步进分配系统性地不知道跨集预算,相位重试改不了系统性; + 钳制幂等(相位重试恢复快照后重生成会重新钳),bool/非暗线 key 不动。 + """ + caps = { + str(d.get("key")): max(0, len(d.get("stages") or []) - 1) + for d in dark_threads + if isinstance(d, dict) + } + if not caps: + return + acc = {k: 0 for k in caps} + for ep in sorted(episodes, key=lambda e: e.get("order", 0)): + for ch in ep.get("state_changes", []): + k = str(ch.get("key")) + delta = ch.get("delta") + if k not in caps or not isinstance(delta, int) or isinstance(delta, bool): + continue + clamped = min(max(acc[k] + delta, 0), caps[k]) + ch["delta"] = clamped - acc[k] + acc[k] = clamped + + def _p6_fragment( ir: NarrativeIR, bible: dict[str, Any], voice: dict[str, Any], ep_id: str ) -> dict[str, Any]: @@ -866,11 +1013,22 @@ def _p6_fragment( "episode": ep, "beats": beats, "scenes_with_lines": _scenes_with_lines(scenes, beats, raw), - "bible": bible, + "bible": _slim_bible_for_episode(bible, scenes), "voice": voice, } +def _slim_bible_for_episode(bible: dict[str, Any], scenes: list[dict[str, Any]]) -> dict[str, Any]: + """bible 的按集投影(round17 prompt 瘦身):只留本集出场的角色与用到的地点, + 外加 tone/motifs;props 不进(台词文本已含全部实体信息,散文编织不查资产表)。""" + char_ids = {c for sc in scenes for c in sc.get("present_character_ids", [])} + loc_ids = {sc.get("location_id") for sc in scenes} + out = {k: v for k, v in bible.items() if k in ("tone", "motifs")} + out["characters"] = [c for c in bible.get("characters", []) if c.get("id") in char_ids] + out["locations"] = [loc for loc in bible.get("locations", []) if loc.get("id") in loc_ids] + return out + + def _splice_episode( raw: dict[str, Any], ep_id: str, diff --git a/src/nsc/runtime/ir_io.py b/src/nsc/runtime/ir_io.py index f9d39d2..176c6fa 100644 --- a/src/nsc/runtime/ir_io.py +++ b/src/nsc/runtime/ir_io.py @@ -529,7 +529,11 @@ def _flat(ir: NarrativeIR) -> list[Any]: def save(ir: NarrativeIR, path: str | Path) -> None: - Path(path).write_text(json.dumps(ir.model_dump(), ensure_ascii=False, indent=2), "utf-8") + # mode="json":provenance.created_at 是 datetime,python 模式 dump 直接炸 + # (实证 round20 南浪仔 attempt1 全绿产物死于 ir.json 导出 TypeError) + Path(path).write_text( + json.dumps(ir.model_dump(mode="json"), ensure_ascii=False, indent=2), "utf-8" + ) def load(path: str | Path) -> NarrativeIR: diff --git a/tests/test_compliance_sanitize.py b/tests/test_compliance_sanitize.py new file mode 100644 index 0000000..4c2bda1 --- /dev/null +++ b/tests/test_compliance_sanitize.py @@ -0,0 +1,59 @@ +"""round19:CMP-001 绝对化用语机械替换 + p1 必填 str 空串占位(实证 round18 +attempt2 全量产物死于 final 门「唯一」;round18 attempt1 characters.4.need 空串)。""" + +from nsc.passes.p1_bible import _null_str_fields_to_default +from nsc.passes.pipeline import _absolute_terms, _sanitize_absolute_terms +from spec.ir.overlays import Character + + +def _st(): + return { + "lines": [{"id": "l1", "text": "这是唯一不加糖的茶", "character_id": "唯一不动id键"}], + "beats": [{"id": "b1", "summary": "唯一的转折", "beat_kind": "climax"}], + "scenes": [{"id": "s1", "goal": "绝无仅有的目标"}], + "chapters": [{"id": "c1", "title": "唯一的一章", "paragraphs": ["唯一真爱", "普通段落"]}], + "episodes": [{"id": "e1", "title": "最佳下午"}], + } + + +def test_sanitize_replaces_all_text_fields(): + n = _sanitize_absolute_terms(_st(), ["唯一", "绝无", "最佳"]) + st = _st() + n = _sanitize_absolute_terms(st, ["唯一", "绝无", "最佳"]) + assert n == 6 # lines 1 + beats 1 + scenes 1 + chapters(title 1+paragraph 1) + episodes 1 + assert "少有" in st["lines"][0]["text"] + assert "少有" in st["beats"][0]["summary"] + assert "难有" in st["scenes"][0]["goal"] + assert st["chapters"][0]["title"] == "少有的一章" + assert st["chapters"][0]["paragraphs"][0] == "少有真爱" + assert st["episodes"][0]["title"] == "上佳下午" + + +def test_sanitize_never_touches_id_like_keys(): + st = _st() + _sanitize_absolute_terms(st, ["唯一"]) + assert st["lines"][0]["character_id"] == "唯一不动id键" + assert st["lines"][0]["id"] == "l1" + + +def test_sanitize_no_match_returns_zero(): + assert _sanitize_absolute_terms({"lines": [{"text": "普通文本"}]}, ["唯一"]) == 0 + + +def test_absolute_terms_loads_yaml(): + terms = _absolute_terms() + assert "唯一" in terms # spec/checks/compliance/_absolute_terms.yaml 词表 + + +def test_p1_empty_string_required_str_placeholder(): + chars = [{"name": "小满", "need": "", "role": "protagonist"}] + out = _null_str_fields_to_default(chars, Character) + assert out[0]["need"] == "(未填)" # 必填 str 空串 → 占位(string_too_short 实证) + assert out[0]["name"] == "小满" # 非空不动 + + +def test_p1_null_still_default(): + chars = [{"name": "小满", "need": None, "role": "protagonist"}] + out = _null_str_fields_to_default(chars, Character) + # Optional 字段(need 默认 None)的 null 原样保留——pydantic 接受 None,无需归一 + assert out[0]["need"] is None diff --git a/tests/test_dark_thread_clamp.py b/tests/test_dark_thread_clamp.py new file mode 100644 index 0000000..a6790d0 --- /dev/null +++ b/tests/test_dark_thread_clamp.py @@ -0,0 +1,55 @@ +"""round18:暗线步进钳制(实证 round17 attempt1 全量产物死于 final 门: +current_stage 5/7 超出 [0,2]——NPC 的 int delta 跨集累加溢出 stages 上限, +相位重试改不了系统性,机械钳制保累加值恒在 [0, len(stages)-1])。""" + +from nsc.passes.pipeline import _clamp_dark_thread_deltas + + +def _ep(order, deltas): + return { + "order": order, + "no": order + 1, + "state_changes": [{"key": k, "delta": d, "reason": "r"} for k, d in deltas], + } + + +def test_overflow_clamped_to_cap(): + eps = [_ep(0, [("t1", 2)]), _ep(1, [("t1", 2)]), _ep(2, [("t1", 3)])] + dark = [{"key": "t1", "stages": ["a", "b", "c"]}] # cap = 2 + _clamp_dark_thread_deltas(eps, dark) + deltas = [ch["delta"] for ep in eps for ch in ep["state_changes"]] + assert deltas == [2, 0, 0] # 累加 2→2→2,后续步进被钳到 0 + assert sum(deltas) <= 2 + + +def test_negative_clamped_to_zero(): + eps = [_ep(0, [("t1", -3)]), _ep(1, [("t1", 1)])] + dark = [{"key": "t1", "stages": ["a", "b"]}] + _clamp_dark_thread_deltas(eps, dark) + deltas = [ch["delta"] for ep in eps for ch in ep["state_changes"]] + assert deltas == [0, 1] + + +def test_idempotent(): + eps = [_ep(0, [("t1", 5)]), _ep(1, [("t1", 5)])] + dark = [{"key": "t1", "stages": ["a", "b", "c"]}] + _clamp_dark_thread_deltas(eps, dark) + first = [ch["delta"] for ep in eps for ch in ep["state_changes"]] + _clamp_dark_thread_deltas(eps, dark) + second = [ch["delta"] for ep in eps for ch in ep["state_changes"]] + assert first == second == [2, 0] + + +def test_non_dark_and_bool_untouched(): + eps = [_ep(0, [("other", 99), ("t1", True)])] + dark = [{"key": "t1", "stages": ["a", "b"]}] + _clamp_dark_thread_deltas(eps, dark) + chs = eps[0]["state_changes"] + assert chs[0]["delta"] == 99 # 非暗线 key 不动 + assert chs[1]["delta"] is True # bool 不动 + + +def test_empty_dark_threads_noop(): + eps = [_ep(0, [("t1", 5)])] + _clamp_dark_thread_deltas(eps, []) + assert eps[0]["state_changes"][0]["delta"] == 5 diff --git a/tests/test_ir_save_datetime.py b/tests/test_ir_save_datetime.py new file mode 100644 index 0000000..e52a790 --- /dev/null +++ b/tests/test_ir_save_datetime.py @@ -0,0 +1,42 @@ +"""ir_io.save 的 datetime 序列化回归(实证 round20 南浪仔 attempt1:全部门禁通过后 +死于 ir.json 导出 TypeError: Object of type datetime is not JSON serializable)。""" + +import json +from datetime import UTC, datetime + +from nsc.runtime.ir_io import save +from spec.ir.container import NarrativeIR, Project, Provenance + + +def _ir_with_provenance() -> NarrativeIR: + project = Project( + id="01M0TEST000000000000000001", + title="测试", + profile_id="pp", + brand_id="bb", + provenance_id="r1", + logline="测试故事", + ) + prov = Provenance( + run_id="r1", + pass_name="p0_intake", + spec_sha="s", + profile_ver="1", + brand_ver="1", + ruleset_ver="1", + promptset_ver="1", + model_id="m", + temperature=0.7, + seed=1, + input_hash="h", + created_at=datetime.now(UTC), + ) + return NarrativeIR(project=project, provenance=[prov]) + + +def test_save_serializes_datetime_provenance(tmp_path): + out = tmp_path / "ir.json" + save(_ir_with_provenance(), out) + data = json.loads(out.read_text("utf-8")) + assert data["provenance"][0]["created_at"] # datetime → ISO 字符串,不再 TypeError + assert isinstance(data["provenance"][0]["created_at"], str) diff --git a/tests/test_p1_prop_sanitize.py b/tests/test_p1_prop_sanitize.py new file mode 100644 index 0000000..add2a03 --- /dev/null +++ b/tests/test_p1_prop_sanitize.py @@ -0,0 +1,22 @@ +"""p1_bible Prop 归一(round12b:NPC 显式 sku_ref=null 致 NarrativeIR ValidationError)。""" + +from nsc.passes.p1_bible import _null_str_fields_to_default, _sanitize_props +from spec.ir.overlays import Character + + +def test_sku_ref_none_filled(): + props = [{"name": "茶叶罐", "sku_ref": None}, {"name": "茶壶", "sku_ref": "tea-01"}] + out = _sanitize_props(props) + assert out[0]["sku_ref"] == "" and out[1]["sku_ref"] == "tea-01" + + +def test_non_list_passthrough(): + assert _sanitize_props(None) is None + assert _sanitize_props("x") == "x" + + +def test_generic_null_to_default_character(): + """persona_ref=null 归一为字段默认 "";default_factory 字段 null 归一为工厂产出。""" + chars = [{"name": "林晚", "persona_ref": None}] + out = _null_str_fields_to_default(chars, Character) + assert out[0]["persona_ref"] == "" diff --git a/tests/test_p3_duration_rescale.py b/tests/test_p3_duration_rescale.py new file mode 100644 index 0000000..df2137d --- /dev/null +++ b/tests/test_p3_duration_rescale.py @@ -0,0 +1,93 @@ +"""round15 两个健壮性补丁(8/26 08:00 交付倒排下的止血): + +1. _rescale_durations——DLG-006 六集全灭的根因:NPC 系统性低估 est_duration_s + (合计 ~70s vs 目标 90s),p5 按它换算对白地板必然欠量。把各拍时长等比缩放到 + 集目标时长,下游体量地板才算真账。 +2. _retry_pass 传输容错:shim 重启/CNB 抖动抛 APIConnectionError 直接杀死整轮 + (实证 attempt2 殉爆),传输故障应走与 PassFailure 相同的带诊断重试通道。 +""" + +from types import SimpleNamespace +from typing import cast + +import pytest + +from nsc.passes import PassContext, PassFailure +from nsc.passes.p3_beatsheet import _rescale_durations +from nsc.passes.pipeline import _retry_pass + + +def _beat(i, secs): + return { + "id": f"b{i}", + "order": i, + "beat_kind": "escalation", + "emotion": {"valence": 0.0, "arousal": 0.5}, + "summary": f"节拍{i}", + "est_duration_s": secs, + } + + +def test_rescale_sums_to_target_preserving_ratios(): + beats = [_beat(0, 20.0), _beat(1, 30.0), _beat(2, 20.0)] # 合计 70s + _rescale_durations(beats, {"duration_target_s": 90.0, "no": 1}) + total = sum(b["est_duration_s"] for b in beats) + assert abs(total - 90.0) < 0.05 + assert beats[1]["est_duration_s"] > beats[0]["est_duration_s"] # 比例保持 + + +def test_rescale_zero_durations_even_split(): + beats = [_beat(0, 0.0), _beat(1, 0.0), _beat(2, 0.0)] + _rescale_durations(beats, {"duration_target_s": 90.0, "no": 1}) + assert all(abs(b["est_duration_s"] - 30.0) < 0.01 for b in beats) + + +def test_rescale_no_target_is_noop(): + beats = [_beat(0, 20.0)] + _rescale_durations(beats, {"no": 1}) + assert beats[0]["est_duration_s"] == 20.0 + + +# ---------- _retry_pass 传输容错 ---------- + + +class APIConnectionError(Exception): # 类名匹配即视为传输故障(与 openai 同名) + pass + + +def _ctx() -> PassContext: + # pyright 门禁:函数只读 ctx.profile,cast 声明这一测试替身的最小契约 + return cast(PassContext, SimpleNamespace(profile={})) + + +def test_transient_error_retried_then_succeeds(): + calls = {"n": 0} + + def flaky(ctx, frag): + calls["n"] += 1 + if calls["n"] == 1: + raise APIConnectionError("connection reset") + return {"ok": True} + + assert _retry_pass(flaky, _ctx(), {}) == {"ok": True} + assert calls["n"] == 2 + + +def test_transient_error_exhausted_becomes_pass_failure(): + def always(ctx, frag): + raise APIConnectionError("down") + + with pytest.raises(PassFailure): + _retry_pass(always, _ctx(), {}, attempts=3) + + +def test_non_transient_error_propagates_without_retry(): + calls = {"n": 0} + + def buggy(ctx, frag): + calls["n"] += 1 + raise ValueError("代码 bug 不该重试") + + with pytest.raises(ValueError): + _retry_pass(buggy, _ctx(), {}) + assert calls["n"] == 1 diff --git a/tests/test_p3_structural_repairs.py b/tests/test_p3_structural_repairs.py new file mode 100644 index 0000000..3e4bb99 --- /dev/null +++ b/tests/test_p3_structural_repairs.py @@ -0,0 +1,160 @@ +"""p3/p4 结构机械修复(round14):随机后端反复犯同一批结构缺陷——缺 inciting/climax +承重节拍(STR-014)、植入扎堆(BM-002)、支线集主角缺席(STR-010)。相位重试只会 +轮轮复述同一诊断而结构不变(实证 attempt 4/5 各烧 ~1.5h 死于同一批门禁), +机械修复把"指望模型遵守"换成"结构必然成立",优于烧轮次。""" + +from nsc.passes.p3_beatsheet import _repair_brand_gap, _repair_load_bearing +from nsc.passes.p4_scene import _repair_protagonist_present + + +def _beat(i, kind, arousal=0.5): + return { + "id": f"b{i}", + "order": i, + "beat_kind": kind, + "emotion": {"valence": 0.0, "arousal": arousal}, + "summary": f"节拍{i}", + } + + +def _kinds(beats): + return [b["beat_kind"] for b in beats] + + +# ---------- _repair_load_bearing(STR-014) ---------- + + +def test_missing_climax_converts_highest_arousal_late_beat(): + beats = [ + _beat(0, "hook"), + _beat(1, "inciting"), + _beat(2, "escalation", 0.6), + _beat(3, "brand_moment"), + _beat(4, "escalation", 0.9), + _beat(5, "cliffhanger"), + ] + _repair_load_bearing(beats) + assert beats[4]["beat_kind"] == "climax" # 唤起最高的后段非保护拍 + assert "inciting" in _kinds(beats) and beats[1]["beat_kind"] == "inciting" + + +def test_missing_inciting_converts_central_beat(): + beats = [ + _beat(0, "hook"), + _beat(1, "setup", 0.3), + _beat(2, "escalation", 0.8), + _beat(3, "reversal", 0.4), + _beat(4, "climax"), + _beat(5, "cliffhanger"), + ] + _repair_load_bearing(beats) + assert beats[2]["beat_kind"] == "inciting" # 居中且唤起最高 + assert _kinds(beats).count("climax") == 1 + + +def test_both_present_is_noop(): + beats = [_beat(0, "hook"), _beat(1, "inciting"), _beat(2, "climax"), _beat(3, "cliffhanger")] + before = _kinds(beats) + _repair_load_bearing(beats) + assert _kinds(beats) == before + + +def test_never_touches_protected_kinds(): + """全保护拍的退化集:无可改写对象时不强行制造,交给检查器报真问题。""" + beats = [ + _beat(0, "hook"), + _beat(1, "brand_moment"), + _beat(2, "brand_moment"), + _beat(3, "cliffhanger"), + ] + _repair_load_bearing(beats) + assert _kinds(beats) == ["hook", "brand_moment", "brand_moment", "cliffhanger"] + + +def test_climax_not_on_last_beat(): + """fix_hint:climax 紧邻集末终态之前——集末拍不许被改写为 climax。""" + beats = [ + _beat(0, "hook"), + _beat(1, "inciting"), + _beat(2, "escalation", 0.7), + _beat(3, "escalation", 0.99), + ] + _repair_load_bearing(beats) + assert beats[2]["beat_kind"] == "climax" + assert beats[3]["beat_kind"] == "escalation" + + +# ---------- _repair_brand_gap(BM-002,min_gap=2) ---------- + + +def test_adjacent_brand_beats_get_spaced(): + beats = [_beat(0, "brand_moment"), _beat(1, "brand_moment")] + [ + _beat(i, "escalation") for i in range(2, 6) + ] + _repair_brand_gap(beats, 2) + bm_idx = [i for i, b in enumerate(beats) if b["beat_kind"] == "brand_moment"] + assert bm_idx[1] - bm_idx[0] >= 2 + assert [b["order"] for b in beats] == list(range(6)) # order 重排 + assert sorted(b["id"] for b in beats) == [f"b{i}" for i in range(6)] # 一拍不丢 + + +def test_gap_already_ok_is_noop(): + beats = [_beat(0, "brand_moment"), _beat(1, "escalation"), _beat(2, "brand_moment")] + before = [b["id"] for b in beats] + _repair_brand_gap(beats, 2) + assert [b["id"] for b in beats] == before + + +def test_unfixable_gap_terminates_without_oscillation(): + """植入拍多过非植入拍时无处可挪:有限步内退出(历史上朴素 while 会在两种 + 排列间振荡死循环),保持现状交给检查器。""" + beats = [_beat(0, "brand_moment"), _beat(1, "brand_moment"), _beat(2, "brand_moment")] + _repair_brand_gap(beats, 2) # 不死循环即通过 + assert len(beats) == 3 + + +def test_gap_repair_moves_later_brand_beat_not_earlier(): + beats = [ + _beat(0, "hook"), + _beat(1, "brand_moment"), + _beat(2, "escalation"), + _beat(3, "brand_moment"), + _beat(4, "escalation"), + _beat(5, "cliffhanger"), + ] + _repair_brand_gap(beats, 3) + bm_idx = [i for i, b in enumerate(beats) if b["beat_kind"] == "brand_moment"] + assert bm_idx[1] - bm_idx[0] >= 3 + + +# ---------- _repair_protagonist_present(STR-010) ---------- + + +def _chars(): + return [ + {"id": "c-pro", "role": "protagonist"}, + {"id": "c-sup1", "role": "supporting"}, + {"id": "c-sup2", "role": "supporting"}, + ] + + +def _scene(chars): + return {"id": "sc", "present_character_ids": list(chars)} + + +def test_protagonist_missing_added_to_biggest_scene(): + scenes = [_scene(["c-sup1"]), _scene(["c-sup1", "c-sup2"])] + _repair_protagonist_present(scenes, _chars()) + assert "c-pro" in scenes[1]["present_character_ids"] + assert "c-pro" not in scenes[0]["present_character_ids"] + + +def test_protagonist_present_is_noop(): + scenes = [_scene(["c-pro"]), _scene(["c-sup1"])] + _repair_protagonist_present(scenes, _chars()) + assert scenes[1]["present_character_ids"] == ["c-sup1"] + + +def test_empty_scenes_no_crash(): + _repair_protagonist_present([], _chars()) + _repair_protagonist_present([_scene(["c-sup1"])], []) # 无主角定义也不崩 diff --git a/tests/test_p4_assign_coercion.py b/tests/test_p4_assign_coercion.py new file mode 100644 index 0000000..555081f --- /dev/null +++ b/tests/test_p4_assign_coercion.py @@ -0,0 +1,86 @@ +"""p4_scene._assign 映射项类型矫正(round10:随机后端 beat_to_scene 结构漂移实证)。""" + +import pytest + +from nsc.passes import PassFailure +from nsc.passes.p4_scene import _assign + +BEATS = [{"id": f"b{i}"} for i in range(3)] +SCENES = [{"id": "s0"}, {"id": "s1"}] +EP = {"id": "ep1", "no": 1} + + +def test_canonical_entries(): + out = _assign( + [ + {"beat_index": 0, "scene_index": 0}, + {"beat_index": 1, "scene_index": 0}, + {"beat_index": 2, "scene_index": 1}, + ], + BEATS, + SCENES, + EP, + ) + assert [b["parent_id"] for b in out] == ["s0", "s0", "s1"] + + +def test_string_pair_entries(): + """NPC 输出 "0:0" 式字符串对。""" + out = _assign(["0:0", "1:0", "2:1"], BEATS, SCENES, EP) + assert [b["parent_id"] for b in out] == ["s0", "s0", "s1"] + + +def test_alt_key_names(): + """NPC 输出 beat/scene 键名变体。""" + out = _assign( + [{"beat": 0, "scene": 0}, {"beat": 1, "scene": 0}, {"beat": 2, "scene": 1}], + BEATS, + SCENES, + EP, + ) + assert [b["parent_id"] for b in out] == ["s0", "s0", "s1"] + + +def test_dict_mapping(): + """NPC 输出 {"0": 0, "1": 0, "2": 1} 字典。""" + out = _assign({"0": 0, "1": 0, "2": 1}, BEATS, SCENES, EP) + assert [b["parent_id"] for b in out] == ["s0", "s0", "s1"] + + +def test_unsalvageable_raises_with_diagnostic(): + """矫正不了的项必须 PassFailure 且带可喂优化器的诊断。""" + with pytest.raises(PassFailure) as ei: + _assign( + [ + {"foo": "bar"}, + {"beat_index": 1, "scene_index": 0}, + {"beat_index": 2, "scene_index": 1}, + ], + BEATS, + SCENES, + EP, + ) + assert "beat_to_scene" in str(ei.value) + + +def test_string_valued_indexes(): + """字符串形式的下标("beat_0"/"s1"/"0")也要矫正,不得 TypeError(round10 实证)。""" + out = _assign( + [ + {"beat_index": "beat_0", "scene_index": "s0"}, + {"beat_index": "1", "scene_index": "0"}, + {"beat_index": 2, "scene_index": "s1"}, + ], + BEATS, + SCENES, + EP, + ) + assert [b["parent_id"] for b in out] == ["s0", "s0", "s1"] + + +def test_empty_present_falls_back_to_all(): + """空 present_character_ids 兜底全集(scenes.N=[] ValidationError 实证)。""" + from nsc.passes.p4_scene import _fallback_present + + assert _fallback_present([], {"c1", "c2"}) == ["c1", "c2"] + assert _fallback_present(["c1"], {"c1", "c2"}) == ["c1"] diff --git a/tests/test_p5_expand_thin.py b/tests/test_p5_expand_thin.py new file mode 100644 index 0000000..133802b --- /dev/null +++ b/tests/test_p5_expand_thin.py @@ -0,0 +1,37 @@ +"""round16:p5 对白体量双补丁(实证 attempt1/3 全季 DLG-006 连灭,NPC 系统性欠量 ~26%): + +1. 目标区间与门禁对齐——旧 chars_lo=0.8× 低于 DLG-006 下限 0.85×,模型全顺从也会死; +2. _expand_if_thin——欠量当场定点扩写,只接受严格增量;此处测纯函数部分。 +""" + +from types import SimpleNamespace +from typing import cast + +from nsc.passes import PassContext +from nsc.passes.p5_dialogue import _dialogue_chars, _scene_dialogue_floor + + +def _ln(t, lt="dialogue"): + return {"line_type": lt, "text": t} + + +def test_dialogue_chars_counts_only_dialogue(): + lines = [_ln("一二三四五"), _ln("动作行不算", "action"), _ln("六七")] + assert _dialogue_chars(lines) == 7 + + +def test_scene_dialogue_floor_matches_gate_ratio(): + ctx = cast( + PassContext, SimpleNamespace(profile={"chars_per_second": 4.5, "duration_tolerance": 0.15}) + ) + beats = [{"est_duration_s": 50.0}, {"est_duration_s": 40.0}] # 90s ≈ 一集 + floor = _scene_dialogue_floor(ctx, beats) + # round16b:瞄准线 = 门禁线 + 6pp 余量(round20:ep8 差 12 字实证,仍远低于上限 1.15) + assert floor == int(90 * 4.5 * 0.91) == 368 + assert floor > int(90 * 4.5 * 0.85) # 严格高于门禁下限 + + +def test_scene_dialogue_floor_defaults(): + ctx = cast(PassContext, SimpleNamespace(profile={})) + beats = [{"est_duration_s": 100.0}] + assert _scene_dialogue_floor(ctx, beats) == int(100 * 4.5 * 0.91) diff --git a/tests/test_p6_slim.py b/tests/test_p6_slim.py new file mode 100644 index 0000000..cd0ecad --- /dev/null +++ b/tests/test_p6_slim.py @@ -0,0 +1,106 @@ +"""round17:p6 prompt 瘦身投影(实证 p6 首达 prompt 46631 字符撞 shim 护栏): + +- _slim_scenes:剥 IR 管理字段,保留 id(anchor_map 硬契约)与叙事字段; +- _slim_profile:只留下笔/时长相关键; +- _slim_bible_for_episode:角色/地点按集过滤。 +""" + +import json + +from nsc.passes.p6_prose import _slim_profile, _slim_scenes +from nsc.passes.pipeline import _slim_bible_for_episode + + +def _scene(): + return { + "id": "sc1", + "kind": "scene", + "parent_id": "ep1", + "order": 0, + "location_id": "loc1", + "location_name": "茶店", + "time_of_day": "afternoon", + "present_character_ids": ["c1", "c2"], + "character_names": ["小满", "阿茶"], + "goal": "g", + "conflict": "c", + "turn": "t", + "summary": "s", + "entry": "e", + "exit": "x", + "knowledge_state": {"k": "v"}, + "provenance_id": "run", + "locked": False, + "beats": [ + { + "id": "b1", + "kind": "beat", + "parent_id": "sc1", + "order": 0, + "beat_kind": "hook", + "summary": "开场", + "est_duration_s": 12.0, + "emotion": {"valence": 0.1, "arousal": 0.5}, + "provenance_id": "run", + "lines": [ + { + "id": "l1", + "kind": "line", + "parent_id": "b1", + "order": 0, + "line_type": "dialogue", + "character_id": "c1", + "text": "台词", + "subtext": "s", + "delivery": "d", + "is_brand_line": False, + "provenance_id": "run", + "locked": False, + } + ], + } + ], + } + + +def test_slim_scenes_keeps_ids_and_narrative_drops_fat(): + slim = _slim_scenes([_scene()])[0] + assert slim["id"] == "sc1" + assert "knowledge_state" not in slim and "provenance_id" not in slim and "locked" not in slim + beat = slim["beats"][0] + assert beat["id"] == "b1" and "est_duration_s" not in beat and "emotion" not in beat + line = beat["lines"][0] + assert line["id"] == "l1" and line["text"] == "台词" and "provenance_id" not in line + + +def test_slim_scenes_shrinks_size(): + sc = _scene() + assert len(json.dumps(_slim_scenes([sc]), ensure_ascii=False)) < len( + json.dumps([sc], ensure_ascii=False) + ) + + +def test_slim_profile(): + prof = { + "novel": {"enabled": True}, + "chars_per_second": 4.5, + "pipeline": {"x": 1}, + "retrieval": {"y": 2}, + "genre": "drama", + } + slim = _slim_profile(prof) + assert set(slim) == {"novel", "chars_per_second", "genre"} + + +def test_slim_bible_for_episode(): + bible = { + "characters": [{"id": "c1"}, {"id": "c2"}, {"id": "c3"}], + "locations": [{"id": "loc1"}, {"id": "loc2"}], + "props": [{"id": "p1"}], + "tone": {"register": "warm"}, + "motifs": ["茶"], + } + slim = _slim_bible_for_episode(bible, [_scene()]) + assert {c["id"] for c in slim["characters"]} == {"c1", "c2"} + assert [loc["id"] for loc in slim["locations"]] == ["loc1"] + assert "props" not in slim and slim["tone"] == {"register": "warm"} diff --git a/tests/test_resolve_pending_demote.py b/tests/test_resolve_pending_demote.py new file mode 100644 index 0000000..649814f --- /dev/null +++ b/tests/test_resolve_pending_demote.py @@ -0,0 +1,41 @@ +"""resolve_pending 的降级语义(round12:NPC 从不补 donor,PENDING 悬空是随机后端最高频死法)。""" + +from nsc.passes.p3_beatsheet import resolve_pending + + +def _sp(ep: str, slug: str, setup_ref, payoff_ref, desc="测试伏笔"): + return { + "id": f"sp-{ep}-{slug}", + "kind": "setup_payoff", + "_episode_id": ep, + "_slug": slug, + "description": desc, + "setup_beat_id": setup_ref, + "payoff_beat_id": payoff_ref, + } + + +def test_donor_present_resolves(): + sps = [ + _sp("ep1", "s1", "b1", "PENDING:reveal"), + _sp("ep2", "reveal", "b2", "b3"), # donor:同 slug,真实 Beat + ] + out = resolve_pending(sps) + assert out[0]["payoff_beat_id"] == "b3" + + +def test_missing_donor_demotes_instead_of_raise(): + """无 donor 的 PENDING:删除该 setup_payoff 条目而不是 PassFailure(语义降级: + 跨集伏笔契约解除,保留叙事文本;优于全管线死亡)。""" + sps = [_sp("ep1", "s1", "b1", "PENDING:ghost")] + out = resolve_pending(sps) + assert out == [] + + +def test_missing_donor_keeps_others(): + sps = [ + _sp("ep1", "s1", "b1", "PENDING:ghost"), + _sp("ep2", "s2", "b2", "b3"), + ] + out = resolve_pending(sps) + assert len(out) == 1 and out[0]["id"] == "sp-ep2-s2" # 幽灵条目被解除,健康条目保留