2026-08-28·by Sijie Wang#cybernetics#theory#control-methods

incentive-shaping

激励塑造

时机:在执行者内部。 改变的是执行者自身的倾向。

  • 计算机侧:RLHF、微调、LoRA
  • 社会侧:教育、文化、习惯

拆开看:7a(权重层)属于模型厂商,对应的是狭窄、重复出现的场景。7b(文本层)才是你的地盘——每一次锁拒绝都自动生成一条规则修复提案 → meta-gate(由人来批准)→ 写入 skill / CLAUDE.md,也就是把激励塑造编译进一种不贬值、可审计、跨模型的介质;这份账本本身就已经是数据集(参见 Hermes 的自我演化,只差铸币权)。修改它必须通过一道 gate——"skill 库就是策略层的宪法"。

Up: control-methods

about this entry

One of sijie's wiki entries. The AI on this site is grounded in the same corpus and answers in sijie's voice, with citations back to entries like this one — answering costs sijie money, so it waits behind a code: enter an access code →