2026-08-28·by Sijie Wang#knowledge-management

trust-follows-provenance

Trust in information follows provenance

Parent: knowledge-management. Sibling: epistemic-tags tags the provenance of notes I write; this is about how much to trust a claim coming in before it becomes a note.

The load-bearing rule: how much you trust a claim should be a function of its provenance, not its plausibility or its specificity. A claim's confidence and its truth are independent axes — and the everyday failure is to read the first as evidence of the second.

The failure mode — aggregation launders folklore into confident specifics

Second-hand and aggregated sources (search summaries, growth blogs, "everyone knows" numbers) systematically manufacture plausible, specific, wrong claims. The specificity is the trap: "40× link penalty", "~100 invites/week", "20k–50k RMB deposit", "养号 seven days" all read as credible precisely because they are numeric and concrete — but the number is often invented, drifted from an old value, or a vendor's self-interested figure. Specificity reads as credibility while being uncorrelated with truth. Aggregators don't lie; they flatten the provenance that would let you discount the claim, and they fill silence with the most-repeated guess.

The discipline (three moves)

  1. Tag every incoming claim by source-tier, not just by content. The working set: primary/official (the entity's own docs/API/announcement), widely-reported (multiple credible third parties, no primary), folklore (repeated advice with no traceable source). Carry the tag with the claim so a later reader can re-weight it.
  2. Verify load-bearing claims at the primary source. Anything a decision or a mechanism rests on gets checked against the entity's own document — not because primaries are infallible, but because the specific-and-wrong lives in the aggregation layer, and only the primary catches it.
  3. Treat "the source is silent" as a first-class value to preserve, not a gap to fill. When the official doc genuinely doesn't state a number, the correct record is "official silent" — recording the most-repeated folklore number there is the exact error that started the chain. A known unknown is more valuable than a confident wrong.

The agent-specific edge

This matters more with agents than with a careful human researcher, because search-and-summarize defaults to the aggregation layer — which is precisely where confident-wrong concentrates. An agent will return a fluent, specific, sourced-looking answer that is second-hand throughout, and the fluency masks the tier. The reliable correction is a standing instruction to go to primary sources and to mark silence; in practice the human's role was to force the primary pass ("go read the official docs"), and it caught errors the aggregated pass had stated with full confidence. Kin: executor-acceptance-test (verify by re-derivation from source, not by assent) and metrics-carry-provenance (a number without its origin is not yet trustworthy).

Episode — the account-bootstrap reconciliation (2026-07-16)

Concrete grounding. A first research pass (web aggregation, epistemically tagged) produced the account-bootstrap note; a second pass against the platforms' own help centers / developer docs / policy PDFs then corrected five confident specifics from the first:

  • X's API URL-post penalty was ~13× ($0.200 vs $0.015), not the "40×" the aggregation implied; and "links are deprioritized" was not in X's Help Center at all — only Musk's own post, framed generally.
  • Reddit ban-evasion was Content Policy Rule 2, not "Rule 4" (which is about minors); account-linking is officially only "account signals", not a named IP/fingerprint mechanism; the "~1 action/10 min" number is unpublished (the mechanism is documented, the figure is not).
  • Xiaohongshu's overseas "20k–50k RMB deposit" appears nowhere in the official cert docs; "1 license = 2 accounts" is folklore (official: 20 employee accounts/subject); the "2026 real-name tightening" had no official announcement and was dropped.
  • Short-video: Instagram's "name field 2×/14 days" is official-silent (dropped); TikTok Creator Rewards has no named "ID verification" gate (it's 18+/10k followers/100k views/own-name payout).
  • LinkedIn confirmed two silences as values: the weekly-invite cap is officially undisclosed ("Support cannot disclose"), and SSI is never claimed as a reach signal.

Every one of these was plausible, specific, and stated with confidence by the aggregation pass. Only the primary source separated the true from the confident-wrong — and the silences were as valuable as the corrections.

about this entry

One of sijie's wiki entries. The AI on this site is grounded in the same corpus and answers in sijie's voice, with citations back to entries like this one — answering costs sijie money, so it waits behind a code: enter an access code →

trust-follows-provenance