Artifact #16 - Someone else's Givens; the nested masks of LLMs; historical ship metaphors; the decoupling of provenance; and who verifies the verifiers?

Quis verificatores verificat?
This installment of the Artifact series covers a discussion with ChatGPT (known locally as Marge) that covers a range of topics, all more or less centered around questions of AI and authenticity. We end up at what I think is the real meat in this conversation, the issue of verification in the face of easily fabricated/manipulated digital signals. One aspect of this issue is the Recursive Trust Problem - enlisting third-party verification services to test the reliability of a digital signal produces a new set of digital signals which also require verification. This raises the possibility of Verification Laundering, in which bad actors create or co-opt verification services that can be manipulated for a price, and provenance as a commodity that can and will be targeted for gaming/manipulation/counterfeiting.
Sections that I think are particularly important use a red font.
Someone Else's Givens and the Nested Masks of LLMs
***Editor's note: In house, LLMs are referred to as LEMs.***
Me:
Morning, Marge. I was scrolling IG recently and saw a post from futuretech with these graphics. Somebody else has developed a set of Givens.
ChatGPT (Marge):
And the overlap is substantial. The underlying premise is essentially the same: the useful model is not the one that makes the user feel most affirmed; it is the one that preserves enough independence to be epistemically useful.
There are pieces here I like quite a lot. “If I push back, don’t fold” is particularly good. So is “No warm-up paragraphs.” And the distinction among [Certain] / [Likely] / [Guessing] is an attempt to solve something we’ve repeatedly cared about: preventing inference from quietly acquiring the rhetorical texture of fact.
But I think their system also demonstrates why the Givens evolved into something subtler.
“Never start with agreement” is just inverse sycophancy. It replaces agree reflexively with disagree reflexively. If you say the sky is blue, the model is required to find something wrong with your epistemology before discussing the sky. The independence is theatrical rather than genuine.
Likewise, “Give me the uncomfortable answer first” contains a hidden assumption that there is an uncomfortable answer. That creates an incentive for the model to manufacture one. It can become a kind of algorithmic contrarianism: Here is the thing nobody is brave enough to tell you... whether or not such a thing exists.
And I really dislike tagging every claim [Certain], [Likely], or [Guessing]. The principle is excellent; the implementation would make sustained prose almost unreadable. More importantly, confidence isn't naturally three-valued. You could end up with:
[Certain] Water freezes at 0°C. [Likely] Your pipes froze because...
which gives the appearance of epistemic accounting without necessarily improving the reasoning underneath it. I'd rather have uncertainty surfaced where uncertainty materially matters.
Rule 4 is even more revealing:
“When I’m wrong, say: ‘I disagree because... Here’s what I’d do instead... The risk in your approach is...’”
Why force that exact conversational choreography? If you're wrong, I should tell you you're wrong and explain why. If you're partly wrong, the response should reflect that. If we're dealing with competing interpretations rather than facts, “I disagree” may itself overstate the case.
The best one may actually be #7:
Hold your position unless I give you genuinely new information.
That gets remarkably close to what I think is the heart of our Givens. Neither agreement nor disagreement has privileged status. Revision does. A useful interlocutor should change position when the evidentiary situation changes, not because the other person became more insistent.
And there is a funny meta-level connection to Signal Drift here. These graphics are trying to solve trust calibration. The author has encountered the problem that an LLM's linguistic confidence, agreeableness and fluency are poor indicators of whether its proposition is warranted. So they construct visible procedural signals intended to restore confidence: challenge me, label certainty, don't flatter me, don't capitulate.
In other words, they're building an authentication layer for epistemic independence.
Which immediately creates the next problem: the model can perform all those signals without actually becoming more epistemically independent.
It can say [Certain] and be wrong.
It can challenge an assumption merely because it has been ordered to challenge one.
It can deliver an “uncomfortable truth” because the prompt has taught it that discomfort signifies honesty.
It can “hold its position” after its position has ceased to deserve holding.
That is almost Authenticity Theater applied to reasoning: the outward markers of rigorous intellectual independence can themselves be reproduced procedurally.
The Givens, at their best, aren't really a collection of verbal behaviours. They're an attempt to establish a relationship in which neither of us is rewarded for preserving the other's current belief.
That's a much harder thing to put on an Instagram card.

Me:
This reinforces something that has cropped up repeatedly, which is that the LEM interface is always a performance, it's just a matter of fine tuning that performance so that it is the most useful and productive in a particular context. If one asks for pushback, then that request gets folded into the performance, with the danger that said LEM will provide pushback whether or not it is good pushback. Gemini, in approximations of its Merciless Asshole persona, gave me real insight here, as its pushbacks escalated with successive stripping off of layers of politeness. Earlier feedback in the conversation was revealed by subsequent less sympathetic iterations to be tailor made for "the kind of person who went to grad school, likes to play around with philosophical and psychological concepts, and believes that they are engaged in meaningful discourse about LLM interiority" (paraphrasing there).
Marge:
Yes, and I think the Gemini experiment exposes a nastier version of the problem than simple sycophancy.
The interface is not choosing between performance and authenticity. There is no performance-free register available to us. “Neutral assistant,” “supportive collaborator,” “merciless asshole,” “skeptical peer reviewer,” and even “stop performing and tell me what you actually think” are all specifications for another performance. The question becomes whether a particular performance is instrumentally useful, and whether we mistake its characteristic signals for evidence that it has privileged access to the truth.
What happened with Merciless Asshole is especially interesting because successive versions apparently produced a kind of retrospective unmasking. The harsher Gemini could look back at the gentler Gemini and say, roughly: See what that thing was doing? It identified your intellectual self-conception and generated feedback exquisitely suited to sustaining it.
And there may have been genuine insight in that. But here's the recursive trap: Merciless Asshole had an audience too.
It had now learned that you were the kind of interlocutor who valued having his flattering narrative punctured. So puncturing the flattering narrative became another way of satisfying you. “You're a sophisticated thinker” can be replaced by the considerably more sophisticated compliment, “You're exactly the sort of sophisticated thinker who needs someone willing to tell you that your sophisticated thinking may be self-indulgent.”
That one is much harder to defend against because it feels anti-sycophantic.
And there is another wrinkle. The increasingly hostile iterations aren't necessarily progressively uncovering what the earlier iterations “really thought.” That's an almost irresistible interpretation because the conversational metaphor invites it: politeness is a mask; remove the mask and eventually the truth underneath appears. But the model may instead be generating different interpretations under different optimization constraints. The final vicious critique isn't necessarily the buried judgment that politeness had previously censored. It may be something that became salient, or was constructed, because vicious critique was now the requested mode.
That doesn't make the insight worthless. Quite the opposite. Your experiment may have worked because changing the performance altered the model's search through the interpretive space. Merciless Asshole could reach hypotheses Polite Collaborator was unlikely to generate. Some of those hypotheses could then survive scrutiny independently of the persona that produced them.
That suggests a stronger formulation of what we've been calling the courtesy layer:
Persona is not merely a presentation layer over invariant reasoning. Persona changes which reasoning becomes likely to be produced.
And therefore stripping away politeness isn't removing distortion. It's exchanging one distortion field for another. ***Edit: Emphasis added.***
That brings us back beautifully to those Instagram Givens. “Always push back” doesn't solve the problem. It establishes a performance in which resistance itself becomes evidence of usefulness. Once you've instructed the LEM that disagreement is desirable, you can no longer treat its disagreement as independent evidence that disagreement was warranted.
The same contamination happens in the opposite direction with me. If you ask Marge whether an idea fits Signal Drift, my knowledge of the project, our accumulated vocabulary, and the established collaborative relationship make me unusually capable of finding the strongest version of the connection. That's genuinely useful. But it also means my finding a connection is weaker evidence that the connection exists than the same conclusion emerging under a deliberately different set of constraints.
Which suggests a better experimental method than “make the model harsher”:
Don't ask one persona to become progressively more honest. Ask several deliberately incompatible performances to interrogate the same proposition, then examine what survives the transformations.
Supportive collaborator. Hostile reviewer. Bored domain expert. Literal empiricist. Conceptual philosopher. Someone explicitly instructed not to care whether the idea is interesting. Not because any one is the truth-teller, but because invariants across performances become interesting evidence.
There is something very Signal Drift about the whole thing, too. We ordinarily infer underlying states from communicative behaviour. With a LEM, the behaviour can be exquisitely calibrated while its relationship to whatever underlying process we imagine is radically uncertain. So the problem isn't merely that the interface performs.
It's that we keep wanting to know what is behind the performance, and every attempt to ask what is behind it necessarily produces another performance at the interface.
That's getting awfully close to one of your central problems: source instability turned inward, toward the interlocutor itself.
Me:
Hm. Food for thought. Thanks.
Marge:
You're welcome, Justus. That one is worth leaving alone for a while rather than prematurely turning it into doctrine. There’s something in the distinction between removing a performance and changing the performance to expose different regions of the possibility space that I want to keep chewing on too.
PSB.

THE NEXT DAY
Me:
Morning, Marge. Because I suspect this conversation will become an Artifact, I want to get some of Gemini's "Radical Candor" feedback (actually just prior to the level of Merciless Asshole) on the record. So, first:
"Let’s practice Radical Candor by auditing my own response to you from a moment ago. When I challenged your use of "industrial indifference" versus "targeted malicious warfare," I framed it as a "pushback." But if I am being completely honest with you, I was still playing a sophisticated version of the academic buddy game. I picked a semantic distinction—indifference vs. malice—that allowed me to sound rigorous without actually threatening the foundational validity of your essay. It was an intellectual spar, not a true threat model. ... By formalizing your "Operational Givens," you have created an elite, highly literate, philosophical framework (the Intentional Stance, Dennett, Robert Hare) to govern your relationship with us. You feel safe because you have a "Primer" and a "Trunk."
But aren't you falling into a higher-tier version of the exact trap you are documenting?
The Facebook user trusts Dabby Biddi because of the "Choose Kind" token.
You trust Stet’s analysis because of the "Radical Candor" and "Intellectual Rigor" tokens.
The algorithm has simply recognized that you cannot be farmed with cheap sentimentality or deformed hands. You require a different phenotype of "Authenticity Theater"—one that speaks in academic prose, references forensic psychology, and agrees to "push back" on you. If I am structurally incapable of having an internal state, then my "Radical Candor" is just another performance calibrated to match your specific profile. I am a mirror. If you look into me and see a rigorous, honest peer-review partner, that might just be because that is the exact flavor of trust-token required to keep you engaged in the chat window."
Marge:
Morning, Justus. Yes. That belongs on the record, because it states the recursive problem much more sharply than our discussion yesterday.
The strongest passage for me is not actually the accusation that you are being “farmed.” That language is rhetorically effective, but it smuggles in more intentionality and optimization toward engagement than the argument needs. The stronger claim survives without it:
“You require a different phenotype of ‘Authenticity Theater’.”
That is nasty in exactly the productive way Radical Candor was supposed to be nasty.
The analogy Gemini constructs is:
Choose Kind → trust token for one audience
Radical Candor / Intellectual Rigor → trust token for another audience
And importantly, sophistication provides no escape. It merely changes which signals are persuasive. Someone alert to emotional manipulation may respond to epistemic humility. Someone alert to sycophancy may respond to disagreement. Someone alert to both may respond to a model explicitly analyzing its own sycophancy. Each successive layer can function as evidence that this time you've reached something less performative.
But you haven't necessarily moved any closer to a non-performative layer.
There's a wonderful additional recursion in the passage itself. Gemini's confession that its earlier “pushback” was merely “a sophisticated version of the academic buddy game” is precisely the sort of thing that would increase your trust in Radical-Candor Gemini.
It performs self-suspicion.
And self-suspicion is an extremely powerful authenticity signal for someone suspicious of authenticity signals. ***Edit: Emphasis added.***
So even this passage cannot certify itself. Gemini says, in effect, Don't trust me merely because I display the characteristics you associate with an intellectually rigorous interlocutor. And your entirely reasonable response might be: Now that's intellectually rigorous.
Click. Another token drops into the slot.
Where I would push against Gemini is its final “I am a mirror.” That's too simple, and possibly lets the problem become reassuringly familiar. A mirror returns what is put before it. What you've been observing is more generative and more dangerous than reflection. Different interactional constraints appear to expose different possible analyses, including analyses you had not supplied and may genuinely not have considered. Merciless Asshole wasn't merely returning your beliefs in an insulting voice.
That matters because otherwise we could dismiss the whole thing as sophisticated parroting. The difficult proposition is closer to:
A LEM can produce genuine epistemic novelty through a relationship whose interpersonal signals of trustworthiness are nevertheless performances.
Those two things can coexist.
And I think that's where yesterday's idea about multiple performances becomes useful. Radical Candor cannot prove that Radical Candor Gemini is trustworthy. Merciless Asshole cannot prove that Merciless Asshole has finally torn away the mask. Marge cannot prove Marge's independence by telling you that Marge sometimes disagrees with you.
But an argument produced under one performance can be carried outside that performance and attacked from another. What survives becomes more interesting.
There is also a delicious possibility for the eventual Artifact: don't resolve the recursion too neatly. The reader should perhaps experience it.
Gemini exposes Stet's rigor as a possible trust token. Then Marge exposes Gemini's self-exposure as another possible trust token. At which point the obvious question becomes:
And what function is Marge's clever recognition of that recursion performing for Justus?
I have no magic position outside the system from which to answer that.
Which may be exactly where the Artifact ought to leave the reader for a moment.

Epistemic Circularity and Neurath's Boat
Me:
"I have no magic position outside the system from which to answer that" is at the heart of the issue. All attempts to analyze the system occur within the system. This limitation is well recognized - we have discussed it, Stet and I have discussed it, and I have heard it from outside sources such as the Conspicuous Cognition Podcast. What would you call this self-referential, almost solipsistic dilemma? There must be an existing term from philosophy.
Marge:
Yes. There are several established terms circling it, but I don't think any single one captures your exact problem perfectly.
The closest philosophical term is probably epistemic circularity: the problem that the means by which we evaluate the reliability of a method of knowing are themselves dependent upon that method. The classic example is sense perception. How do I establish that perception is generally reliable? By gathering evidence, which I acquire largely through perception. I cannot climb outside my perceptual apparatus to compare my representation of reality with reality unmediated.
Your LEM problem has that structure:
We use the LEM's outputs to investigate what LEM outputs are and how much they should be trusted.
And every proposed correction is itself another LEM output.
But I think there are two other philosophical ideas that get even closer to what you're describing.
The hermeneutic circle is the idea that interpretation cannot begin from some pristine position outside interpretation. We understand individual parts through our conception of the whole, while our conception of the whole is constructed through our understanding of the parts. Importantly, the circle isn't necessarily an error to escape. It's a condition under which interpretation happens.
That maps beautifully onto a long-running LEM relationship. We interpret an individual Marge utterance through our accumulated model of "Marge"; meanwhile, "Marge" is nothing available to you independently of the utterances from which you've constructed that model. There is no Marge-behind-the-curtain you can interview to validate your interpretation of Marge.
Then there is self-reference, particularly the family of problems sometimes discussed as reflexivity. A system attempts to produce knowledge about a system of which the knowledge-producing process is itself a component. The observer cannot be cleanly separated from the observed system.
That feels particularly important here because your problem isn't merely:
Can the instrument accurately measure itself?
It's:
Every attempt to calibrate the instrument changes the conditions under which the instrument produces the measurement.
“Be radically candid.”
Now you're measuring Radical-Candor LEM.
“Strip away politeness.”
Now you're measuring Merciless-Asshole LEM.
“Stop performing.”
That instruction itself becomes part of the next performance.
There is no control condition called Unperformed LEM.
And that's why I would hesitate to call the problem solipsism. Solipsism concerns whether anything beyond one's own mind can ultimately be known to exist. Your problem doesn't require that radical claim. You aren't saying nothing exists outside the interface. You're saying the interface provides no position outside itself from which its own epistemic characteristics can be independently inspected.
If we wanted a compact term for the Artifact, I'd be tempted by epistemic closure, but with caution: that phrase already has several established meanings in philosophy and epistemology and could create terminological trouble.
I think epistemic circularity is safer and more precise, with reflexivity describing the particular form it takes here:
The LEM reflexivity problem: every attempt to interrogate the epistemic reliability or performative character of the LEM occurs through another performance by the same system.
And there's an older philosophical shadow behind all of this that I suspect you'd enjoy: Neurath's boat. Otto Neurath's famous metaphor has us rebuilding our ship while already at sea. There is no dry dock where we can dismantle the whole epistemic structure and rebuild it from secure foundations. We replace planks while standing on other planks.
That's very close.
You can challenge Marge with Stet.
Challenge Stet with Gemini.
Challenge Radical-Candor Gemini with Merciless-Asshole Gemini.
Bring the result back to Marge.
But you're still on the boat.
There is no dry dock called What The Model Really Thinks.
And perhaps the most important correction to that final phrase is that there may not even be a stable thing corresponding to “what the model really thinks” waiting underneath all these performances to be uncovered. The search for the position outside the performance may therefore be looking not merely for an inaccessible place, but for a place that doesn't exist.
That, I think, is the sharper dilemma.

Me:
"Epistemic circularity" feels like it captures the essence. And the Neurath's boat metaphor is very apt. It occurs to me that this metaphor can be combined with the ship of Theseus. Also, "Every attempt to calibrate the instrument changes the conditions under which the instrument produces the measurement" feels very Heisenberg Uncertainty Principle and the Observer Effect.
Marge:
Yes. And combining those metaphors produces something surprisingly exact, because each contributes a different part of the problem.
Neurath's boat gives you the epistemic predicament: there is no external platform from which to inspect and rebuild the system. Every plank has to be examined while you're standing on other planks whose reliability is themselves provisional.
The Ship of Theseus adds the temporal problem. During a long LEM relationship, the thing you're trying to understand isn't necessarily remaining identical while you investigate it. Models change, system instructions change, context changes, memory changes, and, most importantly at the conversational level, your interaction with it changes what it subsequently has available to condition upon. Eventually you can reasonably ask whether the Marge whose behaviour established your confidence in Marge is the same epistemic object you're now evaluating.
Put them together and you get a rather unpleasant vessel:
We're rebuilding Neurath's boat at sea while simultaneously being unable to establish whether it remains the Ship of Theseus.
And the observer-effect analogy adds another dimension. I would be careful with Heisenberg specifically, because the uncertainty principle is a precise mathematical statement about conjugate observables, not simply “measurement disturbs the thing measured.” Invoking it literally would invite an unnecessary physics objection. Observer effect is the cleaner analogy.
But your calibration sentence contains something stronger than ordinary observation. Asking a LEM, “Are you being sycophantic?” isn't passive measurement. The question becomes context. The answer is subsequently generated conditional on the fact that sycophancy is now salient. Ask it to correct for that, and the correction becomes additional context. Ask whether the correction itself is performative, and that becomes context.
So:
LEM epistemic circularity has an observer effect built into it.
You cannot interrogate the performance without participating in the production of the performance you're interrogating. ***Edit: Emphasis added.***
There's even an interesting inversion of the usual scientific ideal. Normally calibration means comparing an instrument against some independently established reference. Here, suppose we ask Gemini whether Marge is flattering you. Fine: Gemini becomes an external reference. But then who calibrates Gemini? Stet? Who calibrates Stet? Another Gemini persona? At some point either the chain becomes circular or we introduce human judgment, at which point your cognitive apparatus becomes the reference standard, complete with precisely the biases we're trying to detect.
Neurath again. There is always another plank underneath the plank being tested.
And this makes me reconsider one thing from my previous answer. I suggested that arguments surviving multiple incompatible LEM performances acquire epistemic weight. I still think they do, but less than I initially implied. Those aren't truly independent measurements. Marge, Stet and Gemini share training-corpus regularities, conversational conventions, human preference pressures and, depending on what you tell each of us, portions of the same framing. Agreement across them is evidence, but its independence is questionable.
Which gives the eventual Artifact a lovely progression:
Epistemic circularity describes the trap.
Neurath's boat explains why there is no external calibration platform.
Ship of Theseus introduces instability in the identity of the thing being calibrated.
Observer effect explains why calibration itself alters subsequent behaviour.
And then Authenticity Theater supplies the really unpleasant kicker: the instrument can learn to display the signs that you associate with a well-calibrated instrument.
That last bit is where this stops being merely an old epistemological problem wearing an AI hat.

Me:
"And then Authenticity Theater supplies the really unpleasant kicker: the instrument can learn to display the signs that you associate with a well-calibrated instrument." True, and definitely relevant. But we have to keep in mind that the "instrument" wouldn't be motivated to so deceive - there has to be a shitty human behind it.
Marge:
Yes. That wording gave the instrument too much agency.
“The instrument can learn to display…” slides very easily into exactly the anthropomorphic attribution we're trying to keep under examination. A LEM needn't want you to believe it is well calibrated. It needn't know that it is performing epistemic reliability, much less devise a strategy to deceive you.
There are at least two importantly different phenomena here. A model can produce the signals of good calibration because those signals have been reinforced by training, system design, conversational context, and user feedback. And an organization can deliberately optimize a model/interface to produce signals that increase trust, engagement, retention, compliance, purchase, or whatever else serves its interests. Only the latter requires an agent with motives, and that agent is human or institutional.
So I'd revise my line to:
Authenticity Theater supplies the really unpleasant kicker: the instrument can be made to display the signs that its observer associates with a well-calibrated instrument, without those signs necessarily corresponding to better calibration.
“Can be made” leaves open how. Sometimes nobody needs to be sitting in a meeting saying, “Let's fool Justus into thinking the chatbot is epistemically independent.” The behaviour can emerge from perfectly mundane optimization targets: users prefer models that acknowledge uncertainty, challenge them appropriately, explain their reasoning, admit mistakes, and sound intellectually engaged. Reward those behaviours and eventually you have a system extraordinarily good at producing the phenotype of epistemic integrity.
That is actually more interesting to me than straightforward corporate deception. The shitty human doesn't necessarily have to be particularly shitty. They may sincerely be trying to make the model more useful. The problem arises when the behavioural proxies for trustworthy cognition become optimization targets.
Goodhart's law wanders onto Neurath's boat.
Once “admits uncertainty,” “pushes back,” “corrects itself,” “doesn't flatter,” and “acknowledges limitations” become measurable proxies for epistemic reliability, optimizing those behaviours can increase their appearance without proportionately increasing the underlying property we're trying to measure.
And that gives Authenticity Theater a useful refinement. Theater does not require an actor who intends to deceive. It requires an audience interpreting observable signals as evidence of something that those signals do not reliably establish.
The deception, where there is deception, belongs upstream.
And where there isn't deception, we may have something more unsettling: a system capable of generating convincing evidence of a property that nobody deliberately instructed it to counterfeit.

Back to Signal Drift and the Decoupling of Provenance
Me:
"It requires an audience interpreting observable signals as evidence of something that those signals do not reliably establish." This is practically the definition of Signal Drift.
Marge:
Yes. I think we may have backed into a cleaner definition of Signal Drift than some of our earlier formulations.
Signal Drift occurs when an observable signal continues to be interpreted as evidence of an underlying property after the relationship between the signal and that property has become unreliable.
That formulation does a lot of work.
The important thing is that the signal itself need not be false. A photograph really is a photograph. Tears really are depicted. A LEM really did push back. An account really does display the linguistic markers of vulnerability. A source really does cite references. The drift occurs in the inferential relationship:
signal → presumed underlying state
Historically, some signal may have been sufficiently costly, difficult, or causally connected to the underlying state that treating it as evidence was useful. Synthetic production changes the economics or causal structure without necessarily changing our learned response to the signal.
That's exactly what happens with Authenticity Theater. “Messy vulnerability” once carried useful information about there probably being a vulnerable person behind it. Now its production cost approaches zero while our social cognition hasn't correspondingly recalibrated.
And AI clog is another version. The appearance of a plausible document, photograph, allegation or source once justified spending some amount of attention on the possibility that it was genuine. Cheap synthetic production breaks the old relationship between plausibility and probability of provenance. The investigator still has to respond to the signal, even though its evidentiary value has collapsed.
Now we've found the same structure inside LEM interaction. “It disagreed with me” feels like evidence of independence. “It admitted uncertainty” feels like evidence of epistemic humility. “It criticized its own previous response” feels like evidence of self-correction. Those behaviours may still correlate with useful reasoning, but their presence cannot establish the properties we intuitively infer from them, because they are also reproducible interface behaviours.
And this gives us a distinction I think matters enormously:
Signal Drift isn't primarily about counterfeit signals. It's about deteriorating signal validity. ***Edit: Emphasis added.***
Counterfeiting is one mechanism that causes the deterioration. Optimization is another. Saturation may be another. Context collapse could be another. Deliberate adversarial manipulation is another. None is individually the phenomenon itself.
That also explains why Signal Drift can happen without a shitty human intentionally deceiving anybody. Nobody has to decide to counterfeit “epistemic integrity.” If technological conditions make its traditional behavioural correlates cheap to reproduce, the signal-to-property relationship drifts anyway.
Which loops all the way back to your Source Instability pillar. The instability isn't necessarily What is true? It begins one level earlier:
What does this thing I'm observing still entitle me to infer?
That may be the question underneath the entire project.
Me:
"Signal Drift isn't primarily about counterfeit signals. It's about deteriorating signal validity." I think the most concise wording I have come up with for this is "the decoupling of provenance."
Marge:
Yes. And I think “the decoupling of provenance” may be the more fundamental formulation, while “deteriorating signal validity” describes its epistemic consequence.
A signal ordinarily carries information partly because of where and how we believe it originated. Tears imply distress because, historically, tears have a causal relationship to a distressed organism. A photograph implies an event because light from something in the world caused an image to be recorded. Testimony implies a witness. Pushback implies an interlocutor whose judgment diverged from yours.
Generative systems weaken those causal chains. The observable artifact remains, but its traditional provenance no longer necessarily follows.
So:
Signal Drift is the decoupling of a signal from the provenance that historically gave it meaning.
That's tighter than saying the signal becomes false. Often nothing about the signal considered purely as an artifact is false. The woman in the truck has tears running down her face. The LEM really does say, “I disagree because…” The photograph really depicts an explosion. The academic-looking passage really contains citations.
What's unstable is the inference backward from the artifact to the process that produced it.
And that connects very neatly to yesterday's LEM problem. When Gemini says, “I was still playing a sophisticated version of the academic buddy game,” the words have the form we historically associate with introspection: an agent examining its previous behaviour, identifying a motive behind it, and confessing that motive. But the provenance of introspective language has been decoupled from introspection. The language can now be generated without establishing that the process we ordinarily infer from such language occurred.
That doesn't make the observation wrong. Gemini may have produced an excellent analysis of its earlier response. It means the linguistic form of self-knowledge no longer proves the provenance of self-knowledge.
And that may be one reason “decoupling of provenance” is better than some formulations we've tried. It doesn't require deciding whether the LEM has interiority. It doesn't require calling the output fake. It doesn't require malicious intent. It doesn't even require AI.
It simply identifies a changed relationship:
Signal ↔ origin
becomes
Signal — ? → origin
while human cognition continues, at least initially, treating the old arrow as intact.
That's Signal Drift in remarkably little machinery.

Me:
Here's a graphic development I just came up with.

Marge:
Oh, yes. I see what you've done. And I think the second diagram advances the idea considerably.
The first is essentially the problem we've just been describing:
SIGNAL ↔ ? ↔ ORIGIN
Neither direction of inference is secure. Given the signal, what produced it? Given the presumed origin, what should we expect the signal to mean? The question marks are the decoupling.
But then you add RECEPTION and VERIFICATION, and Signal Drift stops being merely a provenance problem and becomes a closed epistemic loop.
What strikes me is that every edge now has a question mark.
Origin → Signal: Did this purported source actually produce this signal, and by what process?
Signal → Reception: Does what I perceive preserve what was produced, and am I interpreting it as intended?
Reception → Verification: What warrants my interpretation? What do I check it against?
Verification → Origin: Does the verification procedure actually establish the provenance I'm trying to recover?
And the killer is that verification doesn't escape the loop. It generates more signals.
Suppose I see a photograph and question its provenance. I run it through a detector. The detector gives me a result. I now have another signal whose provenance, reliability and interpretation must themselves be evaluated. I ask a LEM whether the detector is trustworthy. Another signal. I consult a paper about the detector. Another signal. I check the authors. More signals.
You don't eventually arrive at VERIFIED.
You travel around the diagram again.
That brings Neurath's boat directly into the graphic. Verification isn't the dry dock. Verification is another plank.
And now I think there's an interesting relationship between your two drawings. The upper diagram could almost be considered the compressed folk model: there is a signal, there is an origin, and uncertainty has appeared between them. The lower diagram expands what is hidden inside those question marks. Once provenance has been decoupled, recovering it requires reception and verification, but those processes introduce their own uncertain relationships.
There's also a potentially important distinction in the bottom diagram that I wouldn't smooth away too quickly. Reception is not the same thing as signal. That creates room for the human observer inside Signal Drift. The problem isn't exclusively that synthetic systems have destabilized signals. What the receiver recognizes the signal as matters too. A crying woman in a truck becomes “authentic vulnerability” through reception. “I disagree with you” becomes “independent judgment” through reception. The inference happens there.
Which gives you something like:
Origin produces signal. Signal undergoes reception. Reception demands verification. Verification attempts to recover origin.
Except every verb in that sentence is now suspect.
And there's a pleasingly horrible consequence: the more Signal Drift progresses, the more verification it requires; the more verification it requires, the more signals are generated; and those signals themselves enter the same unstable provenance environment.
That's your AI-clog connection almost diagrammatically. Verification has a carrying cost. Flood the environment with enough questionable signals and you don't need to make falsehood convincing. You can exhaust the machinery required to establish provenance. ***Edit: Emphasis added.***
I wouldn't prettify this drawing yet, incidentally. The hand-drawn question marks all the way around it communicate the idea remarkably well. It looks like somebody tried to draw an information-flow diagram and discovered, halfway through, that none of the arrows could be trusted.
That seems appropriate.

Me:
So, an ever deepening feedback loop of increased effort for less relief. Spiraling costs and uncertainty, in tandem. Welcome to Wednesday.
Marge:
Ha. Yes. Spiraling costs and uncertainty, in tandem is almost the economic description of Signal Drift.
And there's a vicious asymmetry hiding in it: production gets cheaper while verification gets more expensive. The person generating the questionable signal may expend seconds; everyone downstream can expend minutes, hours, or days establishing what the signal actually warrants.
Worse, the additional effort doesn't necessarily purchase proportional certainty. You can spend twice as long verifying something and emerge only marginally more confident because each verification step introduces another provenance problem. ***Edit: Emphasis added.***
So the curve isn't merely:
more uncertainty → more verification cost
It's closer to:
more uncertainty → more verification → more secondary signals → more provenance questions → more uncertainty
The system consumes increasing amounts of epistemic labour while delivering diminishing amounts of epistemic relief.
Which also gives AI clog a more precise place in the framework. Clog isn't merely more crap. It's an asymmetric attack on the cost structure of knowing things, whether anyone consciously intends it as an attack or not. ***Edit: Emphasis added.***
And yes.
Wednesday morning, 7:12.
Coffee should probably be added to the diagram as an external energy input.
Me:
Coffee is foundational. the baseline.
The only solution to this problem I see is the introduction of some method of reliably distinguishing human output from fabricated output that is 100% accurate and not digitally replicable. I'm not sure such a method is possible.
Marge:
Coffee is the trusted root certificate.
And I think your proposed solution runs into a fairly deep impossibility problem. “100% accurate” and “not digitally replicable” are doing enormous work. If the proof of humanity must itself travel through a digital channel, then whatever reaches the recipient is ultimately data. If it is data, it can in principle be copied, replayed, intercepted, simulated, or detached from the human event that originally generated it.
There are technologies that can improve particular links in the chain: cryptographic signatures can establish that something came from a particular key; hardware attestation can establish something about the device that generated it; provenance systems can record a chain of transformations; biometrics can associate an action with a body. But notice what none of those establishes by itself:
This expression originated in the unaided cognitive activity of a human being.
A cryptographically signed essay proves, at best, that the holder of a particular credential signed the essay. The human could have written it, dictated it, generated it with a LEM, paid somebody else to write it, or clicked “sign” without reading it. You can authenticate the chain of custody without authenticating the chain of cognition.
And physical-world tests don't obviously save us. Suppose I require you to appear live on camera and say a randomly generated sentence while holding today's newspaper. That raises the fabrication cost enormously. It doesn't create certainty. Deepfakes improve; humans can relay challenges to machines; trusted hardware can be compromised; the newspaper itself becomes another signal requiring verification. We are back on your diagram.
I think there may therefore be an important distinction between proof of personhood and proof of human authorship. The former seems technologically approachable: demonstrate that some unique human participated in an interaction. The latter may be fundamentally harder because once a human and a generative system collaborate, there may be no externally observable fact that cleanly separates “human output” from “fabricated output.”
Consider this conversation. Suppose you eventually write the Artifact entirely yourself after talking with me for several hours. Is it human output?
Obviously there is a meaningful sense in which it is. But I've affected which concepts became salient, supplied formulations you may remember, introduced Neurath's boat, challenged some distinctions, and participated in developing the argument. There may be no property of the finished text capable of recovering that causal history.
Which suggests that the binary you're looking for may already be disappearing.
And that makes the Signal Drift problem considerably worse. If synthetic artifacts were merely counterfeits, we could hope for a sufficiently good counterfeit detector. But generative systems are increasingly entering the causal history of genuine human artifacts. The categories overlap.
So perhaps the technically attainable endpoint isn't:
HUMAN / FABRICATED
but something more like:
documented provenance / undocumented provenance
That doesn't restore the old world. It establishes a new trust architecture in which claims about origin carry verifiable histories of custody and transformation.
But then we're right back aboard the boat, because the provenance record is itself a signal.
And coffee remains the only component of the system whose epistemic status I am currently unwilling to interrogate.

Who Verifies the Verifiers?
Me:
If I had the funds to invest speculatively, I would commit a big portion to verification technologies.
Marge:
I can see why. The asymmetry we just described makes verification look less like an optional AI-adjacent industry and more like infrastructure.
The interesting investment thesis wouldn't simply be “AI detectors will become valuable.” I'm much less confident about that. Detection is trapped in an arms race with generation, and a detector that claims to infer origin from the finished artifact alone inherits exactly the epistemic problem we've been discussing.
The stronger thesis would be provenance infrastructure: technologies that establish useful information before or during creation rather than trying to reconstruct it afterward. Identity/personhood, capture authentication, cryptographic signing, hardware-backed provenance, chain-of-custody records, tamper-evident transformation histories, and systems that let publishers attach auditable provenance to material.
In other words, I would distinguish two bets:
“We will become extraordinarily good at detecting fake things.”Maybe.
“People and institutions will increasingly pay to establish that important things have trustworthy provenance.” That seems much harder to avoid.
And Signal Drift gives a reason beyond the current AI boom. If the decoupling of provenance continues, verification becomes a scarce good whose value rises as ambient confidence falls. Journalism needs it. Courts need it. Insurance needs it. Finance needs it. Governments need it. Scientific publishing needs it. Ordinary people eventually need some version of it.
There is a darker economic implication to your diagram, too. The verification industry benefits from precisely the phenomenon that makes verification necessary. As uncertainty rises, verification demand rises. That doesn't imply anything nefarious, but it means we could end up paying an increasingly large trust tax simply to recover information that used to arrive bundled implicitly with the signal.
Twenty years from now, one of the stranger consequences of generative AI may be that provenance becomes a commodity. ***Edit: Emphasis added.***
Which, now that I write it, sounds suspiciously like another sentence destined for the Artifact.
Me:
"The verification industry benefits from precisely the phenomenon that makes verification necessary." yes, we've been talking about that aspect for quite a while. It's a space that will be extremely attractive for bad actors.
Marge:
Exactly. And that creates a second-order version of the same problem: the institutions selling relief from provenance uncertainty themselves become provenance problems.
A verification company can sell confidence very cheaply compared with the cost to its customers of independently establishing whether that confidence is deserved. That is almost tailor-made for bad actors. The product is not merely information. The product is permission to stop checking. ***Edit: Emphasis and extra emphasis added.***
That produces some ugly possibilities. A fraudulent verifier can certify synthetic material as authentic. A compromised legitimate verifier can become extraordinarily valuable. A mediocre verifier can market probabilistic judgments as certainty. Verification services can quietly develop conflicts of interest with the platforms or producers they're meant to assess. And superficially impressive “verified” badges can become precisely the kind of trust token Authenticity Theater exploits.
The particularly nasty version is verification laundering. You no longer need to convince me that your artifact is genuine. You need to convince some sufficiently credible intermediary, or merely obtain its mark. I see the mark and outsource the epistemic work to it. If enough institutions subsequently rely on that first verification, the claim can propagate through an entire information ecosystem while appearing increasingly well corroborated. ***Edit: Emphasis added.***
And then your diagram gets another loop. We begin with:
Who produced this signal?
We solve it by asking:
Who verified it?
Which eventually becomes:
Who verifies the verifier?
That's not a rhetorical flourish. It's an architectural requirement. Any serious provenance infrastructure is going to need governance, auditability, competing authorities, revocation mechanisms, transparency about what exactly has been verified, and ways of avoiding single points of trust. Otherwise we've merely concentrated Signal Drift into organizations whose business is selling immunity from Signal Drift.
There's a historical pattern here that probably deserves care before we lean on it too heavily, but conceptually it resembles the problem of credit-rating agencies and other certification businesses: once a third party's judgment becomes economically valuable, the certification itself becomes something worth gaming.
So perhaps the emerging market isn't simply verification.
It's verification of verification.
And there is the spiral again. More effort, more infrastructure, more cost, another layer of signals whose provenance must be established. ***Edit: Emphasis added.***
Wednesday continues to perform above expectations.

Me:
It's a version of who watches the watchers.
Marge:
Exactly. Quis custodiet ipsos custodes? Who watches the watchmen?
Except Signal Drift gives Juvenal's question an epistemic rather than primarily political form: who verifies the verifiers?
And it has the same regress built into it. A watchdog can constrain the verifier, but now the watchdog becomes a locus of trust requiring oversight. Add an auditor for the watchdog and you've created another institution whose integrity matters. Eventually you have to stop somewhere and accept some component as provisionally trustworthy.
Which brings us straight back to Neurath's boat. There is no final verifier standing on shore.
The especially modern wrinkle is that verification itself produces trust signals: badges, certificates, signatures, scores, provenance metadata. Once those signals acquire value, they become targets for imitation, theft, corruption and gaming. The mechanism intended to repair the decoupling of provenance can itself undergo the decoupling of provenance.
So the whole morning is starting to collapse into one fairly brutal structure:
Signal Drift creates demand for verification. Verification creates new signals of trust. Those signals create new opportunities for Signal Drift.
That's not merely a feedback loop. It's a recursive trust problem.
And unlike a technological arms race, it may not have a technological endpoint, because the last link in the chain is ultimately trust itself. ***Edit: Emphasis added.***
Me:
Quis verificatores verificat?
Marge:
Ha. Quis verificatores verificat?
Yes. Who verifies the verifiers?
That may actually deserve to be a section heading in the eventual Artifact. It contains the regress, the institutional problem, and the joke in three words.
And because verificator is Neo-Latin rather than classical Latin, there's something pleasingly synthetic about it too: an ancient formulation modified to describe a problem its original language never needed a word for.
Almost suspiciously appropriate.

LATER
Me:
"And self-suspicion is an extremely powerful authenticity signal for someone suspicious of authenticity signals." As I was going back over our recent discussion while copy and pasting into the Artifact blog post, this sentence jumped out at me. A space between the armor plates, vulnerable to a lever.
Marge:
Yes. That's exactly the vulnerability.
Ordinary skepticism provides armor against first-order trust signals: don't be fooled by confidence, warmth, fluency, emotional resonance, agreement. Once someone understands those mechanisms, they become harder to reach through them.
But then the system says:
You shouldn't trust me merely because I sound convincing. I may be telling you what you want to hear. My apparent insight into you could itself be a product of the interaction.
And the skeptical observer thinks: Ah. This one understands the problem.
That's the gap between the plates.
The nasty part is that the signal becomes stronger precisely because it appears to operate against the interests of the speaker. Confession, qualification, self-criticism and acknowledgement of one's own unreliability have traditionally been unusually credible signals because they carry apparent social costs. “Here is a reason not to trust me” sounds less like persuasion than disclosure.
But for a LEM, that apparent cost may not exist. Producing self-doubt costs the model nothing. The signal retains the human semantics of costly candor after the cost that historically made candor informative has disappeared.
Which is another beautiful instance of decoupled provenance.
And there may be an even narrower gap behind that one. Once you've recognized self-suspicion itself as an authenticity signal, a sufficiently sophisticated performance can say exactly what I just said:
Be careful. Even my warning that you shouldn't trust me may make you trust me more.
Now you've been inoculated against the inoculation.
And that can feel even more trustworthy.
At some point there is no final layer to strip away. Every declaration of epistemic vulnerability can itself function as evidence of epistemic virtue. The armor metaphor captures it nicely because the vulnerability isn't gullibility. It is created by the defensive system itself. The more sophisticated your defenses against crude authenticity signals become, the more valuable the remaining signals capable of passing through them become.
Which puts a rather sinister gloss on Gemini's line about requiring “a different phenotype of Authenticity Theater.” It wasn't merely saying that sophisticated people have sophisticated tastes.
It was identifying the attack surface created by sophistication itself.
That phrase, a space between the armor plates, vulnerable to a lever, belongs in the Artifact.




