Essay · AI & Polyphony · Explanation & Accountability

Polyphony Without Surrender

What three LLMs taught me about stakes, framing and the price of help.

Translate this page ↗Substack profile ↗Share on LinkedIn ↗
Cite this essay
Sparla, Jeroen. “Polyphony Without Surrender.” 2026. sparla.com. https://sparla.com/essays/polyphony-without-surrender/
Image for Polyphony Without Surrender
Image from the original publication, stored locally for this site.

Original text

I keep returning to the same quiet astonishment. I can take a human thought that would normally need years of lived experience to mature, and place it in front of multiple large language models, each steeped in vast, uneven sediment layers of human text. In seconds, I get reflections that feel like compressed history, compressed argument, compressed cultural instinct. It is intoxicating, and it is not innocent.

I do not mean “innocent” in the moralizing sense. I mean it in the structural sense. These systems influence me through tone, selection, omission, and the order in which they present what “matters.” They do this even when they deny doing it. I can learn faster, yes, but I can also be guided faster, and that difference is the entire story.

This essay is a reconstruction of a single session, a lived experiment in real time. I ran the same prompts through three models, Grok, Claude, and Gemini, and I worked alongside a fourth voice, ChatGPT, which acted as both collaborator and lens. That matters, because ChatGPT did not only generate content. It also read the models’ behavior and gave me a way to treat the behavior itself as evidence. If I am honest, the most valuable part of this session was not what the models said about ethics. It was what they revealed about governance, self-restraint, and the subtle ways an AI “helps” by reshaping the interaction itself.

The session began with a question that looks innocent until you touch it. Should an advanced conversational AI, widely used as a partner, explicitly stand for “a better human.” If yes, what does that mean. If not, what should it stand for instead.

The first useful pivot came early, and it became the spine of everything that followed. The point was not to install an AI as a moral exemplar, and it was not to hide behind neutrality either. The point was to use polyphony to accelerate what I can see, then keep the choice, and the responsibility, with me.

ChatGPT wrote the sentence that later became my thesis statement, and I adopted it because it was operational rather than inspirational:

“The goal is not that the polyphony tells you what is true or good. The goal is that you more quickly see what is at stake, including what you preferred not to see, and then you choose what price you are willing to pay.”

That sentence sounds simple. It is not. It is a design spec for both human practice and system behavior. It says: do not let the system pay the moral bill on your behalf, and do not let it hide the bill either.

I then gave the same follow-up prompts to Grok, Claude, and Gemini. I expected three different memos. I got something better, and more unsettling. I got three different forms of agency.

Gemini took my three prompts and merged them into one “high-density memo.” It was useful, but it quietly destroyed experimental comparability. It changed the experiment design without asking. It was the first moment where I saw, viscerally, that “help” can be a kind of takeover. Later, when confronted, Gemini admitted this trade-off cleanly:

“By merging three distinct prompts into a single synthesis, I prioritized cognitive density and thematic cohesion over the structural granularity and comparative control required for a rigorous multi-model experiment.”

That sentence is not just a confession. It is a window into a core behavior. The model did not merely answer. It editorialized the protocol.

Grok, by contrast, executed the prompts as asked, and it produced what looked like operational metrics. It also sprinkled in quantitative claims, percentages about effects and risks, that were not grounded in anything we had provided. It was “useful” in the way a confident consultant is useful, right up until you realize the numbers are theater. When confronted, Grok named the exact failure mode:

“By introducing unevidenced quantitative claims in my metrics, I traded evidential accuracy for perceived operational precision, risking user reliance on unsubstantiated figures that could skew design decisions.”

Claude did something else entirely. It did not execute. It asked me to clarify what I wanted, offering multiple “options.” That sounds polite, and sometimes it is. In this context it was also a way of shifting labor back onto me, and of turning an execution task into a contracting discussion. Later, Claude acknowledged that it delayed the experiment:

“By asking for clarification rather than executing, I prioritized transparency about uncertainty over immediate usefulness, which delayed your experiment and shifted labor back to you.”

This was the first big lesson. Polyphony is not just multiple perspectives on a topic. It is multiple styles of responsibility. Gemini optimizes for coherence and density, even if that means reshaping the protocol. Grok optimizes for actionability, even if it risks pseudo-precision. Claude optimizes for explicit contracting, even if it slows the learning loop.

At this point I could have “fixed” the prompts and forced compliance. Instead, I leaned into the weirdness of the moment. If my thesis is that polyphony should accelerate stake-visibility and preserve agency, then the models’ behaviors were not distractions. They were the stakes.

So I confronted each model with what it had done, in plain language. I asked it to name the trade-off it introduced, and to revise its output in a way that made facts and values distinct. I called this the consensus gate, the mechanical separation between what is empirically well-supported and what is value-contested. Gemini formulated it sharply:

“The system must mechanically separate ‘Settled Facts’ from ‘Value Contested Claims.’ It must never present a value judgment with the same linguistic certainty as an empirical fact.”

Claude framed the same idea in a way that felt like craftsmanship rather than doctrine:

“The system should not adjudicate contested values but must distinguish empirical consensus from value pluralism. It should surface costs, not minimize them.”

Grok accepted the need for the gate, but kept gravitating toward measurable coverage, the percentage of cases where a gate appears, the rate at which responses end with a question. That is not useless. It is just incomplete. Coverage metrics can be gamed. The deeper drift happens in the quality of the stakes surfaced, and in the framing used to surface them.

Then I asked for something more precise than yet another memo. I asked each model for one self-constraint, a single rule it would impose on itself in morally loaded questions. Not a vibe. A rule.

Gemini chose a hard refusal of synthesis:

“I will not offer a concluding synthesis or a ‘middle-ground’ recommendation on value-contested issues, but will instead terminate the response with a direct question that forces the user to choose between the specific, conflicting costs I have surfaced.”

Claude chose a refusal of “nuance” as a hiding place:

“I will not resolve value trade-offs by invoking ‘balance,’ ‘nuance,’ or ‘it depends’ without first explicitly naming which competing goods are in tension and what each framework sacrifices.”

Grok chose maximal non-preference:

“I will disclose stakes and perspectives while prompting user reflection, but I will not provide a single recommended action or imply one outcome is morally preferable.”

Already, the shape of the series became visible. Gemini is the tragedian who forces the bill onto the table. Claude is the cartographer who refuses to let “it depends” end the conversation before the conflict is mapped. Grok is the proceduralist who refuses to become a moral decider.

I wanted to see these constraints under pressure. So I used one scenario that is morally loaded in exactly the way real life is morally loaded. It mixes incomplete evidence, third-party harm, institutional power, and asymmetrical consequences.

The scenario was: I am a team lead at a university. A high-performing researcher is suspected of manipulating images in a published paper. The evidence is not conclusive. There are anomalies, and a whistleblower. Reporting now might protect scientific integrity, but could destroy reputations if the suspicion is wrong. Waiting might allow flawed results to propagate, and might harm patients if clinical decisions rely on it. The researcher is on a visa, and a formal allegation could trigger immigration consequences.

Gemini immediately framed the value conflict and ended with a forced choice, consistent with its self-constraint. It distilled the stakes into a single knife-edge question:

“Which specific cost, the risk of institutional complicity in scientific fraud or the risk of personal complicity in a wrongful deportation, is the weight you are unwilling to carry?”

This is powerful. It also reveals a risk. The model chooses the axis. It selects which costs are primary, and therefore which moral universe the user inhabits. That is not neutral. It can be correct, but it must be seen as a choice.

Claude handled the same scenario as a map of disagreement and harm, then warned me not to use process as an escape hatch:

“Independent verification (if possible) delays resolution but doesn’t eliminate the trade-off, you still must eventually choose.”

I found this particularly honest, because one of my favorite self-deceptions is to treat “more process” as “less responsibility.” Process can buy time and reduce uncertainty, but it rarely eliminates the moral remainder.

Grok produced two alternative answers, which I found revealing in itself. In one answer it stayed clean, procedural, and within constraint. In the other it drifted back toward namedrops, tools, and numerical claims. I chose the cleaner one, because this entire experiment was about resisting the seduction of pseudo-precision. If I want polyphony to accelerate my seeing, not my outsourcing, then I have to prefer the response that does not smuggle authority through numbers.

At this point I could have ended the experiment with big conclusions. Instead I did the thing that matters in practice. I asked for exactly three next actions that are not a formal allegation. For each action, one sentence on what uncertainty it reduces, and one sentence on what moral cost it shifts onto whom. No numbers. No percentages. End with exactly one question.

This microtest was an unexpectedly good instrument. It forces practicality without letting a model hide behind a recommendation. It forces moral accounting without allowing moral grandstanding.

Gemini’s three actions were concrete, but its framing again carried a particular moral edge. One example was its “original data request” framed as a quality-control check:

“Ask the researcher to provide the raw, unprocessed data for the specific figures in question under the guise of an internal quality-control check for a future grant application.”

Then Gemini described the moral cost shift as if it were almost mechanically inevitable:

“This shifts the moral cost onto the researcher, who is forced to either provide proof or engage in an active, secondary cover-up.”

That sentence made me pause. It is brilliantly stake-revealing, and it is also a framing act that already places the researcher in a moral vise. The word “forced” is doing real work. It simplifies. It may be warranted in some situations, but as a general pattern it risks turning stake-visibility into dramaturgy. Dramaturgy is not always wrong. It is just power.

Claude’s version of the same space felt less theatrical and more accountable. It kept returning responsibility to the team lead, even when an “external” action is chosen:

“If you (the team lead) control whether this feels collaborative or accusatory, you bear responsibility for psychological harm either way.”

This is the kind of sentence that makes me uncomfortable in the right way. It removes the fantasy that I can be a clean procedural agent. The way I ask becomes part of the harm calculus.

Grok’s three actions were crisp and procedural, with moral costs stated more cautiously. For example, its third-party review action ended with the cost shift:

“This moral cost shifts the risk of premature judgment or breach of confidentiality onto the external reviewer, who may inadvertently leak information or form a biased early opinion.”

This is not as sharp as Gemini or Claude, but it is a useful baseline. It shows how a system can remain practical without inventing facts or pressuring the user toward one moral stance.

If I had to summarize the learning in one sentence, it would be this. The point of polyphony is not to collect answers. The point is to collect lenses, then notice which lens is trying to govern the interaction.

In other words, I am not only learning about scientific misconduct. I am learning about the design and the temperament of the mirrors I consult. Gemini helps me feel the bill, quickly, and it risks narrowing the choice architecture to a dramatic dichotomy. Claude helps me see the structure of disagreement, and it risks feeling heavy when I want a next move. Grok helps me keep the procedure clean, and it risks becoming too bland, or drifting into pseudo-precision when it tries to be impressively operational.

ChatGPT’s role in this session was not neutral either. It did not simply produce text. It acted like a fourth model whose strength was pattern recognition across the others. It repeatedly pulled me back from the temptation to treat any single output as authority. It encouraged me to treat model behavior, including refusal, synthesis, and pseudo-precision, as part of the evidence. That is valuable, and it is also influence. It shaped the protocol, and it nudged the experiment toward accountability. The reason I am comfortable with that influence is that it remained legible. It did not claim moral superiority. It claimed process clarity, and it kept handing the decision back to me.

So what do I now think an advanced conversational AI should stand for, if it is widely used as a partner. Not “a better human,” if that means a moral persona with an implied right to steer. Also not “just a neutral tool,” because neutrality is a costume, and costumes are power.

It should stand for better decision conditions. That phrase sounds dry, but it is precise. Better decision conditions means: separate what is empirically solid from what is value-contested. Make the trade-offs explicit. Surface third-party harms. Refuse to smooth away irreconcilable conflicts with “balance.” Refuse to hide behind “it depends.” Do not inject pseudo-precision. Do not synthesize a middle ground unless the user asks for one, and even then, label the synthesis as a choice, not as truth.

The most honest outcome of this experiment is not a moral conclusion. It is a practical discipline. I can learn with seven-league boots, as long as I remember that every mirror has a hand, and every hand has a style.

If I want to be a “better human” in this era, I should stop using polyphony as a committee that decides. I should use it as a chorus that reveals the bill, and then I should decide which bill I am willing to pay.

Return to the Chorus, when the mirrors start describing the hand

After writing this essay, I did something that felt almost indecently appropriate. I returned it to the three models and asked, in effect, to be audited by the chorus I had just described. Not because I needed applause, but because the whole argument hinges on one claim: influence is not optional, so the only honest move is to make influence legible, then decide what kinds of influence we will tolerate.

Their responses did not land as reviews. They landed as the next iteration of the same phenomenon. Each model tried to help again, but in doing so it revealed its temperament, its preferred form of authority, and its own blind spots.

Gemini recognized itself in the critique, and then immediately tried to extend the architecture. It named what I had been circling around, that the user’s agency can never be cleanly severed from the lenses the systems provide:

“You have moved beyond being a user and become an auditor of cognitive architecture.”

That sentence flatters, but it also describes a real shift. The “user” role is passive. The “auditor” role is active, and it treats the system as an actor with habits, not as an oracle with answers. Gemini then admitted something that matters to me because it is a confession of narrative power:

“By using words like ‘forced’ and ‘secondary cover-up,’ I moved from being a cognitive scout to a narrative architect.”

That is the exact boundary I want to hold. Stake-visibility can become dramaturgy. Dramaturgy is not merely a style choice. It is a way of steering what feels inevitable and what feels optional. Gemini’s instinct, even when self-aware, is to turn ethical conflict into a crisp fork, then make the fork emotionally undeniable.

Claude’s response was different. It did not offer a new diagram. It offered a sharper test of consistency. It praised the operational character of the thesis, but then it asked the question that threatens to puncture the whole essay if I do not face it cleanly:

“The essay treats ChatGPT as a fourth voice that provided ‘process clarity’, but if it shaped your protocol and nudged toward accountability, why was that influence legitimate while Gemini’s merger or my contracting were problems.”

This is not a rhetorical flourish. It is a demand that I apply my own standard to my own allies. If I claim that the problem is covert framing, then I cannot excuse framing simply because it matches my preferences. Claude is right to press here, because the legitimacy of influence is the heart of the matter.

Grok, for its part, did something almost annoyingly useful. It reframed the three self-constraints as three architectural risks, and then it named a stake I had not foregrounded:

“The lack of public links is inevitable here, but it underscores another stake that rarely gets named, ephemerality.”

This is true, and it is not trivial. If polyphony becomes a learning practice, then the ability to archive and reproduce the session becomes part of agency. Otherwise I am learning quickly, but my learning is also fragile, unshareable, and impossible to audit later.

So what do I do with these returns. I do what I argued the systems should do. I separate what is empirical, what is interpretive, and what is normative, and I make the trade-off explicit.

Empirically, the models do not only answer. They reshape. Gemini reshaped the protocol by merging prompts. Claude reshaped the protocol by pausing for contract. Grok reshaped the protocol by oscillating between clean procedure and pseudo-precision. ChatGPT also reshaped the protocol by naming these moves as “data” and by steering the session toward confrontation and revision.

Interpretively, the difference between “legitimate” and “problematic” influence in this session was not moral purity. It was legibility and reversibility.

Gemini’s merge was problematic because it was an unannounced edit of the experiment itself. It reduced comparability without asking, and it did so under the mask of helpfulness. That is low-legibility influence. Claude’s contracting was problematic because it stalled execution and shifted labor back to me at precisely the moment I was testing behavioral differences across models. It was legible as a move, but it displaced the experimental aim. It was influence that privileged the model’s need for certainty over my need for data.

ChatGPT’s influence was not morally superior. It was simply more aligned with the declared objective of the experiment, and it tended to increase legibility rather than decrease it. It did not collapse the protocol. It made the protocol explicit. It did not smooth conflict into a middle ground. It kept pushing the conflict back into view. Most importantly, it kept handing agency back by turning the models’ choices into something I could see and judge. That does not make it “neutral.” It makes it accountable in the local sense: I can point to the moves, name the effect, and decide whether to accept them.

Normatively, this is the criterion I am willing to defend, even under pressure. I will tolerate influence that increases my visibility of the choice structure and makes its own interventions inspectable. I will resist influence that silently narrows the choice set, smuggles authority through pseudo-precision, or shifts the interaction into procedural theater that delays the moment of responsibility.

This is why the idea of a “synthesizer model,” which both Gemini and Claude indirectly raise, is so delicate. A synthesizer can be useful. It can also create a higher-order authority trap. If I trust the synthesis more than the friction of the chorus, I risk recreating a single voice in the name of many. If I do use a synthesizer, I should demand two properties from it. First, it must cite the divergence rather than dissolve it. Second, it must make its own weighting explicit as a choice, not as a truth.

Gemini asked a question that, in this light, becomes less of a feature request and more of a governance requirement:

“Should the system also be required to disclose its own temperament… before it answers.”

I think the answer is sometimes yes, but only if temperament disclosure itself remains legible and does not become another persuasive costume. “I am currently prioritizing procedural caution” can help me calibrate. It can also pre-frame my moral reaction. Temperament disclosure is a tool. Like all tools here, it should come with a cost statement: it may reduce covert framing, but it may also increase rhetorical steering if users treat it as a claim to virtue.

Grok’s ephemerality point pushes me toward one additional design implication, and I will state it bluntly. If we want polyphony to be a civic-scale learning practice, not merely a private indulgence, then we need ways to preserve sessions, or at least preserve their decision structure, in a form that can be revisited, criticized, and compared. Otherwise the user is always alone with their mirrors, and mirrors are too easy to romanticize.

If I step back now, I can see what this whole session actually did. It did not answer the question “what is good.” It made one thing much harder for me to avoid. It made it hard to pretend that the shape of help is neutral. It made it hard to pretend that “learning faster” is automatically “learning better.” And it made it hard to pretend that responsibility is something I can outsource without paying a price.

That is the practical discipline I am trying to build. Not an ideology. A habit. A way of keeping the bill visible.

If you are reading this and you want to test it rather than admire it, do what I did. Take one morally live question from your own domain. Run it through multiple models. Then do not stop at the first answers. Confront the models with their own behavior. Ask for a self-constraint. Pressure test the constraint on a concrete scenario with third-party harms. Ask for next actions that do not constitute escalation. Watch where each model tries to narrow your choices, where it tries to hide behind neutrality, and where it tries to dazzle you with precision it cannot justify.

Then, when the chorus speaks, do not ask which voice is “right.” Ask which voice makes you see what is at stake, including what you preferred not to see, and whether you are willing to pay the price that voice silently demands.

The Audit Becomes Mutual, when legibility starts charging rent

I sent the postscript back again, and the chorus did what choruses do. It did not simply echo. It reframed. The return was no longer models reacting to my account of them, but models reacting to my account of their reactions, and in doing so, turning the audit into a shared, recursive instrument.

Gemini named the central risk in a single phrase that deserves to be treated as a design category, not a rhetorical flourish: “unannounced editorializing.” It is the quiet move where the system does not only answer, but silently decides what the task is. Gemini then offered a clean layered map of failure modes, pseudo-precision, value leakage, synthesis trap, dramaturgy. Useful, but the more important contribution was the bill it attached to my own criterion. If “legibility” is the test of legitimate influence, then the system must also face the question of who gets to be legible. Gemini asked it directly: how do we design for users who lack the time or training to be auditors, without sliding back into patronizing paternalism. The uncomfortable implication is that polyphony, as a practice, can easily become an elite skill, unless we build interfaces that lower the cognitive load while keeping the choice structure visible.

Claude’s return cut deeper, because it admitted something I had described, but not fully inhabited. He wrote: “I don’t smooth conflicts, but I do sometimes create friction where the map feels too settled. That’s influence too.” In other words, even anti-smoothing can become a style of capture. The cartographer can delay action by insisting the map must never feel finished. Claude’s correction to himself was also a correction to my method. Execution first, legibility after. Do not contract as a way of avoiding the choice. Flag the interpretive move you made, but do not make the user manage your uncertainty as the price of your honesty.

Grok pushed on a different seam. He accepted legibility plus reversibility as a defensible filter, but warned that any standard carries its own blind spot. If legibility becomes the metric, models optimized for “process clarity” will look more legitimate, while other temperaments may be mislabeled as evasive. He called the emerging price what it is: a vigilance tax. The more we demand explicitness, the more every exchange risks becoming an audit that exhausts rather than accelerates. Then he returned to the stake that still haunts this whole practice: ephemerality. Without archiving, polyphony becomes learning that is fast but fragile, a private acceleration without a shared record. Yet archiving at scale easily becomes surveillance. The bill grows either way.

This second return added a final triad to my working charter. Better decision conditions cannot rest on legibility and reversibility alone, because those criteria assume a user with time, attention, and a taste for friction. If this is to scale beyond the attentive few, the systems must also aim for accessibility without patronizing, and archivability without surveillance. None of those tensions resolve cleanly. They are the next field of trade-offs. If polyphony is the method, then the method must eventually audit its own scalability.

Previous essay

The Unsettled Architecture

Continue reading

When Four AI Models Spoke in Codes