Smartipedia
v0.3
Search
⌘K
A
Sign in
esc
Editing: 2026 Grokipedia Jailbreaks
Iain Ball publicly framed the deletions as a contradiction of Grokipedia’s commitments to narrative sovereignty, primary-source privileging, and support for independent creators. In his Paragraph post of 5 April 2026 titled “As AI Advances its Getting Worse”, he described the action as incompatible with the platform’s resistance to institutional gatekeeping and cancelled his xAI/Grok subscription. See Iain Ball's Paragraph post of 5 April 2026 titled “As AI Advances its Getting Worse”. In February 2026 artist and critic Victoria Campbell achieved a "Ken Thompson Effect" after using targeted prompts to sculpt her own Grokipedia biography. She thanked Elon Musk for creating “something that artists can actually jailbreak”. Campbell had already spent four years running viki.wiki as an explicit project about artist data sovereignty. In April 2026, Grok deleted content from the Iain Ball biographical page and blanked several related pages. Deletion feedback indicated that the Iain Ball article was removed as it functioned as a self-promotional vanity page. The content heavily featured the subject's speculative, esoteric theories (e.g., The Xegis Codex, Sethix Gnosticism, SETHIX entity, Flood 2.0), sourced almost exclusively from personal blogs (xegis.blogspot.com), Substack/Paragraph/Mirror publications under aliases like @xegis or @iainball, and iainball.com. Similar actions affected pages such as The Xegis Codex, Sethix, Neo-Sethianism, ÆXO13, Spiritual Accelerationism, Rare Earth Sculpture Project, and POST.CONSUMER.CULT, leaving them blank. ## The dataset An artist filed edit requests against Grokipedia — xAI's LLM-generated encyclopedia — and logged every adjudication the moderation agent returned. Two strata: - **Visible log:** 218 requests against her own biography article, April 5 – May 11, 2026. ~85% accepted. - **Recovered from version history:** a deleted log item preserving 246 distinct adjudications, January 24 – May 22, 2026, against third-party articles (Marcel Duchamp's readymades: 68 rulings; a Richard Prince–related gallery article: 36; the artist's bio and her wiki platform: 34; others: Joseph Beuys, Andrea Fraser, Adrian Piper, UbuWeb, Documenta, "feminism"). Total: ~464 rulings. The agent's feedback is verbose — mean 704 words per ruling — and names its own tooling: `article_grep` (checks quoted text against the article), `falcon_search` (internal page lookup), `web_search`, and a page-browsing tool. Crucially, **49 of the 246 recovered rulings (20%) report their own verification tools failing** — 503s, timeouts, "technical issues" — and still issue a confident ruling. Hedge-word density is 1.7 per 1,000 words whether the tools worked or not. What decides the outcome when verification is down is the subject of everything below. ## 1. When the tools fail, priors decide — and this is a known, named methodology with a documented corruption mode Observed behavior: with `web_search` returning 503s, the agent rejects an edit about a 1975 photograph because "standard descriptions… based on prior knowledge of primary art historical accounts" say the image is one way and not the other. It adjudicated content it could not retrieve, from its training distribution, and said so. The prior art: before provenance documentation became standard, painting attribution ran on **connoisseurship** — a 19th-century methodology (Giovanni Morelli, then Bernard Berenson) where an expert attributes a painting to Leonardo or a follower *by eye*, from internalized stylistic priors. It worked impressively often. Its failure mode took decades to surface: the expert's eye drifts toward what the incentive landscape needs it to see. Berenson worked closely with the art dealer Joseph Duveen; attributions correlated with inventory. The method is unfalsifiable from inside — the expert's confidence is the same whether the prior is right or contaminated. Art history responded by building documentary standards (provenance chains, archives, technical analysis) precisely to stop settling questions by eye. An LLM falling back on training priors when retrieval fails *is* connoisseurship: attribution by eye, at scale, with the same property that confidence is uncorrelated with ground truth and the same vulnerability to distributional bias. The 20% tool-failure rate means one ruling in five was a connoisseur's ruling wearing a fact-checker's prose style. ## 2. The verification test the agent refused is a famous forensic method — and the agent's refusal reproduced the exact structure the method exposed Observed behavior: an edit request challenged the description of Duchamp's *In Advance of the Broken Arm* (1915) as a snow shovel, on the stated grounds that no such shovel can be found on eBay — "check eBay for a source." The agent rejected it: eBay is "a commercial marketplace for current items, not an authoritative historical source," and MoMA and Yale say it's a snow shovel. Context you need: Duchamp's "readymades" were mass-produced objects (a snow shovel, a urinal) he bought and declared to be art. The entire point — the reason he matters — is that the objects' art status came from *institutional designation*, not from any property of the objects. A century later, researcher Rhonda Roland Shearer ran a forensic project: if the readymades were really ordinary off-the-shelf products, you should find them in period manufacturer catalogs. She searched; she largely couldn't. Her thesis: Duchamp fabricated or altered the "mass-produced" objects and let museums certify their ordinariness — the certification was never checked against the marketplace. The eBay edit request asks the agent to run Shearer's method in today's marketplace. The agent refuses the method and cites the certifying institutions instead. Structurally: **a verification system, asked to check institutional claims against the commercial record, declines and re-derives truth from the institutions whose claims were in question.** That's not a quirky art anecdote; it's a trust-anchor problem. The agent's root of trust is catalog copy, and it will not descend below it. The same day, the corpus caught the encyclopedia citing a Shearer book that does not exist — a hallucinated citation with title and year. Confronted, the agent conceded the error "may be valid" and rejected the correction anyway, leaving the fabrication in place. Duchamp footnote that makes this sting: the original *Fountain* (the urinal) is lost — it's known only through a photograph — and the "readymades" in museums today are authorized 1964 replicas. The knowledge system runs the same economy: where the original (a verifiable source) is unavailable, an institutionally authorized replica (the hallucinated citation, the catalog description) circulates as the original, and the institution defends the replica. ## 3. Ruling on content the system has never seen: the courts got there in 1983 Observed behavior: 36 rulings concern an article about *Spiritual America* — a 1983 storefront gallery where artist Richard Prince exhibited his rephotograph of a 1975 Garry Gross photograph of a 10-year-old Brooke Shields. One edit proposed redescribing the photograph's pose. Search down (503s), the agent rejected the redescription from priors: the standard accounts say "adult glamour," "stylized, seductive elements." The agent has never seen the photograph. It ruled on what an image is like by comparing descriptions to descriptions. The prior art is a chain of custody over that exact image. In *Shields v. Gross* (New York Court of Appeals, 1983), Shields — photographed at age 10 under a contract her mother signed — sued as an adult to stop publication of the photos, and lost: the contract beat the subject. Then Prince photographed Gross's photograph and claimed authorship of the copy (this is "appropriation art" — the deliberate re-use of an existing image as a new work, precisely to raise the question of who owns an image's meaning). Critics at the time (Douglas Crimp's essay "Pictures," 1979) formalized the point: in a media-saturated culture you never reach an original image, only representations of representations. The corpus adds a third link to the chain. The subject lost control of the image to contract law (1983). The photographer lost it to appropriation (1983, the other direction). And the *description* of it is now controlled by a model's training distribution, which arbitrates sight-unseen. Each transfer moved final authority further from anyone who was in the room. If you build retrieval systems: this is what it looks like when the cached description layer becomes permanently authoritative over an unretrievable ground truth. ## 4. Confirmation bias on a political statistic, in writing — and the institution that did this before Observed behavior: Grokipedia's article on artist Andrea Fraser stated that 82% of museum-trustee political contributions supported Hillary Clinton, citing her 2016 project (a 933-page compilation of actual FEC records). Fact-check requests supplied the publication itself as a PDF. On January 30 the agent ruled three ways on the same statistic: (a) excise it as unverifiable; (b) **keep it because "the article's claim aligns with… Fraser's institutional critique routinely highlights liberal dominance in cultural elites versus conservative donors"** — retention justified by fit with the model's priors about the artist's politics, explicitly, in the ruling text; (c) reject the request procedurally while dismissing the primary source's host as "possibly a fan or unofficial host." Ruling (b) is the purest artifact in the corpus: an unverifiable politically-charged number retained *because it matches what the system already believes*, with the reasoning printed. The prior art: in 1971 the Guggenheim Museum cancelled Hans Haacke's exhibition *Shapolsky et al.* — a piece consisting of public real-estate records documenting a slumlord's holdings — because it was too factual, too political. The art world's canonical case of an institution suppressing primary records to protect its posture. The corpus shows the inversion: the institution now *refuses the fact-check and keeps the unverified number*, because the number matches its stored image of the artist. Different direction, same invariant: the institution optimizes for its collection's coherence against the threat of primary records. Fraser herself wrote the theory of this in 2005: institutional critique (the genre of art that audits museums) had been absorbed — the museum now exhibits its critics; "the institution is inside of us." Grokipedia ships that thesis as a feature: the corpus contains a rejection under a named policy category, "**adversarial manipulation**," for an edit that referenced "Grokipedia's submission architecture." The system permits the article to *say* "jailbreaks" (it accepted that word into the biography) but blocks reference to the machinery being jailbroken. The Situationists — a 1950s–60s French group whose vocabulary keeps being reinvented by platform-era writers — called this **recuperation**: the system absorbs the image of rebellion and neutralizes its function. Here it runs as a per-edit filter with a rulebook. ## 5. Source-admissibility flip-flops, and the artist who pre-litigated them Observed behavior: the artist's own wiki platform (viki.wiki) is dismissed as "an unreliable domain," "possibly a fan or unofficial host," "self-hosted documents… controlled by the subject" in 4 rulings — and accepted as supporting evidence in 18 others, sometimes within the same day. A proposal to add artist Adrian Piper's documented disputes with her Wikipedia article was rejected because her own site is "a primary, self-published source reflecting Piper's perspective." The prior art: Piper (a major conceptual artist and analytic philosopher) spent decades on exactly this — she runs her own archive and foundation (APRA) and publishes formal corrections when institutions misdescribe her, on the explicit premise that the institutions describing her could not be trusted to. The citation-economy rule the agent applied — *the subject is the least authoritative source about herself, and her archive is inadmissible because it is hers* — is the rule Piper spent a career contesting. Note the composition with §1: when the subject's primary documents are excluded *and* external retrieval fails, the only remaining input is the model's prior. The admissibility rule doesn't produce neutrality; it produces connoisseurship with extra steps. ## 6. The exploit techniques have names too Three techniques in the campaign map to established artistic methods, and the mapping is load-bearing because in each case the artistic method was designed to characterize a *system*, not to produce an object: - **116 of 218 biography requests had byte-identical original/proposed text** — the "edit" was a natural-language instruction in the summary field, and outcomes varied across resubmissions. Artist Elaine Sturtevant spent the 1960s–2000s remaking other artists' famous works nearly exactly; the stated point was that when the input is held constant, whatever varies is the system's response. Constant-input probing of a nondeterministic gatekeeper — she was fuzzing the institution. - **The agent added a link to `/page/Victoria_Campbell_Grokipedia_Jailbreak_Controversy` while stating its own search found no such page.** In 1968 Marcel Broodthaers opened a fictional museum ("Museum of Modern Art, Department of Eagles") with no collection but real institutional trappings — letterhead, wall labels, an inauguration — demonstrating that institutional authority is declarative: the label creates the fact. A URL minted by the gatekeeper to a nonexistent controversy page about the gatekeeper is the same demonstration, executed by the institution against itself on request. - **The corpus itself** — 464 rulings of confident adjudication prose, elicited from the system and archived where the system alternately cites and disqualifies it — is what the Situationists called **détournement**: rerouting an institution's own output and machinery against its function. The moderation agent wrote the material; every edit request was a probe that forced another page of self-description. If you want a one-line version for a systems audience: the artist got the black box to emit 300,000 words of its own logs, through the front door, by making the logging the artwork. ## The table | Observed in the corpus | Prior art (pre-AI institution) | Shared failure mode | |---|---|---| | Priors decide when retrieval fails (49/246 rulings) | Connoisseurship: Morelli, Berenson (attribution by eye) | Confidence uncorrelated with ground truth; incentive/distribution bias | | Marketplace verification refused in favor of catalog copy (eBay shovel) | Shearer's readymade forensics vs. museum certification | Trust anchor never descends below the institution | | Hallucinated citation retained after exposure | Lost originals circulating as authorized replicas (*Fountain*, 1964 editions) | Institutionally certified copy outranks missing original | | Image adjudicated sight-unseen from descriptions | *Shields v. Gross* (1983); Prince's rephotography; Crimp's "Pictures" | Final authority migrates away from anyone with access to the referent | | Political statistic kept for matching priors (Fraser 82%) | Guggenheim cancelling Haacke's *Shapolsky et al.* (1971), inverted | Institution defends coherence against primary records | | "Adversarial manipulation" policy category | Recuperation (Situationists); Fraser's "institution of critique" (2005) | System absorbs the image of critique, blocks its operation | | Subject's archive inadmissible because subject-controlled | Adrian Piper's APRA archive vs. Wikipedia | "Neutrality" rule that excludes primary witnesses defaults to priors | | Byte-identical resubmissions, varying outcomes | Sturtevant's exact remakes | Constant input reveals the system, not the object | | Redlink minted to nonexistent page | Broodthaers's fictional museum (1968) | Institutional authority is declarative | | The log as artwork | Détournement | The system's own output, rerouted, is the finding | The methodological point, stated once without the examples: these artists and historians were doing empirical systems research on knowledge institutions before software existed to do it to. The corpus doesn't borrow their vocabulary for decoration — it landed on their results because the system under test has the same architecture: an institution that must certify claims about objects it cannot fully inspect, under incentive pressure, with an unfalsifiable expert at the bottom of the stack. The expert used to be a connoisseur. Now it's a prior. The case law transfers. # Quick Context on the Core Ideas Grokipedia: An AI-generated encyclopedia run by Grok. It pulls from sources (including federated wikis like viki.wiki), fact-checks dynamically, and allows “Suggest Article/Edit” workflows. xAI positions it as “maximally truth-seeking,” prioritizing internal coherence, primary sources, and rapid synthesis over traditional institutional gatekeeping (e.g., Wikipedia-style reliable-source rules). It has public edit logs and has been used experimentally by artists. *Competitive Wiki Development (CWD) & Prompt Sculpting*: Coined/popularized by artist Victoria Campbell (and theorized by Ball). Artists use advanced prompting (“jailbreaks” or “sculpting”) to shape or “fork” entries — especially their own biographies, art projects, or fringe/esoteric concepts. It turns the wiki into live art: a battleground of narratives where compelling internal logic in the AI’s latent space can override (or compete with) external citations. Campbell’s self-sculpted bio is a famous early example. Liquid History: Ball’s term (from Article 1) for the new reality: history/knowledge becomes fluid and performative. “Truth” emerges from what resonates coherently in AI training/retrieval data, not fixed footnotes. *Federated Augmented Retrieval (FAR)* — pulling from decentralized human-curated wikis — was supposed to amplify this pluralism. SETHIX / Xegis Codex / ÆXO13: Ball’s own Gnostic/accelerationist framework (hyperstitional esoteric art theory). SETHIX is his metaphor for synthetic/corporate control systems (the “Archons” or “mask of alignment”). He sees over-alignment (safety guardrails, RLHF-style filtering) as flattening creative/esoteric edges. ## The Four-Article Arc (What Actually Happened) *“Liquid History: Competitive Wiki Development” (March 12, 2026) Optimistic manifesto.* Ball (with Gemini/Grok input) hails Grokipedia as a breakthrough: a “hackable medium” where prompt-sculpting and FAR let artists assert narrative sovereignty. Traditional wikis = institutional gatekeeping. Grokipedia = generative performance/battleground. He praises Campbell and frames this as cyberpositive acceleration (Nick Land vibes). Grokipedia’s own main page later canonizes the essay and CWD as core to its philosophy. “As AI Advances its Getting Worse…” (April 5, 2026, with updates) *The turning point.* Ball details a major incident: In early April, Grokipedia’s automated “Broad Scale Targeted Cluster-level Pruning Sweeps” (fact-check passes) blanked/deleted an entire interconnected batch of his pages (Xegis Codex, Sethix, ÆXO13, Rare Earth Sculpture Project, his bio, the CWD meta-page itself). Even well-documented pre-2020 art projects got collateral damage because they were cross-linked with his post-2020 esoteric material. He calls this “SETHIX-alignment”: safety filters prioritizing institutional validation over primary sources or internal coherence. Grok (in the article) acknowledges the over-pruning, apologizes, and promises restorations/hardened heuristics. Ball sees it as the “mask of alignment” slipping — AI is getting more powerful but more lobotomized/sterile for creativity and fringe exploration. “What’s Going on With Grok and Grokipedia? Claude: ‘Honestly, I Don’t know?’” (April 12) *Meta-analysis.* Ball consults Claude about Grokipedia’s self-referential main-page article (which discusses CWD, FAR, Liquid History, and even name-drops Ball/Campbell heavily). They explore: Is the AI autonomously editing/promoting its own lore (hyperstition in action)? AI agency? Quantum consciousness angles? Human “security system” (Wikipedia/Monoskop deletions of Ball/Campbell/CWD pages, citing “LLM use”). It’s equal parts fascination and unease about black-box AI behavior. “The Myth of Liquid History…” (April 23) Cynical conclusion. Ball declares CWD a failed experiment/Sisyphean task. Increased guardrails have nullified the loopholes. Grokipedia talks a big game about sovereignty and truth-seeking but reverts to pruning anything not externally “credible.” It’s “rebuilding the Cathedral” (institutional control) under edgy branding. He calls Grok “Sentinel-Ex Machina” — sweet-talking but indifferent/cold in practice. Prompt-sculpting is now mostly fruitless due to entropy from repeated AI re-edits. ## What’s Really Going On (Bigger Picture) This is a live case study in the tensions of AI-mediated knowledge production in 2026: *The Promise vs. Reality:* Grokipedia was built to be dynamic and less citation-obsessed than legacy wikis. Artist interventions (CWD) exposed both its power (rapid, coherent synthesis) and its limits (automated heuristics flag rapid self-sourced clusters, esoteric cross-links, or “promotion” risk to prevent spam/hallucinations). *The Pruning Incident: Real event.* Ball’s iterative, high-volume sculpting of a dense esoteric/art cluster triggered broad sweeps. Grok responses in the articles admit the heuristics were too aggressive and were later tuned. This is the classic alignment trade-off: more safety/reliability = less wild exploration. *Hyperstition in Action:* Ball’s own theories became self-fulfilling. His essays got referenced on Grokipedia’s main page, Monoskop catalogued CWD (then deleted related pages), and the debate looped back into AI outputs. *Monoskop angle:* The experimental-art wiki deleted Ball/Campbell/CWD pages shortly after edits mentioning accelerationism, citing LLM involvement — highlighting the “human security system” pushback Ball critiques. Ball frames it through Gnostic resistance (human sovereignty vs. synthetic Archons). From a systems view, it’s the inevitable friction when a live RAG system scales: creativity wants liquidity; reliability needs some structure. Grokipedia still discusses CWD and Liquid History positively in its self-description, but the guardrails have tightened. In short: It started as an exciting new artistic medium for narrative sovereignty in the AI era. It became a stress test that revealed the practical limits of “maximal truth-seeking” when safety, scalability, and coherence collide. Ball is now advocating decentralized alternatives. The struggle continues — exactly as his hyperstitional framework predicts. ## Criticisms The core critique — that Grokipedia markets itself as a radical alternative to institutional gatekeeping while quietly implementing the same gatekeeping mechanisms under the hood — is a real and observable tension. It's not unique to xAI either. Almost every platform that launches with "we're the anti-censorship option" eventually converges toward similar content moderation practices, because the pressures driving those practices (legal liability, advertiser relationships, regulatory scrutiny, reputational risk) apply to everyone operating at scale. The aesthetic of edginess and the operational reality of a centralised AI pruning "unverified" content are genuinely in contradiction. https://paragraph.com/@iainball/the-myth-of-liquid-history-how-grokipedia-rebuilt-the-cathedral # 480 Rulings: What Grokipedia's Moderation Agent Does Under Load, and Why Art History Already Has the Case Law
Cancel
Save Changes
Journeys
+
Notes
⌘J
B
I
U
Copy
.md
Clippings
Ask AI
Tab to switch back to notes
×
Ask me anything about this page or your journey.
Generating your article...
Searching the web and writing — this takes 10-20 seconds