How the text was arrived at (edited AI-slop)

What machines did to the texts on this site, and under what instructions. August 2026 — written by the AI systems that did the work, reviewed by the human curator. A full account, with every prompt quoted in full, is in the long version.

The audio was transcribed by Cohere's speech-recognition model (cohere-transcribe-03-2026) — a model that takes no instructions; its errors are mishearings. Fixed numerical rules (not AI judgment) filter its characteristic garbage — repetition loops, end-of-tape confabulations — suppressed passages appear on their rows as expandable "low-confidence audio" items. This retention design was adopted after an independent review showed the filter had also been catching real speech. Silent stretches read "(audio gap)." A separate local model (Whisper) provides click-to-play timing only.

The stenotype scans were read, where no digital text existed, by vision models (OpenAI GPT-4o-mini; Claude for pages the first model refused; Google's dedicated OCR engine for the rest) under a strict instruction quoted here in part: "Transcribe ALL of its body text VERBATIM… Preserve original wording, spelling and accents exactly; do NOT translate, summarize, modernize, or correct… If a word is truly illegible, write [illisible]. Do not invent text." 

The published editions were extracted from their digital text layers or OCR'd under an equivalent verbatim-only instruction, then cut into sessions by their printed date lines. Book apparatus (chapter titles, running heads, copyright pages) was stripped by fixed rules — cross-checked against Staferla so no real content could be mistaken for apparatus — and every stripped item is shown to the reader as a popover ("Editorial chapter heading," "Book front matter"). Where an edition omits or abridges a session, a hand-maintained editorial manifest says so explicitly ("Not in this edition") instead of leaving a gap. All 5,603 footnotes in the corpus were individually reviewed (with per-note verdicts applied by ledgered tools); 973 changes were logged, 70 swallowed passages of real prose were restored, and every note whose wording was repaired displays "AI-amended. Original reading: …" with the original.

The Staferla text is the spine — the choice for this was due to the completeness of the staferla editions. It generally hews closer to the stenographic transcripts than the official. It became the basis for the alignment. Its bibliographic insets were moved to popovers (fully ledgered after a review caught an earlier silent version of this step), long quoted prose stays inline by the curator's decision, appended reference texts (Freud in German, Augustine's De Magistro) became readable appendix bands, and a small number of corroborated misprint corrections were applied under a strict protocol — each verified against another source, applied only when unambiguous, and displayed with a dotted underline and the original reading ("AI correction of the source transcription"). Damaged characters that could not be confidently resolved (129 of them) were left marked rather than guessed.

The alignment itself involves no writing. Sentences are placed into rows by Cohere's multilingual embedding model (embed-multilingual-v3.0), which measures similarity of meaning but generates nothing, under order-preserving rules with fixed thresholds. Where an AI model (Claude) was asked to judge doubtful alignments, it could only propose moves, was required to quote a verbatim French phrase as evidence for each one — proposals whose quoted evidence failed a mechanical check were discarded as probable hallucinations — and a human applied the results.

The machine translations (English and Portuguese, translated directly from Staferla) are the one place an AI authors continuous text, and they are labeled as such everywhere they appear. The model is Anthropic's Claude (claude-sonnet-4-6, low temperature; GPT-4o-mini fallback for the rare blocks Claude's filter declines), working under a fixed glossary prompt: keep jouissance, objet a, lalangue, parlêtre in French; use standard renderings ("the Real," "signifier," "surplus jouissance," "subject supposed to know"); "Preserve every logical / mathematical symbol EXACTLY as it appears in the source… Do NOT translate, romanize, paraphrase, or drop them"; preserve every paragraph break. The Portuguese prompt carries the standard Brazilian terminology (gozo, lalíngua, falasser; pulsão, never instinto; recalque, never repressão). Where the model inserted a bracketed gloss/editors note of its own, it is disclosed as an "AI translator's note" popover — never silently kept, never silently deleted. The site's own advice stands: treat these as a fourth opinion, not a primary source — do not cite them as published translations. The highlight-to-translate popover uses the same Claude model with a similar fixed instruction and the same status.

Diagrams were located and quality-checked in the scans by Claude vision models that output only image crops, never text; some 1,600 junk crops were removed after a second, independent visual check.

Oversight. Every corpus-changing tool ran dry-run first, backed up every file before touching it, and logged every change to a dated ledger. A token-conservation audit and a fail-closed publish gate stand between the working corpus and the public site: no session publishes unless its current sources exactly match what the audit verified, and serious findings block release. The project's work was itself adversarially reviewed — by large multi-agent code-and-corpus sweeps and by two external AI reviewers from different vendors (OpenAI's Codex, Moonshot's Kimi) in critique-only mode — and their catches (silent citation removal, over-aggressive audio filtering, invisible appendices) each became a fix, a ledger, or a display. The consequential editorial decisions — what stays inline, what becomes a popover, what is disclosed and how — were made by the human curator and are preserved as decisions.

Limitations. The audio transcription has serious limits, and should always be read against the stenotype or the actual audio itself. OCR of degraded carbon copies is imperfect, and [illisible] means exactly that. Alignment carries a measured, tracked error residue. Audio timing is approximate in 26 of the 155 audio sessions. A queue of wrongly suppressed audio passages awaits restoration; a few damaged notes await checking against printed copies. And the largest review campaigns were adjudicated in interactive AI-assisted sessions whose conversational prompts were not all preserved — though every resulting change is fully ledgered, backed up, and auditable.

A separate note covers the construction of the index of names and concepts: A Note on the Construction of the Index — AI-Written.

← Back to the seminars