The audio was transcribed by Cohere's speech-recognition
model (cohere-transcribe-03-2026) — a model that takes no
instructions; its errors are mishearings. Fixed numerical rules (not AI
judgment) filter its characteristic garbage — repetition loops,
end-of-tape confabulations — suppressed passages appear on their
rows as expandable "low-confidence audio" items. This retention design
was adopted after an independent review showed the filter had also been
catching real speech. Silent stretches read "(audio gap)." A separate
local model (Whisper) provides click-to-play timing only.
The stenotype scans were read, where no digital text existed, by vision models (OpenAI GPT-4o-mini; Claude for pages the first model refused; Google's dedicated OCR engine for the rest) under a strict instruction quoted here in part: "Transcribe ALL of its body text VERBATIM… Preserve original wording, spelling and accents exactly; do NOT translate, summarize, modernize, or correct… If a word is truly illegible, write [illisible]. Do not invent text."
The published editions were extracted from their digital text layers or OCR'd under an equivalent verbatim-only instruction, then cut into sessions by their printed date lines. Book apparatus (chapter titles, running heads, copyright pages) was stripped by fixed rules — cross-checked against Staferla so no real content could be mistaken for apparatus — and every stripped item is shown to the reader as a popover ("Editorial chapter heading," "Book front matter"). Where an edition omits or abridges a session, a hand-maintained editorial manifest says so explicitly ("Not in this edition") instead of leaving a gap. All 5,603 footnotes in the corpus were individually reviewed (with per-note verdicts applied by ledgered tools); 973 changes were logged, 70 swallowed passages of real prose were restored, and every note whose wording was repaired displays "AI-amended. Original reading: …" with the original.
The Staferla text is the spine — the choice for this was due to the completeness of the staferla editions. It generally hews closer to the stenographic transcripts than the official. It became the basis for the alignment. Its bibliographic insets were moved to popovers (fully ledgered after a review caught an earlier silent version of this step), long quoted prose stays inline by the curator's decision, appended reference texts (Freud in German, Augustine's De Magistro) became readable appendix bands, and a small number of corroborated misprint corrections were applied under a strict protocol — each verified against another source, applied only when unambiguous, and displayed with a dotted underline and the original reading ("AI correction of the source transcription"). Damaged characters that could not be confidently resolved (129 of them) were left marked rather than guessed.
The alignment itself involves no writing. Sentences are
placed into rows by Cohere's multilingual embedding model
(embed-multilingual-v3.0), which measures similarity of
meaning but generates nothing, under order-preserving rules with fixed
thresholds. Where an AI model (Claude) was asked to judge doubtful
alignments, it could only propose moves, was required to quote a
verbatim French phrase as evidence for each one — proposals whose quoted
evidence failed a mechanical check were discarded as probable
hallucinations — and a human applied the results.
The machine translations (English and Portuguese,
translated directly from Staferla) are the one place an AI authors
continuous text, and they are labeled as such everywhere they appear. The
model is Anthropic's Claude (claude-sonnet-4-6, low
temperature; GPT-4o-mini fallback for the rare blocks Claude's filter
declines), working under a fixed glossary prompt: keep jouissance,
objet a, lalangue, parlêtre in French; use
standard renderings ("the Real," "signifier," "surplus jouissance,"
"subject supposed to know"); "Preserve every logical / mathematical
symbol EXACTLY as it appears in the source… Do NOT translate, romanize,
paraphrase, or drop them"; preserve every paragraph break. The
Portuguese prompt carries the standard Brazilian terminology
(gozo, lalíngua, falasser; pulsão,
never instinto; recalque, never repressão).
Where the model inserted a bracketed gloss/editors note of its own, it is disclosed as
an "AI translator's note" popover — never silently kept, never silently
deleted. The site's own advice stands: treat these as a fourth
opinion, not a primary source — do not cite them as published
translations. The highlight-to-translate popover uses the same
Claude model with a similar fixed instruction and the same status.
Diagrams were located and quality-checked in the scans by Claude vision models that output only image crops, never text; some 1,600 junk crops were removed after a second, independent visual check.
Oversight. Every corpus-changing tool ran dry-run first, backed up every file before touching it, and logged every change to a dated ledger. A token-conservation audit and a fail-closed publish gate stand between the working corpus and the public site: no session publishes unless its current sources exactly match what the audit verified, and serious findings block release. The project's work was itself adversarially reviewed — by large multi-agent code-and-corpus sweeps and by two external AI reviewers from different vendors (OpenAI's Codex, Moonshot's Kimi) in critique-only mode — and their catches (silent citation removal, over-aggressive audio filtering, invisible appendices) each became a fix, a ledger, or a display. The consequential editorial decisions — what stays inline, what becomes a popover, what is disclosed and how — were made by the human curator and are preserved as decisions.
Limitations. The audio transcription has serious limits, and should always be read against the stenotype or the actual audio itself. OCR of degraded carbon copies is
imperfect, and [illisible] means exactly that. Alignment
carries a measured, tracked error residue. Audio timing is approximate in
26 of the 155 audio sessions. A queue of wrongly suppressed audio
passages awaits restoration; a few damaged notes await checking against
printed copies. And the largest review campaigns were adjudicated in
interactive AI-assisted sessions whose conversational prompts were not
all preserved — though every resulting change is fully ledgered, backed
up, and auditable.
A separate note covers the construction of the index of names and concepts: A Note on the Construction of the Index — AI-Written.
← Back to the seminars