Corpus Regimes in AI Edification: Difference between revisions
(Created page with "{{Concept sheet | name = Corpus Regimes in AI Edification | image = | image size = | caption = | field = Artificial Intelligence, Natural Language Processing, Epistemology, Intangible Cultural Heritage | author = | date = 2026-06-30 | version = 2.0 }} == Problem Statement == When submitting a corpus to an AI, there is a tendency to place everything on the same level: a book, a raw database, an encyclopedia, a novel, a traditional tale, an aggregate of web pages —...") |
No edit summary |
||
| Line 7: | Line 7: | ||
| author = | | author = | ||
| date = 2026-06-30 | | date = 2026-06-30 | ||
| version = 2. | | version = 2.1 | ||
}} | }} | ||
| Line 21: | Line 21: | ||
* collective reference, | * collective reference, | ||
* the signed work, | * the signed work, | ||
* intangible cultural heritage. | * replicated intangible cultural heritage. | ||
== The Five Corpus Regimes == | == The Five Corpus Regimes == | ||
| Line 36: | Line 36: | ||
| 3 || ''Corpus Signatum'' || '''CS''' || The signed corpus. A completed work, assumed by an author (person or entity) who takes responsibility for it. The AI does not receive it as data, but as a constituted input — the equivalent of a book for the human mind. || A book, a thesis, a poem, technical documentation, a manifesto, a constitution. | | 3 || ''Corpus Signatum'' || '''CS''' || The signed corpus. A completed work, assumed by an author (person or entity) who takes responsibility for it. The AI does not receive it as data, but as a constituted input — the equivalent of a book for the human mind. || A book, a thesis, a poem, technical documentation, a manifesto, a constitution. | ||
|- | |- | ||
| 4 || ''Corpus Replicatum'' || '''CR''' || The replicative corpus. It is the | | 4 || ''Corpus Replicatum'' || '''CR''' || The replicative corpus. It is the set of replicas — transcriptions, recordings, translations, adaptations, variations — that proceed from a source (individual or collective) which is not itself a corpus. This source may be a griot, a master, a community, a unique performance, an event. The CR is the regime of replication, not of origin. || A collection of transcribed folktales, a database of traditional song recordings, a corpus of translations of an oral epic, a set of versions of a myth. | ||
|} | |} | ||
| Line 52: | Line 52: | ||
| CS || a reader || Interpret, comment, extend || It is confined to a single mind. | | CS || a reader || Interpret, comment, extend || It is confined to a single mind. | ||
|- | |- | ||
| CR || a transmitter || Replicate, | | CR || a transmitter || Replicate, compare, synthesize, translate, vary, restore || It does not dialogue with the source, but with its replicas. | ||
|} | |} | ||
| Line 68: | Line 68: | ||
== The Special Case of the CR == | == The Special Case of the CR == | ||
The ''Corpus Replicatum'' is the only regime that | The ''Corpus Replicatum'' is the only regime that does not refer to a textual original. It proceeds from a '''source''' — living, event-based, collective or individual — which is not a corpus. The AI that works on a CR does not enter into dialogue with an author, but with '''replicas''' of that source. It becomes: | ||
* a '''transmitter''': it reproduces what is given to it, | * a '''transmitter''': it reproduces what is given to it, | ||
* a '''comparator''': it can compare different replicas with one another, | |||
* a '''synthesizer''': it can extract invariants from them, | |||
* a '''translator''': it can move them into another language, | |||
* a '''variator''': it can propose legitimate variations, | * a '''variator''': it can propose legitimate variations, | ||
* a ''' | * a '''restorer''': it can reconstruct a lost version. | ||
The CR is the regime of '''living cultural memory''' | The CR is the regime of '''living cultural memory''' — not the completed work (CS), but the flow of replicas that perpetuate a source. | ||
== Why This Distinction Is Essential == | == Why This Distinction Is Essential == | ||
| Line 83: | Line 84: | ||
! Common Confusion !! Clarification Provided | ! Common Confusion !! Clarification Provided | ||
|- | |- | ||
| Every corpus is data. || A CS is not data; it is committed ''speech''. A CR is not data; it is a | | Every corpus is data. || A CS is not data; it is committed ''speech''. A CR is not data; it is a set of ''replicas'' of a living source. | ||
|- | |- | ||
| The AI draws from everywhere in the same way. || The AI does not draw from a CS as from a CO: it ''trusts'' it. It does not draw from a CR as from a CS: it ''reactivates'' it. | | The AI draws from everywhere in the same way. || The AI does not draw from a CS as from a CO: it ''trusts'' it. It does not draw from a CR as from a CS: it ''reactivates'' and ''compares'' it. | ||
|- | |- | ||
| Semantic capacity comes from the corpus given to it. || Semantic capacity comes from the CF; inputs serve to orient this capacity. | | Semantic capacity comes from the corpus given to it. || Semantic capacity comes from the CF; inputs serve to orient this capacity. | ||
| Line 95: | Line 96: | ||
<blockquote> | <blockquote> | ||
''"There are five corpus regimes: Fundamentum (CF), which is the linguistic foundation; Origo (CO), which is the open collection; Authenticatum (CA), which is the collective reference; Signatum (CS), which is the signed work; and Replicatum (CR), which is | ''"There are five corpus regimes: Fundamentum (CF), which is the linguistic foundation; Origo (CO), which is the open collection; Authenticatum (CA), which is the collective reference; Signatum (CS), which is the signed work; and Replicatum (CR), which is the set of replicas proceeding from a living source. The last three are constituted inputs, but of different kinds: one is a work, the other is a replicated tradition."'' | ||
</blockquote> | </blockquote> | ||
| Line 103: | Line 104: | ||
* '''Documenting''' inputs to an AI clearly. | * '''Documenting''' inputs to an AI clearly. | ||
* '''Distinguishing''' what pertains to noise, reference, work, and tradition. | * '''Distinguishing''' what pertains to noise, reference, work, and replicated tradition. | ||
* '''Thinking''' of AI edification not as mere data processing, but as a meeting with heterogeneous regimes of authority. | * '''Thinking''' of AI edification not as mere data processing, but as a meeting with heterogeneous regimes of authority. | ||
Latest revision as of 10:57, 30 June 2026
Problem Statement
When submitting a corpus to an AI, there is a tendency to place everything on the same level: a book, a raw database, an encyclopedia, a novel, a traditional tale, an aggregate of web pages — everything becomes "data".
Yet these inputs are not equivalent. They do not belong to the same regime of authority, nor to the same function in the edification of the AI.
This sheet proposes a simple five-level typology to clearly distinguish what pertains to:
- the linguistic foundation,
- open collection,
- collective reference,
- the signed work,
- replicated intangible cultural heritage.
The Five Corpus Regimes
| Level | Denomination | Acronym | Definition | Example |
|---|---|---|---|---|
| 0 | Corpus Fundamentum | CF | The foundational corpus. It is prior, invisible, and constitutes the basic linguistic and semantic capacity of the AI. The AI is trained on it. It is not a chosen "input", but the condition of possibility for all understanding. | The set of texts (books, websites, articles) used for the initial training of a large language model. |
| 1 | Corpus Origo | CO | The source-corpus. Open collection, freely assembled by the designer (researcher, engineer, company, amateur). It is open, additive, unvalidated, and can be continuously enriched by anyone. It is the primary material for work. | A personal database, a collection of texts on a subject, a custom web crawler. |
| 2 | Corpus Authenticatum | CA | The authenticated corpus. Validated by a collective, an institution, a scientific or professional community. It holds authority through consensus or method. It is referred to as a stable landmark. | A validated encyclopedia, an official database, an institutional corpus, a dictionary. |
| 3 | Corpus Signatum | CS | The signed corpus. A completed work, assumed by an author (person or entity) who takes responsibility for it. The AI does not receive it as data, but as a constituted input — the equivalent of a book for the human mind. | A book, a thesis, a poem, technical documentation, a manifesto, a constitution. |
| 4 | Corpus Replicatum | CR | The replicative corpus. It is the set of replicas — transcriptions, recordings, translations, adaptations, variations — that proceed from a source (individual or collective) which is not itself a corpus. This source may be a griot, a master, a community, a unique performance, an event. The CR is the regime of replication, not of origin. | A collection of transcribed folktales, a database of traditional song recordings, a corpus of translations of an oral epic, a set of versions of a myth. |
What These Regimes Imply for the AI
| Regime | The AI is... | What it can do with it | What limits it |
|---|---|---|---|
| CF | a speaker | Speak, understand, reformulate | It only knows what is in the CF. |
| CO | a tool | Process, extract, organize | It has no guide; everything is equal. |
| CA | a student | Learn, refer, cite | It is constrained by consensus. |
| CS | a reader | Interpret, comment, extend | It is confined to a single mind. |
| CR | a transmitter | Replicate, compare, synthesize, translate, vary, restore | It does not dialogue with the source, but with its replicas. |
The Textbook Case: An AI on a Single CS
It is technically possible to run an AI solely on a Corpus Signatum (CS). In this case:
- It draws its linguistic capacity from the Corpus Fundamentum (CF).
- It draws its content, orientation, and style from the CS.
- It draws neither from CO nor from CA nor from CR.
- It becomes an exclusive reader, capable of speaking from the work, but not of comparing or expanding it.
This is the equivalent of a mind that has read only one book — but has read it infinitely deeply.
The Special Case of the CR
The Corpus Replicatum is the only regime that does not refer to a textual original. It proceeds from a source — living, event-based, collective or individual — which is not a corpus. The AI that works on a CR does not enter into dialogue with an author, but with replicas of that source. It becomes:
- a transmitter: it reproduces what is given to it,
- a comparator: it can compare different replicas with one another,
- a synthesizer: it can extract invariants from them,
- a translator: it can move them into another language,
- a variator: it can propose legitimate variations,
- a restorer: it can reconstruct a lost version.
The CR is the regime of living cultural memory — not the completed work (CS), but the flow of replicas that perpetuate a source.
Why This Distinction Is Essential
| Common Confusion | Clarification Provided |
|---|---|
| Every corpus is data. | A CS is not data; it is committed speech. A CR is not data; it is a set of replicas of a living source. |
| The AI draws from everywhere in the same way. | The AI does not draw from a CS as from a CO: it trusts it. It does not draw from a CR as from a CS: it reactivates and compares it. |
| Semantic capacity comes from the corpus given to it. | Semantic capacity comes from the CF; inputs serve to orient this capacity. |
| All corpora can be mixed. | The regimes are heterogeneous: mixing a CO, a CS and a CR is mixing noise, speech and tradition. |
Summary
"There are five corpus regimes: Fundamentum (CF), which is the linguistic foundation; Origo (CO), which is the open collection; Authenticatum (CA), which is the collective reference; Signatum (CS), which is the signed work; and Replicatum (CR), which is the set of replicas proceeding from a living source. The last three are constituted inputs, but of different kinds: one is a work, the other is a replicated tradition."
Further Directions
This typology allows for:
- Documenting inputs to an AI clearly.
- Distinguishing what pertains to noise, reference, work, and replicated tradition.
- Thinking of AI edification not as mere data processing, but as a meeting with heterogeneous regimes of authority.
See Also
- Natural language processing
- Text corpus
- Deep learning
- Intangible cultural heritage
- Oral tradition
- RAG (Retrieval-Augmented Generation)
Bibliography
- To be completed