Work / Richard Craven / Cravenverse
Richard Craven / Cravenverse
A publishing relationship that kept becoming a different kind of publishing problem.
How the work ran
Publishing work made the information model visible over time.
01
Making the books
The programme began with independent publishing and production work.
02
When a catalogue becomes a world
Recurring entities and relationships needed a navigable structure.
03
Reconstructing Bile
Source, extraction, correction and accepted edition remained distinct states.
04
From archive to corpus
Reviewable, evidence-linked records made a growing corpus workable.
05
Building for editorial judgement
The later tools make review and recovery more explicit, not optional.
A publishing relationship that kept becoming a different kind of publishing problem.
I started working with Richard Craven in late 2023.
At first the work was recognisable independent publishing: reading and research, manuscript work, typesetting, covers, illustrations, website material and preparing books for publication.
That work is still there.
But as more books, editions and archive material accumulated, a second problem appeared. Richard's work does not sit neatly in isolated titles. Characters recur. Places and terms return. Earlier material connects to later material. Drafts, scans, proofs, editions and website records do not all have the same authority.
The publishing programme gradually became a problem of how to work reliably across an interconnected literary world.
Current position
Ongoing paid publishing relationship with multiple published books and a live structured Cravenverse layer. Associated corpus, archive and editorial-knowledge tools are working internal systems and active R&D; they should not be described as public deployed products.
01 / Making the books
The relationship began with the books themselves.
The work included assessment, editorial preparation, proof correction, covers and end matter, typesetting, publication setup and the infrastructure around an independent catalogue.
By October 2024, Amoeba Dick and Pretty Poli had been published.
The catalogue continued to grow. Odour Issues and The Senseless Counterfeit followed in 2025, alongside collected verse, stories, ebooks and further editions.
That matters because the systems work described below did not begin as an abstract knowledge-management exercise.
It emerged while doing the publishing.
02 / When a catalogue becomes a world
A conventional book page can tell you about a title.
It is much less good at showing the things that pass between titles.
As the catalogue developed, the website began to represent characters, places, glossary terms, book-linked objects and relationships as structured information rather than leaving every connection embedded inside prose.
The Cravenverse layer grew around that problem.
WordPress and Pods became a way of giving the literary material a navigable shape: not replacing the books, but making some of their connections easier to inspect and maintain.
Several experiments followed, including graph and flat editorial views. Some worked better than others. The important development was the underlying shift:
the website was no longer only publishing pages about books; it was beginning to carry a model of the world across them.
03 / Reconstructing Bile
Bile presented a different kind of problem.
The useful source material existed across scans and older artefacts rather than as one clean, authoritative manuscript ready for ordinary editing.
The work therefore began with reconstruction.
Pages were converted through OCR, visually checked, reassembled and cleaned. Later production work repeatedly rebuilt and repaired the Scrivener structure when stale or broken project state interfered with a clean final manuscript.
That process made a distinction increasingly important:
- the source;
- what had been extracted from it;
- what had been corrected;
- what had been accepted for the current working edition.
Those are different states.
Keeping them different is what makes reconstruction trustworthy.
04 / From archive to corpus
By summer 2026, the programme had accumulated enough material that ordinary files, folders and memory were no longer a sufficient working interface.
The Extraction Workbench developed as a review environment around the corpus.
Material could be addressed more precisely, extracted into candidate records, linked back to evidence, checked for duplication and reviewed before becoming accepted state.
The system developed:
- paragraph-addressable source material;
- extraction queues;
- evidence-linked candidate records;
- duplicate detection;
- governed merge handling;
- canonical review;
- relationship data;
- corpus exports;
- isolated project workspaces;
- recovery and migration paths.
AI-assisted extraction became useful here, including local-model routes.
But extraction is not publication and a generated candidate is not automatically true.
The system is designed around that boundary.
05 / Building for editorial judgement
The interesting part of the later work is not that more of it became automated.
It is that review became more explicit.
The Workbench was reshaped around an editorial control-room interface. Duplicate records moved into governed merge handling. Canonical information and relationship data could be exported together. Changes gained backup and recovery paths.
Some experiments failed and were rolled back rather than quietly becoming part of the system.
The purpose is not to remove the editor.
It is to give the editor a better view of what the system thinks it knows, where that information came from, and what still needs judgement.
06 / What I did
My contribution spans the publishing and technical layers because, in this project, they became increasingly difficult to separate.
It has included:
- manuscript assessment and editorial work;
- proofing and production;
- cover, layout and illustration coordination;
- Scrivener and edition preparation;
- publication and metadata work;
- website and structured WordPress development;
- modelling characters, places, glossary terms and relationships;
- archive and source-material analysis;
- OCR reconstruction and review;
- corpus extraction;
- evidence-linked candidate/review workflows;
- canonicalisation;
- editorial knowledge infrastructure;
- internal tooling for review, export, migration and recovery.
That does not mean every tool is a public product.
The maturity of each layer matters.
Where it stands
Published work
Real public books and editions produced through the ongoing publishing relationship.
Cravenverse
A live structured WordPress/publishing layer representing books and parts of the literary world around them.
Bile reconstruction
A substantially working and reviewed reconstruction/production workflow.
Corpus and editorial-knowledge infrastructure
Working internal tools and active R&D used to extract, review, connect and maintain literary information.
Related but separate
The Literary Machine reuses lessons and technology from this work, but it is an independent public-domain publishing-system project. It is not Richard's live client corpus.
07 / What the project demonstrates
The unusual thing about this programme is not simply that publishing became technical.
Publishing already was technical.
The important transition was that repeated editorial work made the information model visible.
Books led to a catalogue. The catalogue led to recurring entities and relationships. Archive reconstruction exposed the need for source and review states. Corpus work made provenance, canonicalisation and recoverability explicit.
The software grew out of those editorial requirements.
That is why the publishing work and the knowledge-system work belong on the same page.
Development history
The approved private granular history currently contains:
- 229 meaningful developments
- derived from 552 retained source events
- 110 turning points
- 9 phases
- chronology beginning with located commercial evidence on 31 October 2023
The private history deliberately excludes or protects material such as manuscript prose, scans, private correspondence, live corpus databases and WordPress database/uploads where those are not appropriate evidence for general exposure.
A public development-history projection can be created from the approved history, but should be curated separately rather than exposing the private reader or Drive artefacts directly.
