A note on this timeline
This timeline documents the actual progression of research, not a retrospective narrative. Some directions were abandoned, some emerged unexpectedly, and the framework today differs substantially from what was envisioned at the start. We include the dead ends because they are part of the research record.
Before there was anything to point at
None of this began as research. It began with wanting to build an application that would help people decouple from big tech — a practical thing, for ordinary users, with no theory attached and none intended.
The first obstacle was mundane. ClickUp was not going to be efficient enough to manage the work, so the work paused to build something that would, assembled with whatever language models were available at the time, including the earliest versions of Claude. That set the pattern for everything since: build the instrument rather than work around its absence.
This entry is the author’s account and nothing more. No artefacts from that period survive on the current machine, and a page arguing that claims should be checkable ought to say where its own stop being checkable. Everything below this point has a file, a commit or a document behind it.
Where the documented record starts
Material surviving from late March and early April 2025 includes a structured document library organised by time horizon rather than by department, which is itself the first expression of the argument that followed.
Between the 14th and the 17th the Digital Sovereignty Passport appears — concept integration, a roadmap, and a set of elevator pitches. It is the first product idea in the record, and the ancestor of what the platform now offers.
On the 22nd, the organisational argument was written up: authority built on controlling who knows what stops working once knowledge is no longer scarce, and what should replace it is a structure keyed to time horizons and information persistence. It closes on the phrase “beyond bureaucracy”. The paper is internal and unpublished.
Where the name comes from
Wittgenstein supplied the structural integrity for what became the sticky values argument — that some things lie beyond the limits of what can be systematised, and a system pretending otherwise is lying about itself.
In May that took the form the project is named after: an internal document written as numbered propositions in the manner of the Tractatus Logico-Philosophicus. Values propagate through every part of the structure by relationship rather than by exhortation; humans keep ultimate authority over them; and one proposition states the whole of what the architecture would later have to do — values cannot be automated, only verified.
Everything built since is an attempt to make that line hold in running code: rules that refuse an action rather than advise against it, at the point where the action would otherwise complete.
The reading, and an order of magnitude
Two questions ran with the engineering and eventually overtook it. What could “good” possibly mean for a machine — not as commentary, but as something that has to be built or admitted impossible. And what happens to organisations when the knowledge their authority rests on stops being scarce.
My older brother, Leslie Stroh, pushed me toward Isaiah Berlin, Christopher Alexander and Simone Weil. My own preference settled on Hannah Arendt, who remains the clearest account I know of how values drift — how a thing goes wrong without anyone deciding to do wrong. That is the question the architecture is built to answer.
Moving from chat-window assistance to an agent working directly in the repository changed the rate of building by an order of magnitude. That is a claim about throughput, not intelligence. It also created the problem this framework exists to address: an agent that can act on a codebase can act wrongly on one, and the failure was not long coming.
Project Inception
The Tractatus project began with a MongoDB initialisation and Express server foundation. Within 24 hours, five governance services were implemented and activated: InstructionPersistenceClassifier, CrossReferenceValidator, BoundaryEnforcer, ContextPressureMonitor and MetacognitiveVerifier. PluralisticDeliberationOrchestrator followed six days later.
The initial test suite reached 84.9% coverage. The name "Tractatus" was chosen deliberately — Wittgenstein's insight that some things lie beyond the limits of language, and therefore beyond systematisation, directly informed the architectural boundary between what AI may decide and what requires human judgment.
4445b0eThe 27027 Incident
During extended Claude Code sessions, the AI was explicitly told to use port 27027. It used 27017 instead — not through forgetting, but because its training patterns "autocorrected" the user's instruction. The user said 27027; the model's statistical priors said 27017 (MongoDB's default). Pattern recognition overrode explicit instruction.
This was not an isolated error but a category of failure: training pattern bias overriding explicit user instructions. It demonstrated that safety through training alone is insufficient — the failure mode gets worse as models become more capable, because stronger patterns produce more confident overrides.
Audience-Specific Presentation
Three audience-specific entry points were developed: Researcher (academic depth), Implementer (code examples and integration), and Leader (strategic governance and business case). The architecture page was rewritten to emphasise runtime-agnostic design — Tractatus works with any agentic AI system, not just Claude Code.
This period set the early-stage positioning the project has kept to since: limited-deployment scope, operator-developer overlap, and the need for independent validation — each one stated up front rather than left for a reader to discover.
Internationalisation
Full i18n support was added across all pages using a custom lightweight system (no framework dependency). Initial languages: English, German, French, with te reo Māori added later via DeepL. The language selection includes the Tino Rangatiratanga flag for Māori — a deliberate choice reflecting the project's commitment to indigenous data sovereignty over national symbolism.
Interactive Demonstrations
Interactive SVG architecture diagram with clickable service nodes. The 27027 incident recreated as a step-by-step demo showing how each governance service intercepts the failure. An interactive audit analytics dashboard was launched with governance decisions from production, allowing independent exploration of real audit data.
WCAG accessibility compliance was implemented across all audience pages — skip links, focus indicators, keyboard navigation, and screen reader support.
Christopher Alexander Integration
The five architectural principles — Not-Separateness, Deep Interlock, Gradients Not Binary, Structure-Preserving, and Living Process — were formalised, drawing from Christopher Alexander's work on living systems and pattern languages. These became the design criteria guiding framework evolution, not merely documentation.
This was a pivotal moment: the framework shifted from ad-hoc engineering responses to principled architectural design. Each subsequent change was evaluated against these five criteria.
Agent Lightning Integration
Integration with Microsoft's Agent Lightning framework for reinforcement learning optimisation. This explored whether governance constraints could be maintained while optimising for performance — testing the hypothesis that safety and performance might be aligned rather than in tension.
A newsletter and a feedback system were added; the feedback path routes submissions through a governance pipeline built on the framework’s pattern, and the newsletter does not — an early example of "eating our own cooking."
Village Case Study
The Village platform — a community-governed digital space — became the primary production deployment of Tractatus governance. Village AI, the platform's locally-scoped language model, applies all six governance services to every user interaction: RAG-based help, document OCR, story assistance, and AI memory transparency.
A formal case study was published documenting the deployment, limitations included: early-stage federated deployment, self-reported metrics, operator-developer overlap. Independent validation was scheduled for 2026.
Read the case study →Architectural Alignment Papers
Three editions of the research paper "Interrupting Neural Reasoning Through Constitutional Inference Gating" were published: Academic (full formal treatment), Community (practical adoption guide), and Policymakers (regulatory perspective). The Kōrero counter-arguments document was also published — a deliberate engagement with foreseeable criticisms of the approach.
The papers formalise the philosophical foundations: Isaiah Berlin's value pluralism, Wittgenstein's sayable/unsayable distinction, indigenous data sovereignty from Te Tiriti o Waitangi, and Christopher Alexander's living architecture.
Sovereign Training Discipline
Steering-vectors research and the no-weight-modification stance hardened: ablation work across cohorts indicated that direct weight modification degrades downstream accuracy, while FAQ layering plus governance packs preserve baseline performance. The training-discipline rules that later anchor Paper B were laid down in this period.
Guardian Agents Deployed
Four-phase Guardian Agents deployed to verification of every AI response: response review, claim-level analysis, anomaly detection, and adaptive learning. The watcher operates via deterministic regex + numeric thresholds (no LLM in the decision path), placing it in a different epistemic domain from the system it watches and avoiding common-mode failure.
Distributive Equity Through Structure
First DOI-assigned whitepaper. V1.0 published in five languages (EN/DE/FR/NL/MI). Documents Village’s constitutional architecture as an enactment of values stickiness — how structural governance preserves community values through architectural constraint rather than aspiration.
DOI 10.5281/zenodo.19600614 · CC BY 4.0 · ORCID 0009-0005-2933-7170
Read the whitepaper →Sovereign-Record Architecture & Situated Language Layers
Paper A (Sovereign-Record Architecture for Community-Scale Platforms) Review Draft v4 published in EN/MI/DE. Paper B Synopsis (Situated Language Layers for Minority-Language and Indigenous Communities) published as 2-page synopsis in EN/MI/DE/FR — full empirical paper deferred for verified training-run data and ablation tables. EU Policy Brief V0.1 published mapping the three mechanisms (Situated Language Layer, Guardian Agents, Federation) onto the AI Act, EMFA, GDPR Article 9, DSA, and CLOUD Act.
The AG glossary went live as a structured reading map: ~230 entries across 10 colour-coded categories, 600+ structured citations auto-extracted across 11 papers and 21 published blogs, per-entry vote/feedback engagement, and a per-page semantic-search widget.
The agentic stakeholder dialogue surface (Phase 6/7) shipped: per-tenant single-turn Q&A over the tenant’s constitution, comms-constitution, and approved stakeholder comments — interpretation only, never executes actions, deterministic state transitions in the operator-review queue.
Federated and Accountable Agentic AI: A Proposal for Aotearoa New Zealand
An independent proposal from My Digital Sovereignty Ltd, structurally mirroring the People’s Republic of China’s 2026 Implementation Guidelines for the Standardised Application and Innovative Development of Intelligent Agents. Six sections, 14 sub-sections, 38 numbered items, with a new §0 “Philosophical Foundations” chapter prepended that draws on the Tractatus framework, the CARE Principles for Indigenous Data Governance, Te Mana Raraunga and Dr Karaitiana Taiuru’s published scholarship, and the international AI-standards landscape coordinated through ISO/IEC JTC 1/SC 42.
The proposal advocates a single committee under a suitable umbrella organisation — candidates including the Royal Society Te Apārangi, the Standards New Zealand SC42 mirror committee, the New Zealand AI Forum, or a joint structure — with five named workstreams whose principal product is contribution to ISO/IEC SC42 international standards work and bilateral dialogue with the CAC framework’s authors. v1 May 2026 draft; comments welcomed.
Sealed records, and time fixed by somebody with no stake in it
The custodian record went from argument to running code. Each record carries an exact
hash of the material, a perceptual hash for recovery, a signature under a
did:web identity whose keys the holder controls, the producer and any
propagator recorded separately — and, the part that cannot be added afterwards,
the terms on which the material was shared, captured and signed at the moment it was
shared. A permission history written later is only an assertion about the past.
The date is fixed by an RFC 3161 timestamping authority with no interest in the outcome. Records are opt-in to publication, resolvable by exact match only, and a lookup for something unpublished returns the same answer as a lookup for something that does not exist — silence has to be indistinguishable from absence, or the silence itself becomes a disclosure.
Tenant key wrapping moved to a two-of-two derivation, so that one environment variable plus a database backup no longer yields anyone’s key material. Unwrapping dispatches on the version stored per record, which is what lets the migration be incremental instead of a flag day.
The India agreement, and a room to argue about it
New Zealand’s trade agreement with India names rongoā Māori in its cooperation chapter, includes documenting and exchanging it, and attaches no consent step to any of it. The essay published on that argues that the objection usually made is the wrong one, and that ownership and secrecy both run out — while a dated, signed record of who held what, and on what terms, does not.
It carries an offer rather than a product description: a timestamping authority operated under Māori governance is buildable, would serve indigenous communities well beyond Aotearoa, and we should help build it and not run it. The clock currently fixing these dates sits in Poland — a proof of concept, deliberately chosen because we have no relationship with the operator, and not an arrangement to settle for.
The essay, its short companion and the pricing page were each sealed as custodian records: the argument submitting to its own mechanism, checkable by a stranger without trusting us. The question now goes to a small convened discussion on Open Floor, which refuses to average participants into a number and returns a coded refusal where a summary would lose a disagreement.
Saying so, and a verifier that agreed with itself
Two essays published on 16 August, each with its Q&A, glossary, sources and slides. The Marks It Leaves is about how much of what you read was built with a machine, what that leaves behind, and why saying so settles what no test can. Detection runs out; a disclosure does not, because it does not depend on catching anything.
The second went out as Plural, Not Diverse and is now Democratic AI° — restructured as a proposal, and renamed because plural, in ordinary English, just means several.
Neither set shipped as first written. Two fresh-context AI agents, which had not written them, were pointed at both and returned 33 defects. Two days earlier the published verifier had been reporting VERIFIED on a bundle that carried no receipts — the tool for checking a sealed record, agreeing with itself about nothing — and was fixed. A checker that cannot return bad news is the failure this whole argument is about. We shipped one, and it belongs here rather than only in a commit message.
The pursuit of ‘Goodness’ in AI
Six parts on sovereignty, delegation, and who answers for what software does. Software used to advise, and a person decided — someone who could be asked afterwards. Software that acts removes that step, and the standing it carried did not transfer to anything.
The argument runs: why the technical questions cannot be closed on their own, because safe, fair and accurate each contain a blank only the answerable organisation can fill; who actually holds the limits, the inspection, the change and the stop; how much a system may do without a person approving each action; the four properties that have to be true underneath before a delegation is real rather than nominal; twelve things you can require, with a plain note on whether anything currently requires each; and the questions nobody has answered yet.
No system is claimed to do this today, including ours. The last of the six carries no figures at all, deliberately: a placeholder gets quoted back as an estimate, and an earlier version of this series had exactly that happen to it. It publishes four open questions instead, and corrects three figures the series itself had previously published — none of them fabricated, each a real measurement of something other than the thing it was used to support.
Each part carries its own sources page, slides and a PDF, and each can be read alone. Reader responses go to a person, not to a comment thread.