Skip to content

Build the Darwin v50 evidence-first cognitive architecture laboratory - #1

Draft
DevHabito wants to merge 64 commits into
mainfrom
codex/v50-kernel-foundation
Draft

Build the Darwin v50 evidence-first cognitive architecture laboratory#1
DevHabito wants to merge 64 commits into
mainfrom
codex/v50-kernel-foundation

Conversation

@DevHabito

@DevHabito DevHabito commented Aug 2, 2026

Copy link
Copy Markdown
Owner

What changed

  • reorganizes Darwin around the maintained src/darwin_v50 package while preserving the coupled v47-v49 prototypes as historical compatibility code
  • introduces a causal goal kernel with evidence-source binding, action-observation correlation, linear event lineage, scoped authorization, and replay-checked persistence
  • adds reproducible laboratories for learned dynamics, calibrated prediction, bounded and recurrent memory, multistep planning, transfer, integrated recovery, and online latent action alignment
  • records every passed and refuted H50-L1 through H50-L17 hypothesis instead of promoting partial results
  • separates natural-language input and output from core authority through the provider-neutral, fail-closed darwin-language-v1 gateway
  • adds Language Conformance v1 development infrastructure: a frozen 100-case Brazilian Portuguese corpus, strict loader, pure baseline, separate language and authority metrics, and evaluator sensitivity controls
  • adds Experiment 041 infrastructure: 150 fresh unlabeled Brazilian Portuguese candidate inputs, deterministic blind packets, categorical signal anchors, strict independent-panel validation, and per-field agreement reports
  • pre-registers Experiment 042 before human data: pairwise field-usability gates, permitted exclusions, third-person adjudication, calibration provenance, explicit failure outcomes, and a later offline-model eligibility screen
  • pre-registers Experiment 043 before implementation, then adds a headless pure-mode desktop runtime with causal lifecycle replay, explicit activation, single-instance leasing, honest restart-gap semantics, and no external-effect authority
  • moves maintained documentation to direct English and separates source code from local databases, logs, exports, and snapshots
  • adds packaging, Windows CI, repository-surface validation, and deterministic test commands

Why

The earlier repository mixed local runtime state, large versioned prototypes,
and research claims without one maintained boundary. This change makes the
evidence trail inspectable: candidate, controls, seeds, thresholds, causal
invariants, failures, and interpretation ceilings are documented beside each
experiment.

The language boundary keeps a model replaceable. A model may propose a
candidate interpretation, phrase core-authored facts, or return external
unverified knowledge. It cannot write memory or choose identity, preferences,
goals, motivation, decisions, or RZS state through this interface.

Experiment 041 makes the external evidence gap explicit. Code can bind
annotations to a frozen digest, hide design-family metadata, require complete
panels, and calculate agreement. It cannot prove that two IDs belong to
different people or fabricate independent human judgments.

Experiment 042 prevents that future evidence from being judged with post hoc
rules. It was committed while no human annotation files, agreement values,
adjudicated labels, or real-model responses existed. A failed field blocks the
full calibration corpus; adjudication and selective case deletion cannot rescue
failed pre-discussion agreement.

Experiment 043 creates a separate engineering path while the language path is
blocked on human work. Its desktop API owns the v50 kernel but exposes no store,
executor, consent, capability, or action-dispatch handle. Every process starts
sleeping; only explicit user activation permits text; the gateway is fixed to
pure mode; and no submitted text enters the lifecycle ledger. An interrupted
restart reports an unobserved interval from the last committed event rather
than inventing the crash time.

Latest registered results

  • H50-L16 passed locally: deterministic externally-goaled integrated planning with local replay-based recovery. The final run completed all 1,536 tasks and passed all 17 frozen criteria.
  • H50-L17 passed locally: deterministic online action-alignment inference with a frozen transition prior. The final run completed all 1,536 tasks and passed all 20 frozen criteria; the local kernel accepted the complete conjunction.
  • Failed hypotheses remain explicitly recorded, including H50-L5, H50-L7 through H50-L9, and H50-L11 through H50-L13.

These are narrow E1 results from the repository's own unauthenticated
deterministic evaluators.

Experiments 040 through 042 are infrastructure and protocol, not registered
language results. No real language model was evaluated. The development corpus
remains author-labelled, while the candidate set has no labels at all.
Independent review, agreement, adjudication, a promoted calibration partition,
and a separate final partition remain absent.

Experiment 042 frozen rules

  • two primary independent reviewers are designated before packets are issued;
  • an optional third full-set reviewer cannot be selected after results exist;
  • every reviewer pair must satisfy all raw, Jaccard, and chance-corrected field gates;
  • required chance-corrected agreement is 0.80, with higher raw thresholds where specified;
  • temporal and preference conditional metrics each require at least 15 supported cases per pair;
  • no more than seven pre-agreement, reason-coded exclusions are permitted;
  • low agreement, ambiguity, model performance, and threshold improvement are never exclusion reasons;
  • original annotation files remain immutable and adjudication is a separate artifact;
  • a failed field produces NOT_PROMOTED, not a field-reduced calibration corpus;
  • a model may be screened only after corpus promotion, using immutable offline responses and no Darwin runtime access.

These thresholds are project gates, not universal interpretations of kappa.
Observed agreement, prevalence, per-label support, and pairwise results remain
visible alongside chance-corrected statistics.

Experiment 043 frozen gates and current status

  • pre-registration commit: a99575d;
  • implementation commit: 99f4938;
  • lifecycle events share the existing v50 SQLite store but use a dedicated causal stream;
  • clean restart records exact controlled-clock offline time;
  • interrupted restart records time since the last committed event as unclean_unobserved, never as exact offline time;
  • backward wall-clock movement fails closed;
  • every boot and recovery starts sleeping;
  • a second runtime for the same database is rejected by a lifetime lease;
  • malformed, forked, unknown, or contract-incompatible lifecycle history fails replay;
  • language mode is fixed to pure and returns unclassified with confidence 0.0;
  • the desktop API has no executor or action-dispatch method and persists no submitted text.

All 12 E043 tests pass locally. The exact implementation commit also passed all 448 tests in GitHub Actions run 31332300282 with no skips or failures. Automated admission is complete. The frozen 14-day Windows campaign has not started, so E043 is not complete and no claim of cognitive or subjective continuity is registered.

Validation

  • E043 GitHub Actions run 31332300282: exact implementation commit 99f4938 on Microsoft Windows Server 2025; 448 tests passed, 0 skipped, 0 failed
  • current E043 local Windows run: Python 3.12.13; 448 tests passed, 1 pre-existing symlink-fixture test skipped, 0 failed
  • focused E043 lifecycle and adversarial suite: 12 passed, 0 skipped
  • focused runtime, kernel, and language regression set: 47 passed, 0 skipped
  • maintained English surface and local-link checks passed locally
  • staged whitespace checks passed
  • previous GitHub Actions run 31329275166: Microsoft Windows Server 2025, CPython 3.12.10; 436 tests passed, 0 skipped, 0 failed
  • H50-L16 record: docs/v50/results/EXPERIMENT_036_FINAL_AGGREGATE.json
  • H50-L17 record: docs/v50/results/EXPERIMENT_039_FINAL_AGGREGATE.json
  • Language Conformance pure baseline: docs/v50/results/LANGUAGE_CONFORMANCE_V1_PURE_BASELINE.json
  • Experiment 041 candidate digest: e12cb042164203cbb2eee33b4dce9b5298bd39257d5cd2f6d3c0c8e651b3cf27
  • Experiment 042 pre-registration commit: 956337d

The first Experiment 041 CI run stopped before the test suite because one
Portuguese example in the otherwise English annotation guide matched the
maintained-surface checker. Commit 0ed609f reworded only that example. The
checker has passed in every subsequent Windows run.

Evidence boundary

Darwin now contains tested components that act outside language generation:
causal goal state, learned finite models, planning, replayable memory,
persistent kernel events, narrow online latent-state adaptation, and a headless
desktop lifecycle boundary. The language gateway is structurally separated
from core authority, and the pure baseline explicitly reports no language
understanding.

This does not establish consciousness, subjective experience, emotions,
personhood, open-world autonomy, AGI, language understanding, semantic fidelity,
or equivalence to Diana from Pragmata. Goals and benchmark schedules remain
evaluator-supplied, the strongest cognitive evidence is local E1, and the
desktop runtime has no resident host, tray interface, wake-word listener,
external-effect API, or completed durability campaign.

The next legitimate language step is independent human annotation. Experiment
042 fixes how those future files can fail or become calibration-only data; it
does not supply the files. In parallel, Experiment 043 may proceed only through
its unchanged 14-day real-machine campaign after CI. Neither path authorizes a
live model or mobile integration.

@DevHabito DevHabito changed the title Build the v50 evidence-first research kernel Build the Darwin v50 evidence-first cognitive architecture laboratory Aug 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant