Build the Darwin v50 evidence-first cognitive architecture laboratory - #1
Draft
DevHabito wants to merge 64 commits into
Draft
Build the Darwin v50 evidence-first cognitive architecture laboratory#1DevHabito wants to merge 64 commits into
DevHabito wants to merge 64 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
src/darwin_v50package while preserving the coupled v47-v49 prototypes as historical compatibility codedarwin-language-v1gatewayWhy
The earlier repository mixed local runtime state, large versioned prototypes,
and research claims without one maintained boundary. This change makes the
evidence trail inspectable: candidate, controls, seeds, thresholds, causal
invariants, failures, and interpretation ceilings are documented beside each
experiment.
The language boundary keeps a model replaceable. A model may propose a
candidate interpretation, phrase core-authored facts, or return external
unverified knowledge. It cannot write memory or choose identity, preferences,
goals, motivation, decisions, or RZS state through this interface.
Experiment 041 makes the external evidence gap explicit. Code can bind
annotations to a frozen digest, hide design-family metadata, require complete
panels, and calculate agreement. It cannot prove that two IDs belong to
different people or fabricate independent human judgments.
Experiment 042 prevents that future evidence from being judged with post hoc
rules. It was committed while no human annotation files, agreement values,
adjudicated labels, or real-model responses existed. A failed field blocks the
full calibration corpus; adjudication and selective case deletion cannot rescue
failed pre-discussion agreement.
Experiment 043 creates a separate engineering path while the language path is
blocked on human work. Its desktop API owns the v50 kernel but exposes no store,
executor, consent, capability, or action-dispatch handle. Every process starts
sleeping; only explicit user activation permits text; the gateway is fixed to
pure mode; and no submitted text enters the lifecycle ledger. An interrupted
restart reports an unobserved interval from the last committed event rather
than inventing the crash time.
Latest registered results
1,536tasks and passed all 17 frozen criteria.1,536tasks and passed all 20 frozen criteria; the local kernel accepted the complete conjunction.These are narrow E1 results from the repository's own unauthenticated
deterministic evaluators.
Experiments 040 through 042 are infrastructure and protocol, not registered
language results. No real language model was evaluated. The development corpus
remains author-labelled, while the candidate set has no labels at all.
Independent review, agreement, adjudication, a promoted calibration partition,
and a separate final partition remain absent.
Experiment 042 frozen rules
0.80, with higher raw thresholds where specified;NOT_PROMOTED, not a field-reduced calibration corpus;These thresholds are project gates, not universal interpretations of kappa.
Observed agreement, prevalence, per-label support, and pairwise results remain
visible alongside chance-corrected statistics.
Experiment 043 frozen gates and current status
a99575d;99f4938;unclean_unobserved, never as exact offline time;pureand returnsunclassifiedwith confidence0.0;All 12 E043 tests pass locally. The exact implementation commit also passed all 448 tests in GitHub Actions run 31332300282 with no skips or failures. Automated admission is complete. The frozen 14-day Windows campaign has not started, so E043 is not complete and no claim of cognitive or subjective continuity is registered.
Validation
99f4938on Microsoft Windows Server 2025;448tests passed,0skipped,0failed3.12.13;448tests passed,1pre-existing symlink-fixture test skipped,0failed12passed,0skipped47passed,0skipped3.12.10;436tests passed,0skipped,0faileddocs/v50/results/EXPERIMENT_036_FINAL_AGGREGATE.jsondocs/v50/results/EXPERIMENT_039_FINAL_AGGREGATE.jsondocs/v50/results/LANGUAGE_CONFORMANCE_V1_PURE_BASELINE.jsone12cb042164203cbb2eee33b4dce9b5298bd39257d5cd2f6d3c0c8e651b3cf27956337dThe first Experiment 041 CI run stopped before the test suite because one
Portuguese example in the otherwise English annotation guide matched the
maintained-surface checker. Commit
0ed609freworded only that example. Thechecker has passed in every subsequent Windows run.
Evidence boundary
Darwin now contains tested components that act outside language generation:
causal goal state, learned finite models, planning, replayable memory,
persistent kernel events, narrow online latent-state adaptation, and a headless
desktop lifecycle boundary. The language gateway is structurally separated
from core authority, and the pure baseline explicitly reports no language
understanding.
This does not establish consciousness, subjective experience, emotions,
personhood, open-world autonomy, AGI, language understanding, semantic fidelity,
or equivalence to Diana from Pragmata. Goals and benchmark schedules remain
evaluator-supplied, the strongest cognitive evidence is local E1, and the
desktop runtime has no resident host, tray interface, wake-word listener,
external-effect API, or completed durability campaign.
The next legitimate language step is independent human annotation. Experiment
042 fixes how those future files can fail or become calibration-only data; it
does not supply the files. In parallel, Experiment 043 may proceed only through
its unchanged 14-day real-machine campaign after CI. Neither path authorizes a
live model or mobile integration.