Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
-
Updated
Jul 6, 2026 - Python
Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
Construct-validity audit of the standard blood–brain barrier (BBB) peptide benchmark: an identity-controlled re-evaluation + shared-source provenance/overlap map, with an open, CPU-reproducible evaluation harness. Do these predictors measure penetration, or their benchmarks?
Reproducibility source, frozen configurations, manifests, and provenance for the SROP research programme.
Repository for "Testable or Not? A Pre-Registered Validity Protocol for Architecture Comparisons." Includes pre-registrations, per-seed data, analysis, and paper source.
Supplementary materials for the following publication: Davydenko, A., & Goodwin, P. (2021). Assessing point forecast bias across multiple time series: Measures and visual tools. International Journal of Statistics and Probability, 10(5), 46-69. https://doi.org/10.5539/ijsp.v10n5p46
Local-first construct and operational-definition canvas for research-methods planning.
R and Python replication of a psychometric study validating brief (18-item) versions of the MCQ and DLQ for assessing delay discounting of gains and losses (Wan et al., 2025, The Psychological Record).
Construct validity of sycophancy interventions in Llama-3.1-8B — reading a concept is not controlling it
Five falsification studies of an instrument's own limits — every figure tied to an executed measurement with confidence intervals, or marked BLOCKED. Bounded evidence lattice, critic calibration, domain transfer, verification scaling curve, Goodhart-guarded self-improvement.
Inspect eval: do LLMs writing case notes separate observation from interpretation, and can they evade a lexical validator? Deterministic grader, pre-registered rubric.
Maritime Intent Probe is a Phase 1 research programme on construct validity in neural probing. It introduces BC1 and uses a preregistered maritime routing counterexample to establish the identifiability requirements that motivate a crossed-design Phase 2 validation.
Reproducible benchmark of security-oracle construct validity on 140 real CVE fixes.
Ecological study on administrative diabetes indicators and the care cascade using NDB Open Data, Japan (335 secondary medical areas, FY2023-2024)
R script analyzing the Swahili RCADS-25 among Kenyan adolescents, assessing internal consistency, construct validity, convergent and divergent validity, and measurement invariance to evaluate the psychometric propertiesscale.
Code, per-item results and figures for a study dissociating prompt quality from response compliance in automated prompt-engineering assessment. Four-agent MATLAB evaluator on a locally hosted Qwen 2.5-7B judge, over 498 prompts from IFEval, LMSYS-Chat-1M and WildChat.
Data, code and verification gates for "The Granularity Gap" (arXiv:2606.05183): a continuous-scale audit of sycophancy across 8 Gemini variants and 8,830 responses, with 10,792 per-vote judge logs.
Guided R workflow for scale development and construct validation: item screening, factor retention, CFA, reliability, bifactor and higher-order models, measurement invariance, scoring, and theory-specified nomological networks. Every method is placed in its literature and cited.
Measurement validity, construct validity, and unsupported claims derived from AI agent telemetry.
To associate your repository with the construct-validity topic, visit your repo's landing page and select "manage topics."