Tests whether the AI Foundations principle Belonging ≠ Sameness reduces sycophantic preference-folding in repeated human–AI interaction.
-
Updated
Aug 31, 2026 - Python
Tests whether the AI Foundations principle Belonging ≠ Sameness reduces sycophantic preference-folding in repeated human–AI interaction.
Behavioral evaluation framework for sentience-, emotion-, and welfare-related AI claims, with anti-sandbagging analysis.
CCB behavioral evaluation research snapshot: nine coding-agent workflow experiments, evidence pipeline, and public web report
Open, local-first behavioral evidence and evaluation infrastructure for AI systems, benchmarks, and agents.
Replication-first study of sociolinguistic effects on LLM epistemic judgments, with behavioral validity gates before mechanistic interpretation.
Survey and reproducibility artifact for validity threats in behavioral studies of large language models
Alignment-faking and adversarial-reasoning research artifacts inspired by Redwood Research. Behavioral probes and controlled failure-mode analysis.
QLoRA post-training and evaluation for Qwen3-4B, with a public adapter, dataset, matched evaluation, CLI/API, and reproducible tooling.
Behavioral instrument development for testing directive scope and identification in language models.
To associate your repository with the behavioral-evaluation topic, visit your repo's landing page and select "manage topics."