AI systems engineer. Founder of Aweb.
I build AI systems and study how to check their work.
Telos · Repository · Interface
Research on coding agents that pass their graded tests yet fail additional checks. Current evidence is limited to a fixed task cohort.
Inbar · Repository · Interface
Research on causal diagnosis through physical tests, competing explanations, and independent adjudication. The physical-evidence gate remains blocked.
Sentinel · Repository · Interface
A monitor for frozen driving planners, evaluated in closed-loop simulation. Benchmark gains and the failed transfer are both published.
Odeya · Repository · Interface
Architecture for a research engine that separates proposals, evidence, and verification. Contracts and fixtures exist; the engine is not built.
Reiyah · Repository · Interface
An offline research engine for comparing perception systems under uncertain reference evidence. It computes decision bounds and checks the accompanying certificates.
Nisayon · Repository · Interface
An experiment engine for robot-policy integration, with fresh simulator confirmation and retained failures. Development comparisons have not demonstrated a decision or efficiency advantage.




