Skip to content
View manfromnowhere143's full-sized avatar

Block or report manfromnowhere143

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
manfromnowhere143/README.md

Daniel Wahnich

Daniel Wahnich

AI systems engineer. Founder of Aweb.

I build AI systems and study how to check their work.


Selected work

Telos · Repository · Interface
Research on coding agents that pass their graded tests yet fail additional checks. Current evidence is limited to a fixed task cohort.

Inbar · Repository · Interface
Research on causal diagnosis through physical tests, competing explanations, and independent adjudication. The physical-evidence gate remains blocked.

Sentinel · Repository · Interface
A monitor for frozen driving planners, evaluated in closed-loop simulation. Benchmark gains and the failed transfer are both published.

Odeya · Repository · Interface
Architecture for a research engine that separates proposals, evidence, and verification. Contracts and fixtures exist; the engine is not built.

Reiyah · Repository · Interface
An offline research engine for comparing perception systems under uncertain reference evidence. It computes decision bounds and checks the accompanying certificates.

Nisayon · Repository · Interface
An experiment engine for robot-policy integration, with fresh simulator confirmation and retained failures. Development comparisons have not demonstrated a decision or efficiency advantage.

Elsewhere

Website · Aweb · LinkedIn

Pinned Loading

  1. perceptionproof perceptionproof Public

    Reproducible study: do cheap label-free signals predict human-rated long-tail driving failure where open-loop metrics mis-rank closed-loop safety? Signals + validity statistics + tamper-evident rec…

    Python

  2. sentinel sentinel Public

    Runtime introspective safety monitor for a frozen driving planner — cutting collisions closed-loop on NeuroNCAP. Pre-registered, receipted.

    Python

  3. telos telos Public

    Evidence protocol and benchmark harness for verifying autonomous agent task completion.

    Python

  4. inbar inbar Public

    Physical causal evidence for autonomous fault diagnosis — open-world mechanism hypotheses, safe discriminating tests, independently adjudicated recovery. Authority-separated, receipted.

    Python 1

  5. parameter-golf parameter-golf Public

    Forked from openai/parameter-golf

    Train the smallest LM you can that fits in 16MB. Best model wins!

    Python 1

  6. odeya odeya Public

    Odeya: a governed, replayable research-engine architecture. Question to contract to evidence to independent verification to bounded claim. Architecture foundation only; every gate has a known-bad p…

    Python